Deduplicate

Find and remove duplicate records with confidence.

Group matching records in your CSV or Excel data, decide which record to keep and review proposed removals. Prepare a cleaner import file without changing your original source.

  • CSV and XLSX input
  • Review before removal
  • Browser-local processing

Duplicate review workspace example

Example duplicate review, not a live dataset. Normalized Email and Company matches form 298 groups containing 894 rows. Keeping one row per group proposes 596 removals. No changes have been applied.
RowDesk

Duplicates

Find and remove duplicate records from your dataset.

customer-export.csv12,482 rows 5 columns
298duplicate groups
894duplicate rows
596rows proposed for removal

Match on these fields

EmailCompanyAdd field

Matching options

Ignore capitalization

Trim surrounding spaces

Blank keys do not match.

Keep which row?

Keep most complete (recommended)

Keeps the row with the most populated fields.
You can change this before applying.

Duplicate groups (298)

Review each group and confirm which row to keep. 596 rows are proposed for removal.

Group 13 recordsWhy this group?
These records match after trimming and ignoring case in Email and Company. Completeness counts the five displayed business fields.
SelectionRowNameEmailCompanyPhoneLocationCompletenessAction
108Sarah Johnsonsarah@acme.comAcme Inc.(555) 010-2200Boston, MAKeep (recommended)
245Sarah J. JohnsonSARAH@ACME.COMAcme Inc.EmptyEmptyRemove
617Emptysarah@acme.comACME INC.EmptyEmptyRemove

Keeping Row 108 because it contains 5 populated fields compared with 3 and 2 in the other records.

Group 24 records
Group 32 records
Remove 596 duplicates
How matching works

Your columns.
A clear matching rule.

Define what makes two records duplicates in this file. RowDesk applies that rule consistently, so you can trace each group back to the fields you selected.

Choose the matching columns

Use one field, such as Email, or a combination such as Email and Company. Every selected column must match under your configured rule for records to be grouped.

Set conservative matching options

Use exact matching, or configure normalization to ignore capitalization and trim surrounding spaces. Similar-looking names and approximate spellings do not become matches.

Check the groups, not just the count

A match follows your selected keys, not a guess about identity. Blank key values do not match each other by default. Review the other fields before deciding what belongs in your output.

Same email, different formatting

With surrounding-space trimming and case-insensitive matching enabled:

" SARAH@ACME.COM "matchessarah@acme.com

Normalization is used for comparison. It does not combine field values or establish that an email address is valid.

Keep most complete

Keep the fuller record.
Make the final call.

RowDesk can recommend the record with the most non-empty values across your output columns. Compare the fields behind that recommendation, then keep it or manually choose a different record.

Completeness measures populated fields, not whether those values are correct. When counts tie, the first occurrence wins unless you choose otherwise. Keep First and Keep Last are also available rules.

Group 1: compare the recordsEmail + Company, ignoring capitalization
Illustrative group from the preview above
Three matching records compared across five output fields. Row 108 has five populated fields; rows 245 and 617 have three and two.
Output fieldRow 108Keep (recommended)Row 245Remove from outputRow 617Remove from output
NameSarah JohnsonSarah J. JohnsonEmpty
Emailsarah@acme.comSARAH@ACME.COMsarah@acme.com
CompanyAcme Inc.Acme Inc.ACME INC.
Phone(555) 010-2200EmptyEmpty
LocationBoston, MAEmptyEmpty
Completeness5 / 5 fields3 / 5 fields2 / 5 fields
Why this record?

Row 108 has all five output fields populated, including Phone and Location. The other records have three and two populated fields. RowDesk recommends keeping Row 108 as a whole record; it does not merge values from the other rows.

Review Keep / Remove decisions

Compare each group side by side. A fuller record is a starting point, not a decision you have to accept. Choose the row that belongs in your prepared file.

Make a “Not duplicates” exception

A shared key does not always mean the records represent the same thing. Mark the group as Not duplicates to retain its records instead of applying the proposed removals.

Preview before removal

See what stays.
Know what leaves.

Before you remove duplicate CSV records, check the number of groups, the records involved and the rows proposed for removal. These are different counts. Nothing is removed from the working dataset until you confirm.

298

Duplicate groups

Sets of records matching the selected keys.

894

Records involved

All rows in those groups, including the rows to keep.

596

Proposed removals

Keeping one record per group removes 894 minus 298 rows.

In the illustrated preview: 12,482 input rows minus 596 proposed removals leaves 11,886 output rows. Keeping a different row changes which record survives; keeping an entire group changes the removal count. Review the updated summary before applying.

Review and undo

A removal you can review.
A step you can undo.

Applied duplicate removal becomes a step in RowDesk Changes. Inspect the transformation summary and removed rows, then continue preparing your file with a clear record of what happened in the session.

Explore change review and session undo

Inspect the removed records

Review the rows removed from your working dataset. The removed set is preserved for a separate export when you need to check what was excluded.

Undo within the active session

Undo the most recent transformation while the session remains active. If you have applied later steps, undo those first to return to the state before deduplication.

Keep your source file intact

Removal affects the working dataset and its output, not the original CSV or XLSX file. Export a separate prepared CSV when your review is complete.

Prepare for import

Fewer repeated records.
A clearer import file.

Repeated contacts, companies and operational records can make the next import harder to review. Deduplicate CSV data or clean up Excel duplicates before handing the file to your business system.

Use keys that fit the job

To find duplicate contacts, review an email-based rule. For company records, use your chosen company identifier. Select the columns that make sense for your data rather than treating every repeated name as a duplicate.

Check incoming lists against existing data

Dedupe finds repeated records within one dataset. To check a new list against an existing CRM export, use Compare files. Export your prepared CSV and import it through the destination system's own workflow.

A cleaner file helps reduce repeated data; it does not guarantee that a downstream system will accept every record.

Processed locally in your browser. In local-processing mode, your raw dataset is not uploaded to RowDesk servers.

About local processing
The next step

Keep the right records.
Leave the duplicates behind.

Choose your matching rule, review the groups and export a cleaner CSV.