Define duplicates with your fields
Duplicates are not always identical rows. Group by an SKU, email, or a combination of fields to reflect what makes a record unique in your work.

Bring duplicates together and make each decision clear. Group matching records, inspect conflicts, and choose what to keep.
Import data, confirm the rules, review the result, and export a new copy.
Drag and drop a CSV, TSV, or delimited text file
UTF-8 · Up to 10 MB per fileExplore the available features and choose what your table needs next.
Define duplicate groups with one field or a combination of fields.
Use in the toolConfigure case and surrounding-whitespace handling for matching.
Use in the toolKeep the first record by default, keep the last instead, or mark duplicates without removing them.
Use in the toolInspect differing field values within records that share a matching key.
Use in the toolRemoved records remain available in a separate report for review or export.
Use in the toolApply the result and continue with mapping, filtering, or comparison, with undo available.
Use in the toolConfirm the input, inspect changes, and save the output you need.
Choose a CSV or TSV file, or paste a table. Check the delimiter, headers, and preview before confirming the full import.
Choose one or more fields that define a duplicate, then set the keep policy and comparison rules. Review groups, retained records, and removals separately.
Choose the result or report and CSV or TSV format. Review the export summary and protection count, then confirm download or copy.
Removing duplicates from CSV starts with a decision about identity. Two rows can be identical in every cell, or they can share an identifier while disagreeing about a description or price. Whole-row comparison addresses repeated copies of a record. Selected-field comparison groups records by one or more chosen fields. Neither choice establishes which business record is correct. TableWorkbench makes the grouping and retention policy visible so you can inspect disagreement before accepting a smaller result.
Import the file and confirm its parsing settings. Pay particular attention to identifier columns, because a delimiter mistake or an incorrect header changes what you are comparing. All imported values remain text, so 001 and 1 are different by default. Choose whole-row comparison when every field must match. Otherwise select the fields that jointly identify a record. A combination such as SKU and Region can distinguish valid regional rows that would incorrectly collapse if SKU were used alone.
Keep first retains the first occurrence in the input order. Keep last retains the last occurrence in that same order. Last does not mean the most recent date, the highest revision, or the most authoritative source. If file order is not meaningful, begin with Mark only. It produces duplicate groups without removing any records, letting you examine conflicting fields and decide whether a later preparation step is necessary.
The result preserves the relative order of surviving records. When the final occurrence of a key is retained, that actual occurrence stays in its position relative to other survivors. The tool does not move it to where the first occurrence used to be. This is important when record positions provide useful context. You can later apply a separate stable sort if a different presentation order is required, with its own visible rules and reversible workflow step.
Ignore case and Trim surrounding whitespace are off until you choose them. They alter the comparison key for grouping, not the values retained in the output row. Thus “Blue” and “ blue ” may belong to one duplicate group under both options, while the chosen survivor still contains its original spelling and spacing. If you want those values rewritten as well, use a separate cleaning step and inspect its changes before deduplication.
Selected keys containing an empty or missing component are placed in the review report and retained. They are not all treated as the same person, product, or account. The text zero is a populated value. The strings NULL and N/A are also populated values unless you deliberately clear them earlier. Whole-row mode has a different purpose: a complete row of empty fields can be an exact repeated row. Check which mode is selected when reasoning about blank records.
Composite keys are encoded as separate values, rather than concatenated with a separator that might also appear in the data. A pair containing “a|b” and “c” remains different from a pair containing “a” and “b|c”. You do not need to create a fragile helper column by joining text yourself. Select all relevant fields directly and keep their meaning consistent across the file. The selected combination is evaluated against the complete dataset, not just the rows displayed in the grid.
Consider five records: SKU 001 with price 12.00, SKU 002 with price 8.00, SKU 001 with price 12.50, and two records whose SKU is empty but whose descriptions differ. Group by SKU and use Keep first. The two 001 records form one duplicate group involving two records. The first 001 survives and the later one appears in the removed-record report. SKU 002 survives. Both empty-key rows remain and are listed for review.
The duplicate-group report shows the complete original rows, a group identifier, a decision, and source information. The price disagreement therefore stays visible. The tool does not average 12.00 and 12.50, choose the larger number, or fill fields from different rows to fabricate a combined record. With Keep last, the 12.50 row survives because it occurs later. With Mark only, both price records remain in the result. Those three outcomes answer different questions and should never be described as equivalent cleanup.
Duplicate groups counts the keys that have more than one occurrence. Involved records counts all records in those groups, including survivors. Removed records counts only the rows actually excluded by the selected retention rule. These numbers need not be equal. One group of four records involves four rows and removes three under Keep first, while removing none under Mark only. Review records with missing keys are a separate category and should not be folded into removal totals.
After running, use the result tabs to examine the retained table, duplicate groups, removed rows, and empty-key review. A view search helps find a particular identifier without changing the underlying result. Open export and explicitly choose the scope. A ZIP can keep the final table and supporting reports together, with a file manifest. Spreadsheet protection applies during export, and the verification stage rereads the serialized values before allowing the download. Keep reports with the result when someone else must understand the decision.
Deduplication often follows appending files from several sources. Preserve source information during merging so the duplicate report explains where conflicting records came from. Clean only the fields that require normalization and avoid changing identity semantics casually. Once you accept the deduplicated table, apply it to the workflow. The comparison tool can then use the retained identifiers to compare snapshots, while the column mapper can reshape the table for a destination system.
Undo restores the previous workflow state by replaying accepted steps from the preserved source data. Disabling a step lets you inspect the effect of omitting it. If a later step references a column removed by an earlier change, the workbench blocks that invalid reference instead of quietly selecting another column. Rule files can help repeat a policy on files with matching schemas, but loading one still requires checking the incoming headers and reviewing the actual groups it creates.
Choose the column that defines identity, inspect the duplicate groups, and select a retention policy such as keeping the first or last occurrence. Other fields can differ within a group, so review those conflicts. Keeping the last row means last in the current order, not automatically the most recent date.
No. The available first and last policies refer only to file order. A timestamp-based policy would require an explicit timestamp field, a defined date format, and a decision for ties or invalid dates. Those semantics are not part of this tool. Use Mark only when you need to inspect which occurrence should be considered authoritative.
An empty identifier provides no evidence that two records describe the same entity. Selected-key mode retains them and puts them in the review report. This avoids deleting unrelated rows simply because both are incomplete. Resolve the missing identifiers at the source or consciously choose a different complete key before running again.
Yes. Ignore case normalizes the key used to find groups, while the retained record keeps its stored text. Trimming for matching follows the same principle. To change the stored spelling or spaces, apply an explicit cleaner step. Keeping comparison semantics separate from text transformations makes the result easier to explain.
It evaluates every imported record within the supported limits. The grid renders pages for readability and reports the displayed range separately from the full total. Searching that grid does not narrow the deduplication operation or the default export. Choose the export scope explicitly when you want a report instead of the retained dataset.
It does not perform fuzzy matching, infer that similar names refer to the same person, validate email ownership, merge conflicting fields, or guarantee acceptance by another system. It also does not process XLSX workbooks in this version. Its job is to group exact or explicitly normalized keys, expose the evidence, and apply the retention policy you chose.
Start with a sample to learn the workflow, then process your own files.