BROWSER-BASED TABLE TOOL

Merge Duplicate Records

Merge duplicate rows by exact keys, fill complementary fields, review conflicts, and retain field provenance.

File contents are processed locally in your browser. Refreshing or closing clears the session. The website still requests static resources.

1Import data—2Configure—3Review—4Export

Drop your table here

Drag and drop a CSV, TSV, or delimited text file

UTF-8 · Up to 10 MB per file

Start with your data

File content is processed in your browser, without uploading it.
Kept in page memory · Continue across tools · Refreshing or closing clears the sessionCSV / TSV · UTF-8
Sample library · Normal / Boundary / Error

Clear this session?

Imported data, steps, and results will be released. Export any needed copies first. Original files are unaffected.

Consolidate useful fields instead of discarding rows

Merge Duplicate Records combines complementary information from records that share an explicit exact entity key. One record may contain a phone number and another an address. Ordinary deduplication keeps or removes whole rows, potentially losing a useful field from the discarded record. Consolidation makes a decision for each field, reports conflicting nonempty values, and preserves the source of retained information.

The tool does not infer that similar names identify the same person or organization. Choose a reliable identifier, or use a previously confirmed cross-table entity mapping as a deliberate prerequisite. Two people can share a name, and one person can have several name spellings. A fuzzy candidate score is therefore insufficient to establish a duplicate group. This tool starts after the grouping relationship is explicit and then handles the separate question of which field values should survive.

Select exact keys and retain unidentifiable records

Choose files or Paste a table and confirm its parsing settings. Select Exact entity keys, using several fields when one identifier is only unique within a region, account, or source system. Key components are compared as original text with preserved boundaries. The values 001 and 1 remain different, and embedded punctuation cannot create a collision through naive concatenation. No automatic trimming or case folding is performed during entity grouping.

Records with an empty key component do not enter automatic duplicate groups. They are retained separately and included in Unmerged records, preventing unrelated unidentified records from being combined. The empty definition distinguishes missing cells, empty strings, whitespace, and explicit markers. Missing cells and empty strings are empty initially; zero, false, NULL, and N/A are not universally discarded. Confirm the source convention before adding a marker that might also be a legitimate identifier or field value.

Start with nonconflicting nonempty completion

The default Field retention policy is Fill only nonconflicting values. If a group contains one nonempty phone number and another empty phone field, the known number can fill the result. If several records contain that same nonempty phone number, it remains a single consistent value. If all observations are empty, the result preserves an empty source state rather than inventing a value. Other fields are resolved independently from their own available observations.

Take three records: identifier 001 with phone 123 and no address; identifier 001 with no phone and address East; and identifier 002 with phone 456 and address West. Selecting id as the key produces two main records. The first contains 001, 123, East, and the second contains 002, 456, West. The resulting records still cover all three input sources. Phone and address in the first result correctly cite different donor records rather than claiming that one source supplied the entire consolidated record.

Distinct nonempty values always require a policy

When a group contains phone 123 and phone 456, both nonempty values are retained in the conflict report. The default leaves the field unresolved and prevents applying or exporting the main table as a complete confirmed result. The preview remains available for investigation, and review reports can still be downloaded. A successful-looking main download is not offered merely because the engine was able to construct a provisional row.

Review conflicts manually lets you choose the specific source record to retain for that field. The review panel shows the exact group, field, distinct values, and candidate source locations. Choose the source record to retain records a field-level decision. Changing a related policy or input makes the old confirmation require review again. Selecting one phone does not decide an address conflict elsewhere in the same group, and accepting a group relationship does not authorize arbitrary field overwrites.

Declare source priority explicitly

Explicit source priority can use file names or a selected Priority source column. Enter Priority order with the highest-priority value first. Every relevant donor must have a known rank; an unlisted source cannot silently become the default winner. If several best-ranked donors contain different nonempty values, the priority is tied and the conflict remains unresolved. Source priority expresses your knowledge about data authority, not an assumption that the file read last is newest or most correct.

Run first to inspect conflict counts and retention reasons. The review panel then offers an explicit confirmation of the selected priority before another run can produce a resolved main result. This two-stage process ensures that an automatic policy is reviewed against the actual conflicts it will affect. Keep the conflict and field-provenance reports with the result when another reviewer needs to know why one source won. A priority list without that audit can conceal a consequential choice behind a simple completed status.

Use a date column only when it means recency

Explicit latest date requires a Priority date column and the shared confirmed date-format semantics. File order and logical record position are not substitutes for event time. Calendar dates, local clock readings, and offset timestamps must be interpreted consistently before ordering. Invalid, empty, unsupported, or ambiguous priority dates leave the affected conflict unresolved. A tied latest date with different nonempty field values also requires review instead of an arbitrary first-row choice.

For a phone field with values 123 dated January first, 456 dated January second, and an empty phone dated January third, the latest empty observation does not erase the known phone. Eligible donors are nonempty observations for that field, so the selected phone is 456 when all required dates are valid and the policy has been reviewed. If the source uses an empty field to mean an explicit deletion, that is a different business rule and is not inferred by this nonempty consolidation policy.

Preserve several values reversibly or keep a group separate

All distinct values retains every different nonempty value in a JSON array stored as text in the result cell. For phone values 123 and 456, the output is an array containing those two strings. JSON quoting and escaping preserve boundaries when a value itself contains a comma, quotation mark, line break, or backslash. This is intentionally different from joining with an ambiguous delimiter that could not later be split back into the original values.

A field using this policy may need special handling by the receiving system, since a JSON array inside a CSV cell is still one text field. Review the destination's expected shape before selecting it. Keep this group separate is another valid decision: it retains every original record in the group instead of merging them. Undo all keep-separate decisions restores those groups for review. Choosing not to merge can be safer and more accurate than forcing an unresolved business conflict into one row.

Inspect field provenance and group outcomes

The main preview reports total conflicts, unresolved fields, and groups deliberately kept separate. Field provenance lists the output record, field, retained or reviewed source value, donor record, source file and logical position, and retention reason. Multiple retained values preserve all contributing sources. Even when several input records agree on a value, the source records remain traceable rather than being lost when the result collapses to fewer rows.

Open source links to inspect the underlying group. The source grid keeps complete original records, including unrelated notes or metadata that may help a human decision. A logical record can include quoted multiline text, so its location is not necessarily the same as a physical file line. Search and pagination change only what is visible in the grid. They do not exclude an unseen conflict from the full run or authorize export before the unresolved count reaches zero.

Continue from reviewed matching and validate the result

Fuzzy Match CSV Files can produce Confirmed entity records containing accepted source-target relationships and a shared _confirmed_entity key. Choose that explicit key here to consolidate the confirmed population. Fields align by their exact header names in that prerequisite table. Different source and target names can still become a field conflict, which you must resolve independently. Rejected and pending fuzzy candidates are absent from the entity input and cannot become automatic duplicates through this workflow.

After a completed consolidation, use CSV Data Validator to check the destination's required fields and uniqueness rules. A merged record can still lack a mandatory address if none of its sources supplied one. Likewise, deliberately retained separate groups can remain duplicated under a uniqueness requirement. These are legitimate review outcomes, not reasons to manufacture data. Keep the consolidation rules and validation report together when explaining the final import-ready population to another reviewer.

Export only the scope you actually resolved

Export new copy offers the main table only after blocking field conflicts are resolved. While conflicts remain, Unmerged records, Field conflicts, Field provenance, and Settings are available as review material; a ZIP that would include an unresolved main table is blocked. Completed exports use the shared CSV or TSV serializer and are read back before download. Formula-prefix protection is explicit and may alter formula-like text, including valid negative or plus-prefixed values, so review its impact.

Save rules preserves policies and field bindings. Input-specific manual approvals are cleared when loading a rule file onto another input, requiring renewed review. Apply to workflow supports Undo and later tools; a sample result cannot replace the complete workflow table. Large provenance reports can exceed audit limits even when the source file fits the import limit. Cancel task retains the source and applied steps. Refresh or closing clears the local session, so export needed results and reports before leaving.

Frequently asked questions

How do I combine duplicate contacts when one row has a phone and another has an email?

Group the records by a confirmed exact identifier and choose the field-level consolidation rules. Complementary fields can be brought together, while competing nonempty values require a policy or review. Similar names should not become an automatic duplicate key. Inspect provenance to understand which source supplied each final field.

How is this different from removing duplicate rows?

Removing duplicates chooses whole records to keep or discard. Consolidation chooses values field by field, allowing complementary phone and address information to survive in one result.

What happens when both records contain different phone numbers?

The field becomes a conflict. Choose a source manually, use a reviewed explicit priority, retain all distinct values as a JSON array, or keep the group separate.

Does the last record count as the latest?

No. Latest-date priority requires a selected date column with valid, explicit interpretation. Equal dates or invalid dates cannot silently resolve conflicting values.

Can a similar name be used as an automatic duplicate key?

No. Use a reliable exact key or a previously reviewed and accepted entity mapping. Text similarity alone does not establish identity, and unconfirmed candidates must not enter consolidation.