Find possible correspondences without declaring identity
Fuzzy Match CSV Files helps compare text fields when exact-key joining cannot find every intended correspondence. A misspelled product name such as Notebok may be close to Notebook, and the tool can surface that pair for review. The result is a set of candidates and explicit user decisions. A high text-similarity score does not establish that two customers, companies, or products are the same real-world entity.
Use exact identifiers when they are available and reliable. Fuzzy matching is useful for discovering potential links that need examination, especially after a controlled cleaning step or within a clearly bounded category. It is not an automatic identity-resolution service. The default result contains only mappings you accepted. Unreviewed suggestions, rejected pairs, and records with no candidates stay separate so a later merge cannot silently inherit an unconfirmed relationship.
Import both tables and select corresponding fields
Import the source table through Choose files, then use Add files for the target table. Confirm parsing for each file independently. Select Primary table / Current data and Table B, followed by Source text column and Target text column. Headers and record counts can differ between files. Cell content remains original text, including accents, punctuation, whitespace, multiline notes, and leading-zero identifiers in other fields.
The selected text fields should describe comparable things. Comparing a company name against a complete postal address can produce a numerical score without answering a useful matching question. Preserve stable identifiers and contextual fields in both tables so you can inspect a candidate properly. Reports retain the source file and logical record location for each side. These locations are session identities, not a claim that physical line numbers or names themselves are permanent entity identifiers.
Understand exact comparison and the fixed scoring algorithm
The engine first looks for exact matches after the selected character preprocessing. When an exact name occurs in several target records, all those exact candidates are shown. It does not pick the first target or treat duplicate names as one entity. Sources without exact candidates proceed to approximate comparison within the allowed target population. Exact text candidates still require acceptance; the exact spelling of a common name alone does not prove identity.
Approximate comparison uses the versioned levenshtein-nfc-codepoint-v1 algorithm. Unicode text is normalized to NFC, then compared by code points. Insertions, deletions, and substitutions each cost one. The score is 100 multiplied by one minus the edit distance divided by the longer text's length. Notebok and Notebook differ by one insertion over eight code points, giving 87.5. The score is reproducible under the fixed preprocessing and version; it is not a percentage probability of a correct business relationship.
Choose preprocessing, thresholds, and candidate limits
Ignore case and Trim surrounding whitespace are optional. They change comparison text while preserving the original source values in reports. NFC treats canonically equivalent accent representations consistently, but it does not translate words, expand abbreviations, strip all punctuation, or apply language-specific entity knowledge. Combining marks and emoji are processed under the documented code-point convention, so a displayed symbol may occupy more than one comparison unit.
Candidate threshold ranges from zero to one hundred. Lowering it admits less similar approximate candidates; it does not automatically accept them. Fuzzy candidates per row defaults to three and can be changed within the displayed bound. Equal-scoring candidates remain distinct and follow stable source ordering rather than random selection. Duplicate exact targets are all retained even when their count exceeds the approximate-candidate display setting, because hiding an exact-name ambiguity would create a misleading impression of uniqueness.
Restrict the search with known equal fields
Add equal-field constraint selects an additional source and target field that must match exactly, such as country, product category, or organization type. Several constraints can be combined. These blocking fields reduce irrelevant comparisons and can prevent a plausible-looking name from being suggested across incompatible populations. The extra fields retain strict text comparison; a country abbreviation and a full country name must be explicitly standardized first if they are intended to match.
Empty selected names do not match other empty names. Empty blocking fields also do not create a shared identity bucket. The engine distinguishes an empty key, no candidates in the required block, and available candidates below the threshold. These are understandable outcomes, not necessarily software failures. Review the source conventions and blocking choices before broadening the search. A constraint that excludes a valid match may need a justified correction, while removing it indiscriminately can produce many irrelevant candidates.
Review, accept, reject, or manually specify a pair
Run and review opens the candidate review list. Each item shows both original names, character differences, similarity score, exact or approximate status, and source locations. Accept candidate records an explicit decision. Reject candidate records a negative decision. Undo decision removes that pair's decision. Candidate review status filters pending, accepted, and rejected pairs. After changing decisions, run again to regenerate current mapping and audit reports; the previous export remains stale until then.
Manually specify a match lets you provide the source and target record identifiers shown in reports and the target-record table. Manual acceptance can select a target outside the displayed approximate top candidates, but it must still respect empty-value and equal-field constraints. Inspect the source context before accepting it. Manual choice is recorded as a human decision; it is not disguised as an algorithmically certain match. Record identifiers belong to the current input, so a replacement file requires renewed review.
Enforce one-to-one or many-to-one decisions
Confirmation constraint defaults to One to one. Two different source records cannot both accept the same target under that setting. The review interface explains the conflict and requires undoing or revising an earlier decision. A source also cannot accept two different targets. Many to one explicitly allows several source records to share one target, which can be appropriate when consolidating known duplicate source records after inspecting their details.
For example, two source names Notebok and Notebook may both suggest the single target Notebook. Accepting the first pair reserves that target in one-to-one mode. Attempting to accept the second pair produces a conflict rather than silently replacing the earlier decision. Changing the constraint is a new rule choice that invalidates prior decisions for review. Imported contradictory or repeated decisions are also blocked by the engine, so a malformed rule file cannot create a confirmed mapping inconsistent with the candidate report.
Pass only confirmed relationships into consolidation
Main result content can be Confirmed mappings or Confirmed entity records. The mapping output contains accepted source and target record identities, original names, and source locations. The entity-record output places the accepted source records and their referenced targets into one table with a shared _confirmed_entity key. Corresponding fields are aligned by exact header name, and a target accepted by several sources is included once. Unconfirmed and rejected candidates contribute no entity records.
Apply to workflow can send that confirmed entity table to Merge Duplicate Records. Select its explicit entity key for grouping, then resolve field conflicts independently. Accepting a name correspondence does not authorize overwriting a conflicting telephone number or address. If the source and target names remain different, that field may itself require a manual choice or an all-values policy. A later required-and-unique validation step can check the consolidated result under your actual import contract.
Export decisions, preserve scope, and respect budgets
Export new copy offers confirmed mappings, confirmed entities when that main output mode is selected, candidate reports, accepted and rejected decisions, no-candidate records, and settings. Candidate scores retain their meaning and algorithm version in the report bundle. CSV and TSV output is read back before download. Formula-prefix protection remains an explicit choice. Save rules includes decisions, but Load rules deliberately clears input-specific confirmations so a similarly shaped replacement file is not automatically authorized by an old acceptance.
Candidate work is limited to fifty thousand compared pairs, five million character-work units, two hundred fifty-six code points per compared name, and the shared execution timeout. Exact-name indexing avoids unnecessary approximate comparisons where exact candidates exist. If a budget is exceeded, add reliable equal-field constraints or reduce the population; the tool does not secretly sample and call the result complete. A selected sample remains labeled as a sample. Cancel task terminates active work, input edits invalidate previous results, and refresh or closing clears the local session.
Frequently asked questions
Can I match names that are spelled slightly differently in two files?
The tool can propose text-similarity candidates for the fields you select. Use blocking fields where appropriate, review the candidate values, and confirm the intended relationship. Similarity is not proof of identity. Inspect ambiguous cases and the one-to-one or many-to-one choice before downloading confirmed matches.
How should I choose a threshold?
Use the fixed score definition and review representative candidate pairs, including likely false matches. A threshold filters text similarity; no threshold guarantees that two records describe the same entity.
Does an exact name get accepted automatically?
No. Exact candidates also need confirmation, and all duplicate exact targets are shown. A common or reused name can identify several different entities.
Can two source records point to one target?
Only when Many to one is selected. One to one blocks that conflict. A source record still cannot accept several targets in either mode.
Why did my old confirmations disappear after loading rules?
Decisions are tied to reviewed inputs and rules. Loading a template onto another input clears those input-specific approvals and requires review again, preventing an old relationship from silently moving to a new source.