Entity matching is the process of determining whether two or more records, across separate datasets, refer to the same real-world object. It answers one question: is “Acme Corp., 123 Main St” in your CRM the same entity as “ACME Corporation, 123 Main Street” in your ERP? Without a shared identifier, the answer requires algorithms. This article explains how those algorithms work, where they break, and how to configure a pipeline that scales.
Start a free trial of Match Data Pro and run your first entity matching job in minutes.
What Is Entity Matching and Why Exact Joins Fail
Exact matching works when every record uses the same unique identifier: a Social Security Number, a DUNS number, a UUID assigned at creation. In practice, that condition holds for maybe 30–40% of enterprise data. The rest arrives from systems that never shared a key. A customer record ingested from a web form, a vendor record keyed manually in an ERP, and a contact enriched by a third-party provider all describe the same company but carry none of the same identifiers.
The result is duplicate and fragmented records. A company with five offices might appear as fourteen distinct entities across your data warehouse. A customer who changed her name at marriage becomes two separate contacts in your CRM. A vendor registered under three subsidiary names maps to three separate spend buckets in your procurement system.
Exact joins miss every one of these. Entity matching algorithms do not.
Deterministic vs. Probabilistic Matching
The two fundamental approaches to entity matching are deterministic and probabilistic. Deterministic matching applies explicit rules: if two records share the same tax ID and the same phone number, they match. It is fast and fully auditable, but brittle. One transposed digit in a phone number breaks the rule.
Probabilistic matching assigns a weighted similarity score across multiple fields. A name similarity of 0.92 plus an address similarity of 0.88 plus a phone similarity of 0.60 produces a composite score of, say, 0.85 — enough to route the pair to a confirmed match or a review queue. Most production pipelines combine both: deterministic rules handle the easy cases, probabilistic scoring handles the ambiguous ones.
The Four Core Algorithms Used in Entity Matching
No single algorithm solves every field type. The table below shows which algorithms apply to which data.
| Algorithm | Best For | Example | Typical Threshold |
|---|---|---|---|
| Levenshtein (edit distance) | Short strings, codes, names | “Smith” vs “Smyth” → 0.80 | 0.80+ |
| Jaro-Winkler | Names with prefix similarity | “Johnson” vs “Johnston” → 0.93 | 0.88+ |
| Phonetic (Soundex, Double Metaphone) | Names with spelling variants | “Schmidt” vs “Schmitt” → match | Exact phonetic key match |
| Token-based (TF-IDF, Jaccard) | Company names, addresses | “Acme Corp Ltd” vs “ACME Corporation” → 0.85 | 0.75+ |
The choice of algorithm matters as much as the threshold. Levenshtein on a company name field will penalise word-order differences that token-based methods handle well. Jaro-Winkler on an address field will over-reward records that share a street number prefix. Applying the right algorithm to the right field type cuts false positives and false negatives significantly.
Composite Scoring: Combining Algorithms Across Fields
Production entity matching uses composite scoring. Each field gets its own algorithm and weight. A name field might carry a weight of 0.40, an address field 0.30, a phone number 0.20, and an email domain 0.10. The weighted average of per-field scores produces a single match probability for the pair.
Match Data Pro’s AI-powered fuzzy matching engine lets you configure field weights and algorithm selection per field type. The platform surfaces AI-suggested match pairs with confidence scores, so you can see exactly which fields drove a match decision before approving it. That transparency is what separates configurable platforms from black-box tools.
The Entity Matching Pipeline: Six Stages
A reliable entity matching pipeline runs six stages. Skipping any one of them reduces recall or precision.

Stage 1: Data Profiling
Before any matching runs, AI data profiling scans every field for completeness, format consistency, and value distribution. A field with 40% null values needs different handling than one with 99% population. Profiling surfaces these gaps so you can weight incomplete fields lower or exclude them from scoring entirely.
Stage 2: Standardisation
Raw records arrive inconsistently formatted. “Street” appears as “St”, “St.”, and “Street”. Company suffixes appear as “Inc”, “Inc.”, “Incorporated”, and “Corp”. A standardisation pass normalises these variations before scoring so that two records representing the same entity score high rather than being penalised for formatting differences alone.
Match Data Pro’s data cleansing and standardisation engine handles company suffix normalisation, name parsing (first/last/middle), address component parsing, and phone number formatting automatically as a pipeline step.
Stage 3: Blocking
Comparing every record against every other record in a dataset of one million rows means 500 billion candidate pairs. That is computationally impossible at scale. Blocking reduces the candidate space by grouping records that share at least one attribute — the first three characters of a name, the zip code, the phone area code. Only records within the same block are compared.
A well-designed blocking strategy retains 99%+ of true matches while eliminating 95%+ of non-matches from the comparison set. Match Data Pro uses adaptive blocking that combines multiple blocking keys, reducing missed matches caused by single-attribute blocking failures.
Stage 4: Multi-Algorithm Scoring
Each candidate pair inside a block is scored across all configured fields using the appropriate algorithm per field. The result is a composite score between 0 and 1. Records with a score above the auto-match threshold are linked immediately. Records in a middle band go to a review queue. Records below the no-match threshold are kept separate.
Stage 5: Survivorship Rules
When two records match, you need one canonical golden record. Survivorship rules determine which field value wins. Common rules include: most recently updated, most complete, most frequently occurring, or source-system priority. Match Data Pro lets you configure field-level survivorship rules independently — so the CRM wins on contact name, but the ERP wins on billing address, without either record overwriting all fields from the other.
Stage 6: Golden Record Output
The output is a clean golden record — one unified entity per real-world object, with a lineage trail showing which source records contributed each field value. This record flows downstream to your CRM, ERP, data warehouse, or analytics layer via Match Data Pro’s import/export connectors.
Real-World Entity Matching Challenges
Several categories of data create persistent matching difficulty.
Name Variants and Nicknames
A person named “William” may appear as “Bill”, “Will”, or “Billy” across source systems. Standard edit-distance algorithms will score “William” vs “Bill” poorly — 0.44 on Levenshtein. Phonetic algorithms and nickname lookup tables solve this. Match Data Pro applies configurable nickname expansion as a pre-processing step before scoring.
Transposed Digits and OCR Errors
A phone number of “555-1234” versus “555-1243” has a Levenshtein score of 0.88 — a transposition of two adjacent digits. A ZIP code of “90210” versus “90201” looks like a non-match on exact comparison but scores 0.89 on normalised edit distance. Entity matching pipelines need digit-transposition tolerance built into numeric field scoring.
International Address Formats
Address standardisation that works for US records breaks on international formats. A UK postcode appears after the city. Japanese addresses run province-first. German street numbers come after the street name. Address matching at scale requires format-aware parsing, not a single regex pattern.
Company Name Variations
A single company may appear as “Acme Corp”, “Acme Corporation”, “ACME Corp.”, “Acme Corp Ltd”, and “Acme (UK) Ltd” across datasets. Token-based algorithms with stopword removal — stripping “Ltd”, “Inc”, “Corp” before scoring — significantly improve recall on company name matching.
Entity Matching vs. Entity Resolution: The Distinction
Entity matching determines whether two specific records refer to the same object. Entity resolution goes further: it groups all matching records across an entire dataset into clusters, one cluster per real-world entity. Entity resolution is entity matching applied at dataset scale, with graph-based clustering to handle transitive matches.
If Record A matches Record B, and Record B matches Record C, entity resolution knows that A, B, and C are the same entity even if A and C do not score above the match threshold directly. Match Data Pro integrates Senzing entity resolution — a probabilistic, graph-based engine that handles exactly this transitive linkage problem at millions of records per hour.
Ready to see entity matching in action on your own data? Book a demo with the Match Data Pro team or start a free trial to run a matching job against your own records today.
Frequently Asked Questions
What is entity matching in data management?
Entity matching is the process of identifying records in separate datasets that refer to the same real-world object — a person, company, or address — without relying on a shared unique identifier. It uses similarity algorithms to score record pairs and produces linked or merged records, enabling accurate analytics and a clean master dataset.
What algorithms are used for entity matching?
The most common algorithms are Levenshtein edit distance for short strings and codes, Jaro-Winkler for names, phonetic algorithms such as Double Metaphone for name variant detection, and token-based methods such as Jaccard similarity and TF-IDF cosine for multi-word fields like company names and addresses. Production pipelines use combinations of these with per-field weights.
How is entity matching different from deduplication?
Deduplication removes duplicate records within a single dataset. Entity matching links records across separate datasets that represent the same real-world entity. Deduplication is typically a step within an entity matching or entity resolution pipeline, not a substitute for it when working across multiple source systems.
What is a blocking strategy in entity matching?
Blocking reduces the number of record pairs that need to be scored by grouping records that share at least one common attribute — such as the first characters of a name, a postal code, or an area code. Only records within the same block are compared, cutting the candidate pair space from billions to millions without significantly reducing match recall.
How do I know if my entity matching thresholds are correct?
Threshold calibration requires labelled ground truth pairs — records you know are matches or non-matches. Run your scoring pipeline against those pairs, then measure precision and recall at different threshold values. A threshold producing 95% precision with 90% recall is a reasonable starting point for most enterprise datasets. Adjust field weights first before changing the global threshold.