The CRM report said 12,847 customers. The finance team insisted it was closer to 11,200. The marketing team, who had run a dedupe script last quarter, optimistically claimed 11,900. They were all wrong, and they were all right. The problem wasn’t counting. The problem was that “Acme Inc” appeared in the database as “Acme Inc,” “ACME, Incorporated,” “Acme (London),” “Akme Ink,” and ten other variants that human eyes recognized instantly and software treated as strangers.
This is entity resolution. It’s the unglamorous, unmentionable problem behind every “single source of truth” initiative that quietly fails. It sits upstream of personalization, segmentation, forecasting, and basically every data-driven decision your company makes. If your customer count is a lie, your LTV is fiction. Your churn rate is poetry. Your “AI-powered” recommendation engine is recommending to ghosts and duplicates.
Entity resolution is the process of identifying records that refer to the same real-world entity despite variation, and merging them into a canonical representation. It sounds simple. It is not. But it is solvable, with discipline, a pipeline, and a willingness to accept that 90% accuracy today beats 99% accuracy never.
The 14 Spellings of Acme Inc
I once worked with a B2B services firm that had merged two CRMs after an acquisition. The combined database held 34,000 contact records. After a naive deduplication pass (exact match on email) they declared victory at 28,000 unique contacts. Then they tried to run a campaign and discovered that 3,200 records shared no email because they were created by sales reps who typed phone numbers into the email field when they were in a hurry.
The real duplicate rate, after proper entity resolution, was closer to 22,000 unique entities. That meant their “clean” database was still 27% fiction. Their segmentation was targeting phantom companies. Their account-based marketing was treating subsidiaries as separate enterprises. Their support history was fragmented across records that referred to the same customer but never met in a query.
Those 14 spellings aren’t an edge case. They’re the norm. Human data entry is inconsistent. Acquisitions import databases with different schemas. Web forms autocapitalize, strip punctuation, or append country codes. A typical mid-size firm with 50,000 customer records should expect 8–15% duplication even after a basic dedupe. Without entity resolution, that’s not just messy data. It’s bad decisions, multiplied.
Why This Breaks Everything
Duplicate records don’t just inflate your CRM license count. They create cascading failures across the business:
Metrics lie. If “Acme Inc” and “ACME, Incorporated” are separate records, your customer count is wrong. Your average revenue per customer is wrong. Your churn calculation is wrong. I’ve seen a SaaS company panic over a “sudden churn spike” that was actually a customer record merge finally catching up to reality.
Support history fragments. A frustrated customer calls about an issue they reported last month. The agent searches by company name, finds the record, sees no prior tickets. Escalates. The customer, now furious, mentions they spoke to “Dave” three weeks ago. Dave exists in the other record. The customer experience is broken by data architecture, not by the support team.
Personalization becomes impossible. Your marketing automation sends three onboarding sequences to the same person because they signed up with a personal email, a work email, and a typo. Instead of feeling welcomed, they feel spammed. The build vs buy decision for data quality tools often starts here: someone realizes the AI can’t work until the data is clean. Tracking data quality metrics early catches these duplications before they poison downstream systems.
The Matching Pipeline
Entity resolution isn’t a single algorithm. It’s a pipeline. Skip a step and the output degrades predictably. Here’s the sequence that works:
1. Normalization. Lowercase everything. Strip legal suffixes (Inc, LLC, Ltd). Standardize punctuation. Convert “St.” to “Street.” This isn’t glamorous. It’s janitorial. It also eliminates 30–40% of your trivial duplicates before any matching begins.
2. Blocking. You can’t compare every record to every other record in a large database. Blocking creates candidate pairs using cheap heuristics: same postal code, same phone prefix, shared token in the name. A good blocking strategy reduces the comparison space from O(n²) to O(n) without missing obvious matches.
3. Similarity scoring. For each candidate pair, compute similarity on multiple fields. Name gets Jaro-Winkler. Address gets a geo-aware comparison. Email gets exact match with domain normalization. Each field contributes to a composite score. No single field decides; the ensemble does.
4. Threshold. Pairs above a high threshold (say 0.85) auto-merge. Pairs below a low threshold (0.60) auto-reject. The zone between, the ambiguous matches, goes to a human review queue. This is where most projects fail: they try to automate the edge cases instead of routing them to people.
5. Human review loop. A simple interface showing two records side by side. A reviewer clicks “same” or “different.” Their decisions feed back into the scoring model. Over time, the threshold zone shrinks. The system gets better without getting more complex. Production pipelines for data processing need this same feedback architecture to survive contact with reality.
Fuzzy Matching Without the Black Box
Every data team eventually encounters the siren song of ML for deduplication. Resist it, at first. Fuzzy string matching is powerful, explainable, and sufficient for most business data.
Levenshtein distance measures how many single-character edits turn one string into another. It’s great for catching typos: “Acme” vs. “Acne” scores 1. But it’s brittle on name reordering (“John Smith” vs. “Smith, John” scores poorly) and slow on large datasets.
Jaro-Winkler weights matching prefixes more heavily, making it ideal for names and short strings. “Martha” and “Marhta” score 0.94. It’s fast, well-understood, and available in every standard library. For 80% of entity resolution tasks, this is the workhorse.
Phonetic codes like Soundex and Metaphone encode words by how they sound rather than how they’re spelled. They catch “Smith” vs. “Smyth” and “Jon” vs. “John.” Use them as a blocking strategy, not a primary score: if two records share a phonetic code, they’re worth comparing more carefully.
Transparency is the key. When the system flags two records as potential matches, you should be able to explain why. “The model said so” is not an explanation. “Jaro-Winkler score of 0.91 on normalized names, matching address tokens, and shared phone prefix” is. Observability for data pipelines starts with this kind of explainability.
When Machine Learning Helps (and When It Doesn't)
ML enters the picture when rules stop scaling. A rules-based system with Jaro-Winkler, token overlap, and phonetic blocking typically hits 85–92% accuracy on clean domains with tens of thousands of records. That’s often enough.
ML-assisted entity resolution shines when:
At 100,000+ records, the comparison space explodes, and learned embeddings can compress similarity computation.
Labeled training data changes the game. Human reviewers have already decided on thousands of pairs. A classifier can learn the boundary between “same” and “different” better than a manually tuned threshold.
Rich features beyond text also help. Transaction history, website visits, and behavioral signals can distinguish “two employees at the same company” from “the same person with two roles.”
ML doesn’t help when you have 5,000 records and no labels. The model overfits. The complexity obscures errors. You spend three weeks tuning a classifier that performs worse than Jaro-Winkler plus common sense. Start with rules. Graduate to ML when the data justifies it.
The FDE Approach: Ship a 90% Solution
The perfect entity resolution system is a mirage. Chasing it is how projects die in month six with nothing shipped. The FDE approach is different: resolve the obvious duplicates first, build a review loop for the edge cases, and improve over time.
A typical engagement looks like this. Week one: normalize and block. Find the trivial matches: exact emails, nearly identical names with the same phone number. Auto-merge them. Show the client a before-and-after count. That early win funds the rest of the work.
In week two, add fuzzy scoring. Tune the threshold so that high-confidence matches auto-merge and the rest queue for review. Build the review interface. It doesn’t need to be beautiful. It needs to be fast: two records, a yes/no/maybe button, and a keyboard shortcut.
By week three, train the client team on the review queue. Their domain knowledge, knowing that “Acme UK” and “Acme Global” are actually separate entities, is irreplaceable. The system learns from their decisions. The queue shrinks.
At week four, 70–80% of duplicates are resolved. The remainder (the genuinely ambiguous cases, the edge cases with partial information) stay in the queue for ongoing review. The client has cleaner data, a working process, and a system that improves without requiring more engineering hours.
That’s the 90% solution. It’s not perfect. But it’s real, it’s running, and it’s making decisions better today instead of waiting for a theoretical 100% that never arrives. In the world of entity resolution dedupe business data, shipped and improving beats perfect and postponed every time.