Tables, time series and transactional data
Entity Resolution Training and Evaluation Data from Real Master Data
Quick answer
Entity resolution training data that holds up in production comes from real master data: customer, company, vendor and counterparty records exported with their duplicates intact, plus gold match clusters derived from merge histories and data-steward decisions. Public benchmarks such as Cora (1,879 records) are small and of uneven quality [1][2]. Specify cluster-level labels, merge provenance, entity-level train/test splits and a privacy method that preserves matching signal before you license anything.
By SourceX Editorial · Updated
Why public entity matching benchmarks stop being enough
Public benchmarks are useful for sanity checks but too small, too clean and too label-noisy to train or certify a commercial matcher. Cora, the classic citation deduplication set, has 1,879 records with a gold set of matching ids [2]. Researchers re-evaluating the benchmarks commonly used for deep learning-based matching found that their quality had largely gone unexamined, and several turn out to be easy or unrepresentative once inspected [1]. Label noise compounds this: across widely used test sets in other ML domains, Northcutt and colleagues estimated an average error rate of at least 3.3% [6], which is enough to reorder models whose F1 scores sit within a point of each other.
Real master data fails differently. A vendor master in SAP or Oracle EBS carries "ACME CORP", "Acme Corporation (DO NOT USE)" and "ACME CORP - REMIT TO" with different tax IDs, addresses and bank details. A CRM account table holds parent/child hierarchies, re-orged territories and leads converted twice. Card transaction feeds truncate merchant descriptors to "SQ BLUEBTL COF 4412" or "AMZN MKTP US2K4". Those distributions are what your customers will send you, and benchmarks rarely contain them.
What an entity resolution training record should contain
A licensable entity resolution dataset is a set of source records, a cluster assignment per record, and the evidence behind each assignment. Pairwise labels alone are not enough; entity-centric labelling, where every record is assigned to a resolved entity, produces representative and reusable benchmarks without complex sampling schemes [3]. Pairs can always be derived from clusters, but clusters cannot be reliably reconstructed from a sampled set of pairs.
At minimum, request these fields per record:
- Source keys:
source_system,source_record_id, and the table or object the record came from (for example SalesforceAccount, NetSuiteVendor, SAPLFA1/KNA1). - Matching attributes as captured: legal and trading names, addresses (raw and standardized if both exist), phone, email domain, tax IDs (EIN, VAT), DUNS or LEI where present, website, industry code.
- Temporal fields:
created_at,last_modified_at, and status flags such as inactive, blocked or "do not use". - Gold label:
entity_id(the cluster), withlabel_source(merge log, steward decision, golden-record survivorship, manual review) andconfirmed_by_role. - Merge history: the surviving record, the losing record, merge timestamp and any un-merge events. Un-merges are some of the most valuable hard negatives you can buy.
For merchant and counterparty name matching, the record is the raw descriptor plus its resolved canonical company entity. That mapping is an entity resolution task; categorizing the transaction itself into spend categories is a separate problem covered on the financial transaction data page.
Where gold labels in real master data come from
The strongest gold labels come from decisions a business already made and lived with, not from labelers looking at pairs after the fact. Merge logs in MDM hubs (Informatica, Reltio, Stibo, SAP MDG) record survivorship and steward approvals. CRM deduplication tools leave merge audit trails, and Salesforce merges keep the winning record id. ERP vendor consolidations often follow audit findings on duplicate payments, so clusters were typically checked against payment and tax records.
Each source has a characteristic failure mode you should ask about:
| Label source | Strength | Typical failure mode | What to ask for |
|---|---|---|---|
| MDM merge log with steward approval | Human-confirmed, timestamped | Auto-merge rules above a threshold were never reviewed | Flag separating rule-based from steward merges |
| CRM merge audit | Large volume, real sales context | Reps merge for convenience (territory, ownership), not identity | Merge reason codes, user role |
| ERP vendor consolidation | Verified against payments and tax IDs | Only covers vendors someone noticed | Date range of the cleanup project |
| Golden-record survivorship tables | Full cluster coverage | Clusters reflect the vendor's own matcher, so you inherit its errors | Share of clusters touched by a human |
| Migration crosswalks | Explicit old-to-new id mapping | One-to-many splits are often dropped | Unmapped and split records |
The last row matters because a CRM or ERP migration is often when a company builds a clean crosswalk; the scenario is covered in licensing data before a CRM migration. Treat any label produced by an existing matcher as weak supervision, not gold, and keep a human-confirmed subset for evaluation.
How to split and evaluate entity resolution data without leakage
Split by entity, never by pair. If records of the same entity appear on both sides of a split, a model memorizes the entity's name variants and test F1 overstates real performance. A 2026 practice paper on a self-serve ER pipeline evaluates on six deduplication benchmarks, splits at the entity level and caps training and validation at 10K records each [4]; the entity-level split is the part to copy regardless of size.
Evaluate blocking and matching separately. Blocking recall bounds everything downstream, and the WDC Block benchmark exists precisely because blocking deserves its own test [5]. Report pair completeness and reduction ratio for the candidate generator, then pairwise precision and recall plus a cluster metric (B-cubed or cluster-level F1) for the matcher. A model can score well on pairs and still produce chained clusters that merge unrelated companies through transitive links.
Also hold out by time and by source system. Train on records created before a cutoff and test on later records to mimic production drift, and keep one source system entirely out of training to test transfer. The point-in-time discipline is the same one described in point-in-time correct training data.
How de-identification interacts with matching signal
De-identification and entity resolution pull in opposite directions, because names, addresses, phone numbers and emails are the matching signal. Generic PII tools such as Microsoft Presidio detect and replace these fields, and the project itself cautions that it cannot guarantee finding all sensitive information [7]. Independent redaction of each record destroys the very variation (typos, abbreviations, transpositions) the model needs to learn.
Practical options, in rough order of fidelity:
- Business-entity records only. Company and vendor masters contain fewer personal fields than consumer customer masters; sole proprietors and contact names still need treatment.
- Consistent perturbation. Replace each real token with a surrogate through a keyed mapping so that "Jon Smyth" and "John Smith" become surrogates with a comparable edit distance. This preserves the shape of variation but must be documented and tested.
- Field-class substitution. Replace personal names with sampled names while keeping company names, which suits B2B matching.
- Controlled-environment access. Train inside an environment the supplier controls when raw identifiers cannot leave; this is a deal-structure question to raise early.
Health-adjacent records, such as patient or provider matching, fall under HIPAA, where de-identification follows Safe Harbor (removing 18 listed identifiers) or Expert Determination [8]. Safe Harbor removes most of what a patient matcher uses, so Expert Determination is the usual route when signal must survive.
Entity resolution data request template
A precise request lets a supplier check fit without exposing records first. Describe the data, the labels and the evaluation you plan, not the companies you want.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: entity resolution training and evaluation data
record_types: [vendor_master, customer_account] # B2B, US entities
source_systems: [ERP vendor master, CRM accounts]
volume_target: 200k-1M records, >= 15% in multi-record clusters
fields_required:
- source_system, source_record_id, created_at, last_modified_at
- legal_name, trading_name, address_raw, tax_id_present_flag
- status (active, inactive, blocked)
labels:
unit: cluster (entity_id), not pairs
provenance: [steward_merge, rule_merge, unmerge, migration_crosswalk]
human_confirmed_subset: required for evaluation split
hard_negatives: unmerge events, parent/child pairs kept separate
splits: entity-level; plus time cutoff and one held-out source system
privacy: personal fields replaced with consistent surrogates; method documented
quality_checks: 500-cluster audit sample; report disagreement rate
intended_use: train and evaluate blocking + matching models
Run your own audit sample before acceptance. A double-labeled sample of a few hundred clusters, scored for agreement, tells you whether the supplier's merges are trustworthy, and ISO/IEC 5259-4 offers a process vocabulary for documenting labelling quality in training and evaluation data [9].
How SourceX sources entity resolution data
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Data is not held in stock, so a request does not guarantee a match; buyers describe the data they need, and SourceX looks for US businesses that hold it. Relevant material often sits inside sales and support histories, finance workflows and document sets, such as the account data described in sales CRM histories.
Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. That last point is why the surrogate strategy above belongs in the request from day one; you can start one at sourcex.si/buyers.
For neighboring problems, see the structured data buyer's guide, B2B relationship graphs from vendor and spend records, ERP transaction and master data for AI training and building a golden evaluation dataset from business records. Product matching on catalog identifiers and attributes is a different signal set, covered under product attribute data, and entity resolution should not be confused with named entity recognition, which finds mentions in text rather than merging records.
Request entity resolution training data
If you need real master records with merge-derived match clusters, describe the record types, label provenance and privacy method you require. SourceX assesses data and licensing permissions with candidate suppliers, and nothing is contracted until a supplier agrees. Submit a buyer request at sourcex.si/buyers.
Sources
- Université Paris Cité (ICDE 2024), "ICDE 2024 paper re-evaluating benchmark datasets for learning-based entity matching" (2024). https://helios2.mi.parisdescartes.fr/~themisp/publications/icde24-dlmatching.pdf
- cleanzr, R-universe, "Cora: duplicate citation records (R package manual)". https://cleanzr.r-universe.dev/cora/doc/manual.html
- arXiv, "How to Evaluate Entity Resolution Systems: An Entity-Centric Framework with Application to Inventor Name Disambiguation" (2024). https://arxiv.org/pdf/2404.05622
- arXiv, "Entity Resolution in Practice: Lessons from a Self-Serve Pipeline" (2026). https://arxiv.org/pdf/2607.26298
- University of Mannheim, Data and Web Science Group, "WDC Block: A large Blocking Benchmark released". https://www.uni-mannheim.de/dws/news/wdc-block-a-large-blocking-benchmark-released/
- Northcutt, Athalye, Mueller (arXiv / NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Microsoft (presidio project), pkg.go.dev, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- U.S. Department of Health and Human Services, OCR, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and ML - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.