Industry-specific operational data
E-discovery review coding decisions for document-review AI
Quick answer
E-discovery training data that actually teaches first-pass review is a set of documents paired with the reviewer's calls: responsiveness, issue tags, privilege and confidentiality designations, plus QC overturns and the review protocol that defined each tag. Public options (TREC Legal Track collections, Enron email, LegalBench) are dated or measure something else. Licensed review sets exist inside law firms, corporate legal departments and review providers, but protective orders and client rights decide whether they can be used at all.
By SourceX Editorial · Updated
What a usable review coding dataset contains
A usable dataset is the document, its family context and every coding decision made on it, tied to the protocol version in force when the call was made. A responsiveness label without the protocol is noise, because "responsive" means whatever the requests for production and the negotiated scope said in that matter. Buyers building classifiers, review agents or TAR seed models should expect these layers:
- Documents and metadata: extracted text, native type, custodian, date sent or modified, family ID (parent email and attachments), near-duplicate and email-thread group IDs, load file fields such as BegBates and EndBates if produced.
- First-level coding: responsive or not responsive, issue tags, privilege (attorney-client, work product, partial with redaction), confidentiality tier, "needs further review" flags, reviewer ID and timestamp.
- Second-level and QC results: the QC reviewer's call, whether the first-level call was overturned, and the reason code.
- Privilege log entries: the basis asserted and the description, which train privilege-log drafting as well as classification.
- Protocol artifacts: the coding manual, tag definitions, decision trees for edge cases, and change logs as the protocol evolved.
- Validation outputs: for TAR matters, the control set, the recall estimate, elusion sample results and the stopping decision.
The QC overturn layer is the most underrated. Reviewer disagreement shows where the protocol is ambiguous, and label noise matters: an audit of widely used benchmarks found an average test-set label error rate of at least 3.3% [6]. Review sets coded by contract reviewers at speed are rarely cleaner than that, so an overturn rate per tag is a quality signal you should ask for.
Why Enron, TREC and LegalBench fall short
Public legal data either predates modern review practice or measures legal reasoning instead of review decisions. The NIST TREC Legal Track began by testing whether retrieval methods help lawyers find electronic business records for civil litigation [2] and by 2009 ran Interactive and Batch tasks on responsive review of electronically stored information [1]. Those collections and their relevance judgments remain the best-known public review benchmark, and Grossman and Cormack's analysis of the 2009 Interactive data compared technology-assisted review against exhaustive manual review using recall, precision and F1 [3].
The limits are practical. The topics were mock requests, assessments were made for a research evaluation rather than a production under court deadlines, and the email era they reflect has little Slack, Teams, mobile messaging or modern attachment types. Check the license on any copy before commercial use; an audit of popular dataset hosts found license omission above 70% and license error rates above 50% [7].
LegalBench is a collaboratively built benchmark of tasks hand-crafted by legal professionals; it measures legal reasoning, not whether a reviewer would tag a given email responsive to Request No. 14 [4]. Open corpora such as Pile of Law draw on court opinions, filings, agency publications and statutes [5]. None of these contain produced or reviewed business documents with reviewer coding. For collaboration content without review labels, see the workplace email and chat datasets page.
Rights: protective orders, clients and privilege
Review coding data is usable only when the client that owns the documents, and every order governing the matter, permit reuse. Most of the documents in a review set belong to the producing party or to third parties whose communications were collected, not to the firm or vendor that coded them. Protective orders commonly restrict produced material to use in that litigation, and many require return or destruction at the close of the case.
Questions to settle before any data moves:
- Whose documents are these? A party's own collected documents differ from documents received from the opposing side, which are usually the most restricted.
- Is the matter closed? Open matters add discovery and confidentiality risk; the related question of whether licensing data creates discovery obligations is covered in our litigation discovery explainer.
- Who owns the coding? The work product (tags, privilege calls, protocols) may belong to the client, the firm or the review provider depending on engagement terms.
- What about privilege? Federal Rule of Evidence 502, enacted in 2008, limits the scope of waiver from disclosure in federal proceedings [8], but it was written for litigation productions, not commercial licensing. Treat privileged documents as excluded unless counsel concludes otherwise.
- Personal data: review sets are dense with names, emails, phone numbers, account numbers and sometimes health records; plan de-identification before delivery.
Copyright risk around legal content is live as of October 2026. The Third Circuit affirmed on September 29, 2026 that the Westlaw headnotes at issue were copyrightable and that ROSS's use of them to build a non-generative legal research tool was not fair use [9]. Review data is not headnotes, but the case is a reminder to document provenance and rights for every layer, including protocols written by outside counsel.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Specification template for a review-coding request
A good request names the review task, the labels and the rights conditions, not a particular firm or matter. Suppliers can only check fit against a clear description, and the fields below are what both sides need to assess it.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example entry |
|---|---|
| Target task | First-pass responsiveness plus 6 issue tags; privilege screen as a separate head |
| Matter types | Closed commercial contract and employment matters, US federal and state courts |
| Document mix | Email with attachments, chat exports, Office documents; families intact |
| Labels required | First-level call, QC call, overturn reason, privilege basis, reviewer ID (pseudonymized) |
| Protocol | Coding manual and tag definitions per matter, with version dates |
| TAR artifacts | Seed set, control set, recall and elusion results where available |
| Volume per matter | Full review population, not responsive-only extracts |
| De-identification | Names, emails, phones, account numbers replaced; method documented |
| Rights evidence | Client authorization, order review, confirmation the matter is closed |
| Intended use | Classifier training and held-out evaluation of review accuracy |
Two details change model quality more than volume. Ask for the full reviewed population, because responsive-only extracts destroy the negative class and inflate apparent precision. And keep at least one matter entirely held out, since tag meanings are matter-specific and random document-level splits leak protocol knowledge into test.
Evaluating models trained on review decisions
Evaluate on the metrics courts and review teams already use: recall, precision and elusion against a held-out, independently coded sample. TREC-era work established recall as the headline measure for review effectiveness and highlighted how hard it is to estimate when a review is complete [1][3]. A model that hits high document-level accuracy can still miss rare issue tags or whole families, so report per-tag recall and family-level consistency.
Common failure modes to test for:
- Family splits: an attachment coded responsive while its parent email is not, which a production cannot ship.
- Protocol drift: labels from early in a review contradicting later calls after the protocol changed.
- Privilege leakage: a responsiveness model that never learned to route probable privilege to a lawyer.
- Reviewer bias: a model reproducing one reviewer's lenient calls, detectable when reviewer ID is in the data.
For adjacent decision datasets with a similar structure, compare KYC and CDD case review files and legal billing and bill-review data; for long, multi-document test sets see long-context evaluation on real documents. The wider set of industry decision data lives in the industry-specific operational data guide, and contract-centric legal data is on the legal AI training data page.
How SourceX approaches review-data requests
SourceX sources operational datasets, including document and legal workflow records, from US companies and manages the licensing process; data is sourced on request, so a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, and the supplying company approves every release. Names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Legal teams can see the industry view on buyers in legal, or describe the review data you need.
Sourcing e-discovery training data with SourceX
Describe the review task, labels and rights conditions you need, and SourceX looks for US businesses that hold that data, assesses the data and licensing permissions, and agrees pricing and allowed uses in a license before anything is contracted. Nothing moves until a supplier agrees. Start an e-discovery data request.
Sources
- National Institute of Standards and Technology (NIST), "TREC 2009 Legal Track overview" (2009). https://pages.nist.gov/trec-browser/trec18/legal/overview
- National Institute of Standards and Technology (NIST), "TREC 2006 Legal Track overview" (2006). https://pages.nist.gov/trec-browser/trec15/legal/overview
- Richmond Journal of Law and Technology (Grossman and Cormack), "Technology-Assisted Review in E-Discovery Can Be More Effective and More Efficient Than Exhaustive Manual Review" (2011). http://jolt.richmond.edu/jolt-archive/v17i3/article11.pdf
- Stanford Hazy Research, "LegalBench: a collaboratively built benchmark for legal reasoning" (2023). https://hazyresearch.stanford.edu/legalbench
- arXiv (Henderson et al.), "Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset" (2022). https://arxiv.org/pdf/2207.00220
- arXiv / NeurIPS 2021 (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Crowell & Moring, "Congress passes new Federal Rule of Evidence to address privilege issues" (2008). https://crowell.com/en/insights/client-alerts/congress-passes-new-federal-rule-of-evidence-to-address-privilege-issues
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (3d Cir.)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.