Evaluation and benchmarking datasets
Contract review evaluation datasets: clause-level gold labels from real redlines
Quick answer
A contract review AI evaluation dataset is a held-out set of real agreements with lawyer-verified answers: clause spans and types, key-term values, playbook deviations and, for redlining, the markup a lawyer actually sent and the text finally signed. Public benchmarks such as LegalBench and CUAD test clause recognition and legal reasoning on published text. A private set built from negotiated drafts tests what a deployed reviewer faces: counterparty paper, your playbook, multi-turn redlines and unseen documents.
By SourceX Editorial · Updated
What LegalBench and CUAD measure, and where they stop
Public legal benchmarks measure whether a model can classify clauses and reason about published legal text, not how it reviews a live draft against one company's negotiating positions. The 2023 LegalBench paper describes 162 tasks across six types of legal reasoning, drawn in part from earlier sets including CUAD and ContractNLI, and evaluates 20 models [1]; the project describes its English-language tasks as hand-crafted by legal professionals [2]. A 2025 review describes CUAD as a pre-LLM benchmark with more than 13,000 expert annotations across 41 clause types [3].
What they leave out:
- Negotiation. Labels sit on finished agreements, so nothing records who proposed a clause or which positions were conceded.
- Your playbook. A cap of three months' fees suits one customer and is a walk-away term for another.
- Messy inputs. Real review runs on Word drafts with tracked changes, comments and amendments.
- Held-out status. Open legal corpora exist to feed pretraining; Pile of Law gathers about 256 GB from 35 sources [4]. One study notes that many LLMs' training data is contaminated with test data and that most benchmarks lack a private holdout [5].
Keep public sets for comparability; for the trade-offs, see private evaluation sets vs public benchmarks and the evaluation datasets hub.
Four review tasks and the gold label each needs
Each review capability needs its own gold label, taken from a record a lawyer produced or approved, never from a model's output.
| Task | Model sees | Gold label | Label source | How to score |
|---|---|---|---|---|
| Clause extraction | Full agreement (DOCX or PDF) | Spans and clause types from your taxonomy, including "absent" | Lawyer annotation; contract-system clause tags | Span precision and recall under a stated overlap rule; separate score for missing clauses |
| Key-term extraction | Agreement or clause | Normalized governing law, term, renewal notice, cap basis | Metadata entered at signature, checked against executed text | Exact match after normalizing dates, durations and currency |
| Playbook deviation flagging | Counterparty draft plus playbook | Deviates or acceptable, severity, position breached | What the lawyer changed or escalated; approval records | Precision and recall per clause type; high-severity misses reported separately |
| Redline proposal | Counterparty draft at turn n plus playbook | Lawyer rubric grade; actual next-turn markup as one reference | Tracked changes in the next draft; executed version | Rubric on position reached, scope of change and drafting, not string overlap |
Three choices separate a useful set from a demo: include contracts where a clause is genuinely absent, since invented caps and indemnities are a common failure; include document families so the model must apply an order-of-precedence clause; and label a deviation accepted with documented approval differently from one the lawyer missed.
Where the answers already exist in a legal team's records
Most gold labels need not be written from scratch: a negotiated agreement leaves them in the draft chain, the approvals and the executed copy.
- Draft chain. Versions in a document management system (for example iManage or NetDocuments) or a shared drive. Ask for native DOCX: a PDF print flattens tracked changes and loses who changed what.
- Playbook and approvals. Standard, fallback and walk-away positions per clause type, plus sign-offs showing which deviations were accepted on purpose.
- Contract-system metadata. Governing law, renewal dates and notice periods keyed in at signature. Verify them against the executed text: an audit of 10 widely used image, text and audio test sets estimated average label errors of at least 3.3% and showed such errors can reorder model rankings [6].
- Executed version. Which position held; see outcome-labeled evaluation data.
Record structure is on the contract redline datasets page; training-side labeling is in contract clause annotation data. SourceX sources operational datasets, including documents and legal workflows, from US companies on request, not from stock, so a request does not guarantee a match. To check whether such records exist, describe the contracts and labels your evaluation needs.
Illustrative eval item: a limitation-of-liability turn
One item packages a decision point: the text the model sees, the playbook it applies and answers taken from what lawyers did next.
Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.
{
"item_id": "cr-eval-0412",
"split": "held_out",
"task": "playbook_deviation_flag",
"contract_type": "saas_subscription",
"paper": "counterparty",
"reviewing_side": "customer",
"governing_law": "US-NY",
"turn_index": 2,
"turns_total": 5,
"input": {
"document_sha256": "9f2c...e41a",
"clause_ref": "Section 9.2",
"text": "In no event shall VENDOR_1's aggregate liability exceed the fees paid in the three (3) months preceding the claim."
},
"playbook": {"standard": "12 months of fees", "fallback": "6 months of fees", "walk_away": "under 6 months"},
"gold": {
"clause_type": "limitation_of_liability",
"deviation": true,
"severity": "escalate",
"reference_markup": "three (3) months -> twelve (12) months",
"executed_outcome": "settled_at_fallback",
"approval": "fallback accepted; sign-off recorded by role only"
},
"labeling": {"annotators": ["L1", "L2"], "agreement": "match", "adjudicated": false},
"slices": ["counterparty_paper", "liability_cap", "mid_negotiation"],
"redaction": {"party_names": "consistent typed placeholders", "amounts": "kept visible, as agreed in scoping"}
}
Ship items as JSON Lines, one UTF-8 JSON value per line [7], referencing each source document by hash.
Defensible gold labels: two lawyers, adjudication and agreement
A gold label is only as reliable as the agreement between qualified reviewers, so double-label scoring-critical items, adjudicate disagreements and report agreement per clause type. Krippendorff's alpha measures agreement among annotators assigning values to the same items, where 1 is perfect reliability and 0 means none beyond chance [8]. Low alpha on one clause type often signals a taxonomy problem, such as a consequential-damages exclusion and a liability cap sharing a section; fix definitions before scoring models.
Record each annotator's qualifications and every adjudication. For redline items, treat the lawyer's actual markup as one acceptable answer, not the only one, and grade model markups with a lawyer-written rubric; see designing evaluation rubrics with domain experts. If an LLM judge applies that rubric at scale, validate it first: the LiveBench authors note that LLM judges can introduce significant biases and break down on hard questions [9]. See the LLM-as-a-judge calibration set and gold-label audit at acceptance guides.
Slices that show where contract review models break
Report scores by slice, because one accuracy number hides where review models fail.
- Contract type: NDA, master services agreement, SaaS subscription, data processing addendum, supply, lease.
- Paper, side and governing law: your paper or theirs; customer or vendor; the state named, which can change what your playbook allows.
- Negotiation intensity: turns, changes per turn, contested clauses versus untouched boilerplate.
- Document family and condition: amendments and order forms that override base terms; native DOCX versus scanned exhibits.
- Absent and unusual clauses: no cap, carve-outs, cross-referenced definitions.
- Date: agreements executed after your candidate models' training cutoffs.
Rare slices need deliberate allocation; see stratified evaluation sets for rare and high-risk cases and how many examples an eval set needs.
Keeping a private legal eval set out of training data
A private contract eval set stays useful only while no model under test has trained on it, so control exposure at the API, the curator and the license. In a 23 February 2026 post, OpenAI said it had stopped reporting SWE-bench Verified because score gains increasingly reflect training-time exposure to that benchmark [10].
- Hosted models. API evaluation sends every item to the provider, so use terms that exclude training on inputs; FTC staff wrote in 2024 that model-as-a-service companies may be liable if they break promises not to use customer data for training [11].
- Curator independence. A 2025 paper examines transparency and conflict-of-interest risks when private data curators build or run evaluations [12]. Ask whether the same items or annotators also feed training data the curator sells. SourceX does not train AI models.
Item hashes, canary strings and access logs are covered in keeping a private eval set private; see also contamination checks for licensed evaluation data.
Rights, confidentiality and de-identification for licensed contracts
Contracts carry duties to counterparties and sometimes to clients, so confirm the supplier may share each agreement and its drafts for evaluation before labeling starts.
- Confidentiality terms. Check whether each agreement's confidentiality clause covers its own terms and how counterparty-authored markup may be used. If the supplier is a law firm, ask how client authority was obtained per matter. Exclude internal comments giving legal advice rather than redacting them.
- De-identification that keeps tests valid. Replace party names with consistent typed placeholders, including inside defined terms, or "who indemnifies whom" items become unanswerable; masking a cap amount breaks key-term extraction. SourceX removes or replaces personal details such as names, emails and account numbers before delivery, records the method and checks a processed sample; no de-identification method is perfect. State company-name and deal-value handling separately; see de-identifying evaluation data without breaking the test.
- Labels you did not create. Annotations copied from a legal publisher's products are its editorial work. On 29 September 2026 the Third Circuit, in a precedential opinion, held the Westlaw headnotes at issue copyrightable and ROSS's use of Westlaw material to train a non-generative legal-research tool not fair use [13]. The case concerned training, but it is a reason to document each label's origin.
- License scope. Every SourceX dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and is delivered under a license defining included records, permitted uses, term and delivery. Negotiate evaluation-only use, score publication and retention explicitly; see evaluation-only data license terms.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What to specify when requesting a contract review eval set
Naming tasks, taxonomy, contract mix, label method and license scope lets a supplier judge whether its records can support the set.
| Specify | Contract review detail |
|---|---|
| Tasks | Which of the four tasks; the draft turn the model sees |
| Taxonomy | Your clause types and definitions (CUAD-style categories as a start), plus "absent" |
| Contract mix | Types, paper, reviewing side, governing-law states, minimum items per slice |
| Documents | Native DOCX draft chains, executed copies, amendments and order forms; execution date range |
| Ground truth | Next-turn markup, approvals, executed outcome; who supplies playbook positions |
| Labeling | Annotator qualifications, double-labeled share, adjudication rule, agreement per clause type |
| Privacy and rights | Party placeholders, visible values, privilege exclusion, counterparty text |
| License | Evaluation-only scope, no training on items, score publication, retention |
| Delivery | JSON Lines items, documents by hash, a datasheet on motivation, composition and collection [14] |
See also golden evaluation datasets from business records and, for training data, legal AI use cases.
Sourcing negotiated contracts for a private legal eval set?
Describe the contract types, review tasks, label method and license scope your evaluation needs. SourceX looks for US businesses that hold matching agreements and draft histories, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Describe your contract review evaluation needs.
Sources
- Guha et al., NeurIPS 2023 Datasets and Benchmarks, "LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models" (2023). https://proceedings.neurips.cc/paper_files/paper/2023/hash/89e44582fd28ddfea1ea4dcb0ebbf4b0-Abstract.html
- Stanford Hazy Research, "LegalBench: a collaboratively built large language model benchmark for legal reasoning" (project site). https://hazyresearch.stanford.edu/legalbench
- arXiv:2509.19580, "LLMs4All: A Review of Large Language Models Across Academic Disciplines" (2025). https://arxiv.org/pdf/2509.19580
- Henderson et al., arXiv:2207.00220, "Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset" (2022). https://arxiv.org/pdf/2207.00220
- arXiv:2410.09247, "Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts" (2024). https://arxiv.org/pdf/2410.09247.pdf
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- White et al., ICLR 2025, "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- arXiv:2503.04756, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
- Gebru et al., "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.