Document AI data
Contract Clause Annotation Data for Contract Review Models
Quick answer
A contract clause extraction dataset pairs full agreements with lawyer-applied labels: the clause type, the exact text span, normalized metadata such as dates, parties and governing law, and an explicit "absent" label when a clause is missing. CUAD is the standard public starting point [1], but it covers SEC-filed contracts only. Production review models usually also need private agreements, a taxonomy mapped to your product and separate metrics for span extraction and clause presence.
By SourceX Editorial · Updated
This page covers the clause annotation spec itself. For redline and negotiation history, see the contract redline datasets page; for the wider use case, see legal AI training data from real contract work. It sits in the document AI data hub.
What public clause datasets cover, and where they stop
Public clause datasets give you a benchmark and a taxonomy, but not the distribution of contracts your customers actually upload. CUAD, from The Atticus Project, labels clauses in commercial contracts drawn from public filings across 25 contract types and 41 categories such as Governing Law, License Grant, Change of Control and Insurance [1]. Its public release is packaged as extractive question answering (SQuAD-style question, context and answer spans). Confirm the license file in the exact release you download, since loader code and data can carry different licenses.
Other public sets fill narrower gaps. ACORD is an expert-annotated clause retrieval dataset aimed at drafting, with categories like limitation of liability and indemnification [2]. RealKIE includes US non-disclosure agreements and resource contracts for key information extraction [3], LegalBench tests legal reasoning across 162 tasks, some contract-related [4], and 3CEL covers Spanish contract clauses [6].
The limits matter for a commercial product:
- Selection bias. Contracts filed as exhibits are material agreements of public companies. They underrepresent vendor paper, click-through terms, NDAs between private firms, order forms and amendments.
- Clean text. Filings are mostly born-digital HTML or text. Your users upload scanned PDFs with signatures, stamps, redaction boxes and exhibit attachments.
- Fixed taxonomy. CUAD's 41 categories will not match your product's playbook, for example splitting "Cap on Liability" into general cap, super-cap and carve-outs.
- No negotiation context. Labels describe the final text, not which party's paper it was or which positions moved.
For a broader license comparison, see public document AI datasets and commercial use.
The annotation spec buyers should agree before sourcing
A clause annotation spec should fix four things before any labeling starts: the taxonomy, the span rules, the absent-clause rule and the metadata normalization. Without them, two suppliers' "indemnification" labels will not be comparable, and your model will learn the disagreement.
Taxonomy mapping. Write your categories as a table that maps each one to a CUAD category where one exists, with a definition, inclusion and exclusion notes, and a parent for hierarchy (for example, Termination > Termination for Convenience). Mark categories that are many-to-one, such as several CUAD categories collapsing into a single "Assignment and Change of Control" label.
Span rules. Decide whether a span is the minimal operative sentence, the full numbered section or every passage that contributes, including definitions referenced by the clause. Multi-span labels are common: a liability cap often sits in one section with carve-outs in another. Record character offsets against a frozen text layer, plus page and bounding box if you train on PDFs.
Absent-clause rule. "No label" must not mean "not reviewed." Require one of three explicit states per category per document: present, absent after full review, or not applicable to the contract type. Absent labels are what train a model to say "no non-compete found" with confidence.
Metadata normalization. Dates become ISO 8601, with a flag for relative dates ("12 months after the Effective Date"). Parties carry role labels (licensor, customer, vendor). Governing law maps to a jurisdiction code, and renewal terms record type (auto-renew, evergreen, fixed) plus notice period in days.
The document annotation schema guide covers how these layers sit alongside layout, reading order and tables.
An illustrative clause annotation record
A single record should carry the document identity, the clause label with its offsets, the review state and who applied it. The structure below is a pattern to put into a request or acceptance spec.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "agr-000417",
"contract_type": "SaaS subscription agreement",
"paper": "vendor",
"text_layer": {"source": "ocr", "engine_version": "recorded", "sha256": "..."},
"family": {"parent_doc_id": "msa-000112", "relation": "order_form"},
"labels": [
{
"category": "cap_on_liability",
"cuad_map": "Cap On Liability",
"state": "present",
"spans": [{"start": 18842, "end": 19310, "page": 9}],
"normalized": {"cap_basis": "fees_paid_12_months", "carve_outs": ["confidentiality", "indemnity"]}
},
{
"category": "non_compete",
"cuad_map": "Non-Compete",
"state": "absent_after_review",
"spans": []
},
{
"category": "renewal_term",
"state": "present",
"spans": [{"start": 4021, "end": 4388, "page": 2}],
"normalized": {"renewal_type": "auto_renew", "notice_days": 60}
}
],
"annotation": {"annotator_role": "licensed_attorney", "adjudicated": true, "guideline_version": "v3.2"}
}
The family field matters for amendments and order forms, which often override the master agreement. See contract families linking master agreements to amendments for that structure.
Using CLM metadata as labels
Metadata already stored in contract lifecycle management (CLM) systems can serve as weak labels, but only after someone verifies it against the signed text. Typical CLM fields include effective date, expiration, renewal type, notice period, counterparty, governing law and contract value, often entered by a paralegal at intake.
Treat these fields carefully:
- Field-level, not span-level. CLM records rarely store where in the document the value came from, so they help train metadata extraction but not span extraction.
- Drift from the document. Fields are often entered once and not updated when an amendment changes the term, so pair them with the amendment chain.
- Free-text overrides. "See SOW" or "per Section 8" in a date field is a common failure; require a parsing rate and a sample audit before accepting.
The same pattern of pairing documents with system-of-record values is covered in documents paired with system-of-record entries.
Measuring span extraction and clause presence separately
Span extraction and clause presence are different tasks and should be scored separately, because a model can find the right clause and still return the wrong boundaries. Report at least two numbers per category:
| Metric | What it measures | Typical failure it catches |
|---|---|---|
| Presence precision and recall | Whether the model says a category exists in the document | Hallucinated non-competes; missed auto-renewal |
| Span overlap (token F1 or Jaccard) | How much of the gold span the prediction covers | Returning the heading but not the operative sentence |
| Exact or boundary-tolerant match | Whether boundaries align within a set tolerance | Splitting a cap and its carve-outs into separate hits |
| Normalized field accuracy | Whether extracted values match after normalization | "Sixty (60) days" parsed as 6; wrong governing-law state |
| Absent-label accuracy | Whether "not found" is correct | Under-reviewed documents labeled absent |
CUAD's authors framed evaluation around precision at high recall because a reviewer cares most about not missing clauses [1]. Weight your categories the same way: a missed indemnity costs more than a false positive on a notice clause. Gold labels themselves contain errors, and label errors in widely used test sets have been shown to change model rankings [5], so audit a sample before you trust a score. The annotation quality audit guide and contract review evaluation datasets cover acceptance testing in more depth.
Rights and confidentiality checks for private contracts
Private agreements carry rights questions that public filings do not, because the counterparty usually never agreed to the contract being shared. Most commercial contracts include a confidentiality clause covering the agreement's terms, so the company holding the contract may need to assess whether it can license the text at all, or only derived labels and redacted versions.
Questions to put to any supplier:
- Which contracts carry confidentiality terms covering the agreement itself, and how was that assessed?
- Are counterparty names, signatory names, emails, addresses and bank or account details removed or replaced, and how is the method recorded?
- Are pricing schedules and commercially sensitive exhibits included, redacted or excluded?
- Who applied the labels (licensed attorneys, law students, contract analysts), under what agreement, and who owns the annotations?
- Does the license permit training, evaluation and fine-tuning, and does it cover derived models?
For commissioned annotation, read IP assignment vs license in data collection contracts, and for annotator documentation see provenance for human-annotated data.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
How SourceX approaches licensed contract data
SourceX sources operational datasets, including documents and legal workflows, from US companies and manages the commercial process from licensing to ongoing purchases. Data is sourced on request rather than held in stock, so a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery.
Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, with the method recorded and a sample checked; no method is perfect. Every release is approved by the supplying company, and diligence materials on source, rights, preparation and allowed use are prepared per dataset. You can describe your clause taxonomy and contract mix to SourceX to start the Find and Assess steps.
Request contract clause extraction data
Describe the contract types, clause categories, label format and intended uses you need, rather than the companies you want it from. SourceX looks for US businesses that hold matching data, assesses data and licensing permissions, and agrees pricing and allowed uses in a license; nothing is contracted until a supplier agrees. Start a contract clause data request at sourcex.si/buyers.
Sources
- arXiv (Hendrycks et al.), "CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review" (2021). https://arxiv.org/pdf/2103.06268
- arXiv, "ACORD: An Expert-Annotated Retrieval Dataset for Legal Contract Drafting" (2025). https://arxiv.org/pdf/2501.06582
- arXiv, "RealKIE: Five Novel Datasets for Enterprise Key Information Extraction" (2024). https://arxiv.org/pdf/2403.20101
- Stanford Hazy Research, "LegalBench: A collaboratively built large language model benchmark for legal reasoning". https://hazyresearch.stanford.edu/legalbench
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/pdf/2103.14749
- arXiv, "3CEL: a Corpus of Legal Spanish Contract Clauses" (2025). https://arxiv.org/html/2501.15990v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.