Industry-specific operational data
Litigation docket and motion-outcome data: linking filings to rulings for litigation AI
Quick answer
Litigation outcome data for AI is a set of court filings joined to the rulings that resolved them: each motion paired with its order, a normalized disposition label, and case context such as court, judge, nature of suit and dates. Public dockets supply the filings and orders, at a per-page cost and with uneven state coverage. The drafts, internal assessments and client outcomes behind those filings exist only inside law firms and require privilege, client and rights review before any license.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What a motion-to-ruling record has to contain
A usable outcome record links one moving document to the specific order that disposed of it, not just to the case's final judgment. On a federal CM/ECF docket, that means pairing a motion entry (for example, entry 45, "MOTION to Dismiss for Failure to State a Claim") with the later entry that resolves it ("ORDER granting in part and denying in part" with a cross-reference back to entry 45). The bracketed cross-reference is the most reliable join key, but clerks and chambers do not apply it consistently.
The minimum fields buyers should require:
- Case identity: court, case number, nature-of-suit code, filing date, assigned judge and any reassignment.
- Motion identity: docket entry number, filing date, filer role (plaintiff, defendant, intervenor), motion type, and the attached memorandum, declarations and proposed order.
- Opposition and reply: entry numbers and dates, so a model sees the full briefing cycle rather than one side.
- Ruling: entry number, date, ruling type (written opinion, text-only order, minute entry, oral ruling reflected in a transcript), and the disposition label.
- Censoring flags: motion withdrawn, mooted by amended pleading, mooted by settlement, stayed, or still pending at snapshot date.
The common failure mode is treating "no order found" as "denied." Many motions are terminated as moot when an amended complaint is filed, and many cases settle with motions pending. Those records must be labeled as censored, or an outcome-prediction model learns a false denial rate.
What public dockets and open corpora give you
Public sources give you opinions and filings at scale, but not clean outcome labels and not the work product behind the filings. PACER bills federal docket sheets and documents per page, with a per-document cap that does not cover every report type or transcripts. Pulling full docket sheets plus every motion, opposition and order across thousands of cases is a real line item, so verify the current PACER fee schedule ($0.10 per page, increasing to $0.12 on January 1, 2027) before budgeting [8].
Free and open alternatives cover different slices:
- Opinions: The Caselaw Access Project's original access restrictions expired in March 2024, and its data can now be released without restriction on access or use [2]. CourtListener publishes per-court coverage for its opinion collection, which varies widely by court and date range [1].
- Pretraining corpora: Pile of Law collects 256GB from 35 sources, including court opinions and filings, under open licenses [3]. MultiLegalPile covers 689GB across 24 languages with mixed licenses, and its authors frame pretraining use as fair use, which is their position rather than a settled rule [4].
- State courts: There is no national equivalent of PACER. Access ranges from statewide portals to county-by-county clerk systems, many without bulk export, and redaction practices differ.
These corpora suit pretraining and RAG over opinions. They rarely carry motion-level joins, so outcome labeling is still your pipeline's job. For other legal and industry data types, see the industry-specific operational data hub and the broader AI data hub.
What only law firms hold
The data that improves litigation drafting and strategy models lives inside firms, not on the docket. A filed brief is the last version; the firm also holds the drafts, partner redlines, research memos that shaped the argument, internal case assessments with estimated win probabilities, settlement authority and valuations, and the client-reported final result when a case resolves out of court.
That material is what turns a docket record into a supervised example. A pair like "first-year associate draft → partner-edited filed brief → order granting the motion" teaches drafting quality against a real outcome. Our guide to legal LLM fine-tuning data, work product and privilege covers redline pairs in depth; finished briefs and memos as documents are covered under licensing legal briefs and memos.
Firm data carries constraints public data does not. Attorney-client privilege belongs to the client, work-product protection can be asserted by client and counsel, and engagement letters, outside-counsel guidelines and protective orders under the federal civil rules can restrict reuse of discovery material. A firm is a service provider holding client information, so the authorization questions in client data held by service providers apply directly.
Sourcing decision: public docket, firm work product or both
Most litigation AI programs need both: public dockets for coverage and outcome labels, firm data for drafts and unfiled context. The table below maps common applications to the source that actually carries the signal.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Application | Public docket and opinions | Firm work product | Main risk to check |
|---|---|---|---|
| Motion-outcome prediction | Motion, briefing and order entries; judge and court fields | Internal assessments, settlement values for censored cases | Mislabeled moot or settled motions |
| RAG over dockets | Docket text, orders, opinions | Rarely needed | Residual personal data in filings |
| Brief drafting SFT | Filed briefs only (final versions) | Drafts, redlines, research memos | Privilege and client consent |
| Drafting evaluation | Briefs paired with rulings | Partner edits as reference answers | Leakage of test cases into training |
| Pretraining | Open corpora and opinions | Not usually licensed for this | Publisher headnotes and annotations |
Rights and privacy review for litigation data
Court opinions are generally public domain in the United States, but the editorial layer around them is not. On September 29, 2026 the Third Circuit affirmed that the Westlaw headnotes at issue were copyrightable and that copying them to build a legal-research tool was not fair use; the tool was non-generative, and the opinion addresses that setting [5]. Strip or exclude publisher headnotes, key numbers, synopses and annotations from anything sourced through commercial research services, and confirm what your scraping or export terms allow.
Do not trust license tags on aggregated legal datasets without checking. An audit of 1,800+ text datasets found license omissions above 70% and license errors above 50% on popular hosting sites [6]. Trace each subset of a legal corpus to its original source and terms.
Public filings are not free of personal data. Federal Rule of Civil Procedure 5.2 [9] limits filers to the last four digits of Social Security, taxpayer-identification and financial-account numbers, the birth year and a minor's initials, but it exempts categories such as official state-court records, and the duty to redact rests with the filer rather than the clerk. Expect unredacted identifiers in exhibits, medical details in personal-injury and disability matters, and names of non-party witnesses. Run your own detection and pseudonymization pass before training, and record the method.
Labeling outcomes without leaking the answer
Outcome labels should come from the order, and model inputs should stop at the moment the court ruled. Build each example from entries filed before the ruling date, and exclude later entries such as notices of appeal, settlement filings or a second motion that quotes the first ruling.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"case_id": "D-XX-1:24-cv-00000",
"court": "District court (anonymized)",
"nature_of_suit": "Contract: other",
"judge_id": "J-0412",
"motion": {"entry": 45, "filed": "2025-03-14", "type": "motion_to_dismiss_12b6", "filer_role": "defendant"},
"briefing": [{"entry": 51, "type": "opposition"}, {"entry": 55, "type": "reply"}],
"ruling": {"entry": 62, "date": "2025-07-02", "form": "written_order", "disposition": "granted_in_part", "leave_to_amend": true},
"censoring": null,
"input_cutoff_date": "2025-07-01",
"pii_method": "names and identifiers replaced with role tokens; sample reviewed"
}
Use a closed disposition vocabulary: granted, denied, granted in part, moot, withdrawn, held in abeyance, and pending at snapshot. Split train and test sets by time and, for judge-level analytics, check that the same judge's later rulings are not leaking into the training window. Document all of this in a data card covering sources, labeling rules and known gaps [7].
Request checklist for litigation outcome data
A precise request saves back-and-forth with any supplier, and it is the same brief you would submit to SourceX for buyers. Before contacting a vendor, firm or marketplace, write down:
- Courts and case types: federal districts, state courts, nature-of-suit codes, and date range.
- Unit of record: motion-to-ruling pair, whole docket, or draft-to-filed pair.
- Required joins: motion to order, opposition and reply to motion, docket to final judgment or client-reported resolution.
- Label vocabulary: your disposition codes, including how moot, withdrawn and settled outcomes are flagged.
- Use: pretraining, SFT, RAG, or evaluation only, and whether outputs are commercial.
- Exclusions: sealed material, publisher annotations, discovery under protective order, matters for specific clients.
- Privacy handling: de-identification method, sample review, and how residual identifiers are reported.
- Delivery: format (JSONL, Parquet, PDF plus extracted text), document-to-entry mapping, and refresh cadence.
Licensing litigation outcome data for AI through SourceX
SourceX sources operational datasets, including legal workflows and documents, on request from US companies that hold them, and every release is approved by the supplying company; a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded, and delivery runs under a license defining records, uses, term and delivery. For legal-sector requests, see the legal industry buyer page, or describe the litigation data you need.
Sources
- Free Law Project (CourtListener), "Data Coverage: What Case Law Does CourtListener Have?". https://courtlistener.com/help/coverage/opinions/
- Harvard Law School Library Innovation Lab, "Transitions for the Caselaw Access Project" (2024). https://lil.law.harvard.edu/blog/2024/03/26/transitions-for-the-caselaw-access-project
- Henderson et al., "Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset" (2022). https://arxiv.org/pdf/2207.00220
- Niklaus et al., "MultiLegalPile: A 689GB Multilingual Legal Corpus" (2023). https://arxiv.org/abs/2306.02069
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Pushkarna, Zaldivar, Kjartansson, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- United States Courts, "Electronic Public Access Fee Schedule" (2026). https://pacer.uscourts.gov/announcements/2026/06/26/temporary-fee-increase-effective-jan-1-2027
- Legal Information Institute, "Rule 5.2. Privacy Protection for Filings Made with the Court". https://www.law.cornell.edu/rules/frcp/rule_5.2
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.