Skip to content

Industry-specific operational data

Litigation docket and motion-outcome data: linking filings to rulings for litigation AI

Quick answer

Litigation outcome data for AI is a set of court filings joined to the rulings that resolved them: each motion paired with its order, a normalized disposition label, and case context such as court, judge, nature of suit and dates. Public dockets supply the filings and orders, at a per-page cost and with uneven state coverage. The drafts, internal assessments and client outcomes behind those filings exist only inside law firms and require privilege, client and rights review before any license.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What a motion-to-ruling record has to contain

A usable outcome record links one moving document to the specific order that disposed of it, not just to the case's final judgment. On a federal CM/ECF docket, that means pairing a motion entry (for example, entry 45, "MOTION to Dismiss for Failure to State a Claim") with the later entry that resolves it ("ORDER granting in part and denying in part" with a cross-reference back to entry 45). The bracketed cross-reference is the most reliable join key, but clerks and chambers do not apply it consistently.

The minimum fields buyers should require:

  • Case identity: court, case number, nature-of-suit code, filing date, assigned judge and any reassignment.
  • Motion identity: docket entry number, filing date, filer role (plaintiff, defendant, intervenor), motion type, and the attached memorandum, declarations and proposed order.
  • Opposition and reply: entry numbers and dates, so a model sees the full briefing cycle rather than one side.
  • Ruling: entry number, date, ruling type (written opinion, text-only order, minute entry, oral ruling reflected in a transcript), and the disposition label.
  • Censoring flags: motion withdrawn, mooted by amended pleading, mooted by settlement, stayed, or still pending at snapshot date.

The common failure mode is treating "no order found" as "denied." Many motions are terminated as moot when an amended complaint is filed, and many cases settle with motions pending. Those records must be labeled as censored, or an outcome-prediction model learns a false denial rate.

What public dockets and open corpora give you

Public sources give you opinions and filings at scale, but not clean outcome labels and not the work product behind the filings. PACER bills federal docket sheets and documents per page, with a per-document cap that does not cover every report type or transcripts. Pulling full docket sheets plus every motion, opposition and order across thousands of cases is a real line item, so verify the current PACER fee schedule ($0.10 per page, increasing to $0.12 on January 1, 2027) before budgeting [8].

Free and open alternatives cover different slices:

  • Opinions: The Caselaw Access Project's original access restrictions expired in March 2024, and its data can now be released without restriction on access or use [2]. CourtListener publishes per-court coverage for its opinion collection, which varies widely by court and date range [1].
  • Pretraining corpora: Pile of Law collects 256GB from 35 sources, including court opinions and filings, under open licenses [3]. MultiLegalPile covers 689GB across 24 languages with mixed licenses, and its authors frame pretraining use as fair use, which is their position rather than a settled rule [4].
  • State courts: There is no national equivalent of PACER. Access ranges from statewide portals to county-by-county clerk systems, many without bulk export, and redaction practices differ.

These corpora suit pretraining and RAG over opinions. They rarely carry motion-level joins, so outcome labeling is still your pipeline's job. For other legal and industry data types, see the industry-specific operational data hub and the broader AI data hub.

What only law firms hold

The data that improves litigation drafting and strategy models lives inside firms, not on the docket. A filed brief is the last version; the firm also holds the drafts, partner redlines, research memos that shaped the argument, internal case assessments with estimated win probabilities, settlement authority and valuations, and the client-reported final result when a case resolves out of court.

That material is what turns a docket record into a supervised example. A pair like "first-year associate draft → partner-edited filed brief → order granting the motion" teaches drafting quality against a real outcome. Our guide to legal LLM fine-tuning data, work product and privilege covers redline pairs in depth; finished briefs and memos as documents are covered under licensing legal briefs and memos.

Firm data carries constraints public data does not. Attorney-client privilege belongs to the client, work-product protection can be asserted by client and counsel, and engagement letters, outside-counsel guidelines and protective orders under the federal civil rules can restrict reuse of discovery material. A firm is a service provider holding client information, so the authorization questions in client data held by service providers apply directly.

Sourcing decision: public docket, firm work product or both

Most litigation AI programs need both: public dockets for coverage and outcome labels, firm data for drafts and unfiled context. The table below maps common applications to the source that actually carries the signal.

Illustrative example: invented to show structure; it does not describe an available dataset.

ApplicationPublic docket and opinionsFirm work productMain risk to check
Motion-outcome predictionMotion, briefing and order entries; judge and court fieldsInternal assessments, settlement values for censored casesMislabeled moot or settled motions
RAG over docketsDocket text, orders, opinionsRarely neededResidual personal data in filings
Brief drafting SFTFiled briefs only (final versions)Drafts, redlines, research memosPrivilege and client consent
Drafting evaluationBriefs paired with rulingsPartner edits as reference answersLeakage of test cases into training
PretrainingOpen corpora and opinionsNot usually licensed for thisPublisher headnotes and annotations

Rights and privacy review for litigation data

Court opinions are generally public domain in the United States, but the editorial layer around them is not. On September 29, 2026 the Third Circuit affirmed that the Westlaw headnotes at issue were copyrightable and that copying them to build a legal-research tool was not fair use; the tool was non-generative, and the opinion addresses that setting [5]. Strip or exclude publisher headnotes, key numbers, synopses and annotations from anything sourced through commercial research services, and confirm what your scraping or export terms allow.

Do not trust license tags on aggregated legal datasets without checking. An audit of 1,800+ text datasets found license omissions above 70% and license errors above 50% on popular hosting sites [6]. Trace each subset of a legal corpus to its original source and terms.

Public filings are not free of personal data. Federal Rule of Civil Procedure 5.2 [9] limits filers to the last four digits of Social Security, taxpayer-identification and financial-account numbers, the birth year and a minor's initials, but it exempts categories such as official state-court records, and the duty to redact rests with the filer rather than the clerk. Expect unredacted identifiers in exhibits, medical details in personal-injury and disability matters, and names of non-party witnesses. Run your own detection and pseudonymization pass before training, and record the method.

Labeling outcomes without leaking the answer

Outcome labels should come from the order, and model inputs should stop at the moment the court ruled. Build each example from entries filed before the ruling date, and exclude later entries such as notices of appeal, settlement filings or a second motion that quotes the first ruling.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "case_id": "D-XX-1:24-cv-00000",
  "court": "District court (anonymized)",
  "nature_of_suit": "Contract: other",
  "judge_id": "J-0412",
  "motion": {"entry": 45, "filed": "2025-03-14", "type": "motion_to_dismiss_12b6", "filer_role": "defendant"},
  "briefing": [{"entry": 51, "type": "opposition"}, {"entry": 55, "type": "reply"}],
  "ruling": {"entry": 62, "date": "2025-07-02", "form": "written_order", "disposition": "granted_in_part", "leave_to_amend": true},
  "censoring": null,
  "input_cutoff_date": "2025-07-01",
  "pii_method": "names and identifiers replaced with role tokens; sample reviewed"
}

Use a closed disposition vocabulary: granted, denied, granted in part, moot, withdrawn, held in abeyance, and pending at snapshot. Split train and test sets by time and, for judge-level analytics, check that the same judge's later rulings are not leaking into the training window. Document all of this in a data card covering sources, labeling rules and known gaps [7].

Request checklist for litigation outcome data

A precise request saves back-and-forth with any supplier, and it is the same brief you would submit to SourceX for buyers. Before contacting a vendor, firm or marketplace, write down:

  1. Courts and case types: federal districts, state courts, nature-of-suit codes, and date range.
  2. Unit of record: motion-to-ruling pair, whole docket, or draft-to-filed pair.
  3. Required joins: motion to order, opposition and reply to motion, docket to final judgment or client-reported resolution.
  4. Label vocabulary: your disposition codes, including how moot, withdrawn and settled outcomes are flagged.
  5. Use: pretraining, SFT, RAG, or evaluation only, and whether outputs are commercial.
  6. Exclusions: sealed material, publisher annotations, discovery under protective order, matters for specific clients.
  7. Privacy handling: de-identification method, sample review, and how residual identifiers are reported.
  8. Delivery: format (JSONL, Parquet, PDF plus extracted text), document-to-entry mapping, and refresh cadence.

Licensing litigation outcome data for AI through SourceX

SourceX sources operational datasets, including legal workflows and documents, on request from US companies that hold them, and every release is approved by the supplying company; a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded, and delivery runs under a license defining records, uses, term and delivery. For legal-sector requests, see the legal industry buyer page, or describe the litigation data you need.

Sources

  1. Free Law Project (CourtListener), "Data Coverage: What Case Law Does CourtListener Have?". https://courtlistener.com/help/coverage/opinions/
  2. Harvard Law School Library Innovation Lab, "Transitions for the Caselaw Access Project" (2024). https://lil.law.harvard.edu/blog/2024/03/26/transitions-for-the-caselaw-access-project
  3. Henderson et al., "Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset" (2022). https://arxiv.org/pdf/2207.00220
  4. Niklaus et al., "MultiLegalPile: A 689GB Multilingual Legal Corpus" (2023). https://arxiv.org/abs/2306.02069
  5. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  6. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  7. Pushkarna, Zaldivar, Kjartansson, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  8. United States Courts, "Electronic Public Access Fee Schedule" (2026). https://pacer.uscourts.gov/announcements/2026/06/26/temporary-fee-increase-effective-jan-1-2027
  9. Legal Information Institute, "Rule 5.2. Privacy Protection for Filings Made with the Court". https://www.law.cornell.edu/rules/frcp/rule_5.2

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data