Skip to content

Industry-specific operational data

Legal billing data: LEDES invoices, UTBMS codes and bill-review adjustments

Quick answer

Legal invoice data for AI is most useful when it joins three layers: the LEDES invoice the law firm submitted, the UTBMS task, activity and expense codes on each line, and the reviewer's actions afterward (reductions, reason codes, appeals and the amount finally paid). Invoices alone teach extraction and code classification. Bill-review and guideline-violation models need the adjustment layer, tied to the governing billing guideline version, with privileged narratives and firm identities handled before licensing.

By SourceX Editorial · Updated

A training-grade record is one invoice line with its narrative, codes, rate and every downstream decision made about it. The submitted layer usually arrives as LEDES 1998B, a pipe-delimited ASCII format with 24 fields that is the most common e-billing standard in US legal work; LEDES.org publishes the field specification and sample files, so check samples against it. Firms billing across borders often use LEDES 1998BI, which extends the record to 52 fields with timekeeper classification, tax rate and addresses, or LEDES XML for richer structure.

The 1998B fields that matter most for modeling are LINE_ITEM_DESCRIPTION (the time-entry narrative), LINE_ITEM_TASK_CODE, LINE_ITEM_ACTIVITY_CODE, LINE_ITEM_EXPENSE_CODE, LINE_ITEM_NUMBER_OF_UNITS, LINE_ITEM_UNIT_COST, LINE_ITEM_ADJUSTMENT_AMOUNT, LINE_ITEM_DATE, TIMEKEEPER_ID and TIMEKEEPER_CLASSIFICATION. Note the trap in LINE_ITEM_ADJUSTMENT_AMOUNT: it records the firm's own discount at submission, not the client's review reduction. Reviewer cuts live in the e-billing or spend-management system, so a LEDES dump without that system's audit export cannot train a bill-review model.

The review layer is where the labels sit. Ask for line-level reduction amount, reduction reason (often a coded list such as block billing, vague description, excessive staffing, minimum increment, rate above approved rate, non-billable administrative task), reviewer type (automated rule, staff reviewer, attorney reviewer), appeal text from the firm, appeal outcome, and final paid amount. For accrual prediction, add matter phase, budget by UTBMS phase, invoice submission and approval dates, and matter close date.

How UTBMS codes shape classification datasets

UTBMS codes give each line a hierarchical label, which makes task-code classification a well-posed supervised problem. The litigation code set groups tasks by phase: L100 case assessment, development and administration (L110 fact investigation, L120 analysis and strategy, L160 settlement and non-binding ADR), L200 pre-trial pleadings and motions (L240 dispositive motions), L300 discovery (L310 written discovery, L330 depositions), and later phases for trial preparation and appeal; confirm the current tables against the published UTBMS code set. Activity codes A101 through A111 describe the kind of work, such as planning (A101), research (A102), drafting (A103) and communication split by audience (A105 in-firm, A106 client).

Two structural rules affect label quality. Task codes may appear without activity codes, but activity codes should not appear without a task code, so expect sparse activity labels in older or corporate data. Clients decide whether codes are required, so a sample from a carrier that mandates codes will look very different from one where firms coded voluntarily.

Firm-assigned codes are noisy labels, not ground truth. Timekeepers default to a catch-all code, pick the phase code instead of the task code, or file deposition-prep work under L120 analysis and strategy instead of L330 depositions. Even curated public test sets carry measurable label error [4], so budget for an adjudicated subset where a reviewer re-codes a sample, and report agreement between firm codes and adjudicated codes before you train on them.

Which labels support guideline-violation and bill-review models

Guideline-violation detection needs labels that cite the rule broken and the guideline version in force on the line date. Corporate legal departments apply outside counsel guidelines; insurance defense work follows carrier litigation guidelines, which are often reported to be more prescriptive about staffing, research caps and task pre-approval. A "vague entry" label from one guideline can be acceptable under another, so the guideline identifier and effective date belong on every label.

The strongest bill-review datasets keep the full decision chain: the rule-engine flag, the human reviewer's accept or override, the firm's appeal and the final outcome. Overrides and successful appeals are the hard negatives that keep a model from simply copying the rule engine. Block-billing detection specifically needs narratives where several tasks share one time increment, paired with the reviewer's split or reduction decision.

Public legal AI benchmarks will not fill this gap. LegalBench, for example, is built from hand-crafted legal-reasoning tasks [3]; it does not cover legal-operations records such as invoices, codes or reviewer adjustments. Evaluation sets for billing agents usually have to come from real, licensed operational data with a held-out time window.

Confidentiality, privilege and pseudonymization in invoice narratives

Time-entry narratives are the most sensitive field in the record, because they can disclose litigation strategy, the subject of advice and the identity of witnesses or experts. Under ABA Model Rule 1.6, which states adopt with their own variations, lawyers owe a duty not to reveal information relating to a representation without informed consent or an applicable exception, and must make reasonable efforts to prevent unauthorized disclosure [1]. Disclosure can also raise waiver questions for attorney-client privilege and work product, which Federal Rule of Evidence 502 addresses in federal proceedings [2].

In practice, this means the data holder's counsel, and often the client whose matters appear, should approve release. Treat narratives as privileged-by-default until reviewed; strip or replace party names, case captions, docket numbers, opposing counsel, witness and expert names, and claimant details (insurance defense lines often name the injured party). Medical detail in personal-injury or workers' compensation narratives may also bring health-privacy rules into scope.

Rates and identities are commercially sensitive in their own right. Pseudonymize LAW_FIRM_ID, TIMEKEEPER_ID, TIMEKEEPER_NAME and CLIENT_ID with stable tokens so models can still learn per-firm and per-timekeeper patterns, and consider banding or normalizing LINE_ITEM_UNIT_COST when rate disclosure is the supplier's concern. Keep timekeeper classification (partner, associate, paralegal) because staffing-mix violations depend on it.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

A precise request describes records, labels and controls, not suppliers. The template below shows the level of detail that lets a data holder decide quickly whether it can help.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample specification
Use caseTrain a line-level bill-review model and a UTBMS task-code classifier; hold out the latest 6 months for evaluation
Source formatLEDES 1998B or 1998BI files plus e-billing audit export (CSV or JSON)
Matter typesLitigation (L-codes) and insurance defense; exclude bankruptcy and project codes
Required fieldsNarrative, task, activity and expense codes, units, rate, line date, timekeeper classification
Review labelsReduction amount, reduction reason code, reviewer type, rule-engine flag, override flag
AppealsFirm appeal text, appeal decision, final paid amount per line
Guideline contextGuideline identifier and version effective on each line date
PseudonymizationStable tokens for firm, timekeeper and client; parties, captions and claimants removed from narratives
ApprovalsData holder counsel sign-off on narrative release; documented redaction method and sample check
DeliveryAccess-controlled transfer after license execution; no email attachments

An illustrative joined record might look like this:

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "invoice_id": "INV-TOKEN-0481",
  "law_firm_token": "FIRM-07",
  "timekeeper_token": "TK-1132",
  "timekeeper_classification": "AS",
  "line_date": "2025-03-14",
  "task_code": "L330",
  "activity_code": "A103",
  "units": 3.4,
  "narrative": "Draft deposition outline for [WITNESS_1]; review production; call with [CLIENT_CONTACT] re strategy",
  "review": {
    "rule_flags": ["block_billing", "multiple_activities"],
    "reduction_amount": 0.4,
    "reduction_reason": "block_billing",
    "reviewer_type": "attorney",
    "guideline_id": "OCG-v2024-02"
  },
  "appeal": { "submitted": true, "decision": "partially_upheld", "final_paid_units": 3.2 }
}

Quality checks before you license a sample

Run structural, label and leakage checks on a sample before agreeing to a full extract. The checklist below catches the failures that most often surface after delivery.

  • Parse rate. Every LEDES file parses with the correct header row and field count; reject files with free-text pipes in narratives that shift columns.
  • Code validity. Task, activity and expense codes belong to the declared code set; count lines with blank or catch-all codes by firm and year.
  • Review linkage. Every reduction joins to a submitted line by invoice number and line number; orphan adjustments signal export gaps.
  • Label coverage. Reason codes exist for reductions, not just amounts; overrides and appeals are present, not only automated flags.
  • Guideline versioning. Each label maps to a guideline version effective on the line date.
  • Redaction residue. Scan narratives for names, captions and docket-number patterns that survived pseudonymization.
  • Temporal split. Rate increases and guideline changes make random splits optimistic; evaluate on a later time window.
  • Concentration. Report how much of the data comes from the largest firms and reviewers so one reviewer's habits do not become the model's policy.

Related pages cover adjacent review workflows: workers' compensation and auto medical bill review data for a parallel adjustment-and-appeal structure, e-discovery review coding decisions for privilege-sensitive legal operations data, and invoice line-item extraction data for document-level parsing. For licensed legal text such as case law, see licensing legal content for AI.

SourceX sources operational datasets, including finance and legal workflows, from US companies on request, and manages licensing and ongoing purchases; nothing is held in stock and a request does not guarantee a match. Buyers describe the data they need, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and no method is perfect.

Legal departments and insurers can see sector context on the legal buyers page and the insurance buyers page, and invoice-focused requests on licensing invoices and receipts for AI training. Broader context sits in the industry-specific operational data hub and the AI data guide. To scope a request, describe your legal billing data needs to SourceX.

SourceX works from Find through Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Each dataset is delivered under a license that defines the records, allowed uses, term and delivery, through private, access-controlled workflows after an executed agreement. Start a legal billing and bill-review data request.

Sources

  1. American Bar Association, "Model Rule 1.6: Confidentiality of Information". https://www.americanbar.org/groups/professional_responsibility/publications/model_rules_of_professional_conduct/rule_1_6_confidentiality_of_information/
  2. Legal Information Institute, "Federal Rule of Evidence 502: Attorney-Client Privilege and Work Product; Limitations on Waiver". https://www.law.cornell.edu/rules/fre/rule_502
  3. Stanford Hazy Research, "LegalBench: A collaboratively built large language model benchmark for legal reasoning". https://hazyresearch.stanford.edu/legalbench
  4. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data