Skip to content

Industry-specific operational data

Freight claims data for AI: licensing OS&D and cargo claim files

Quick answer

Freight claims data for AI training is the complete claim file behind a cargo loss or damage event: the claim filing and amount, bill of lading, delivery receipt with OS&D notations, inspection photos, invoice proving value, carrier acknowledgment and the final paid, declined or compromised disposition. As of October 2026 there is no known public dataset of these files, so buyers license them from shippers, brokers and 3PLs that hold them, and the value depends on intact packets, coded outcomes and consistent join keys.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What a freight claim record should contain

The unit you license should be a claim packet, not a row in a claims log, because triage, drafting and liability models learn from the relationship between documents and the outcome. A claims-management export (date filed, amount, status) trains almost nothing on its own. Vendors that automate evidence assembly already treat the packet as the working object [1], and freight-document extraction guides stress that fields must agree across the bill of lading, packing list and commercial invoice [4].

A complete packet usually holds the claim form or letter with the claimed amount, the bill of lading (including any released-value declaration), the delivery receipt or proof of delivery with exceptions noted, inspection reports and damage photos, the commercial invoice or replacement-cost evidence, the freight bill, the carrier's acknowledgment and correspondence, and the final disposition with amount paid. Ask whether salvage records travel with the packet, since salvage value changes the net settlement.

Separate three record types before you specify anything. Cargo claims against carriers (shipper or broker versus carrier) follow the US carrier-liability framework below. Cargo-insurance claims and subrogation recoveries follow policy terms and insurer workflows; if those are what you need, see subrogation recovery files and the broader insurance claims datasets.

The US liability frame your model has to learn

Interstate motor carrier claims are governed by the Carmack Amendment, 49 U.S.C. 14706, which generally makes the carrier liable for actual loss or injury to the property, subject to released-rate limits the shipper agreed to on the bill of lading. Under section 14706(e), carriers cannot set a claim-filing period shorter than nine months or a suit period shorter than two years from written notice of disallowance; confirm the current statute text before you encode these rules. A drafting or triage model that ignores the released value or the carrier's tariff deadline will produce claims that look correct and get declined.

49 CFR Part 370 sets the processing rules that create your timestamps and labels. Section 370.3 treats a written or electronic communication as a claim only when it identifies the shipment, asserts liability and demands a specified or determinable amount, so a delivery-receipt notation or a phone call is not a claim on its own. The part also sets carrier acknowledgment and disposition windows (acknowledgment within 30 days under section 370.5 and payment, declination or a firm compromise offer within 120 days under section 370.9); verify the intervals against the eCFR as of October 2026 before relying on them.

These rules give you free, high-quality supervision. Fields such as claim_received_at, acknowledged_at and disposition_at support deadline-compliance models, and the presence or absence of a sum-certain demand is a clean classification target for "is this a valid claim filing."

Label quality: dispositions and declination reasons

The disposition is the most valuable label in a claim file, and it is also the field most likely to be inconsistent. Paid, declined and compromised are easy to code; the reason behind a declination usually is not. Common reasons include concealed damage reported after a clean delivery receipt, inadequate packaging, shipper load and count, released-value limits, late filing and missing proof of value, and many claims systems store them as free text in adjuster notes.

Ask each supplier for its reason taxonomy, how compromise settlements are coded (percentage of claimed amount, flat offer, freight-charge credit), and whether reopened or appealed claims overwrite the original outcome. Treat outcome fields as noisy until verified; the method in verifying outcome labels in operational records applies directly. Logistics practitioners report that historical data is plentiful but rarely structured for training, and that validating it for gaps and outliers is the slowest step [2].

Join keys that turn claims into causal training data

Claims become causal training data when you can link them to the shipment events that preceded them. Specify PRO number, BOL number, shipment or load ID, and carrier SCAC (or a tokenized stand-in) as join keys, tokenized consistently across every table so the joins survive de-identification.

With those keys, a claim can be connected to tracking exceptions, appointment changes, temperature logs and dispatcher messages. That lets you train models that predict claim likelihood at the exception stage, not only after delivery. Related sources include logistics exception records, the general supply chain and logistics datasets page and freight broker and carrier communications data.

Illustrative claim packet record

The record below shows the structure to request; field names are examples you can adapt to your extraction schema.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueNotes for the buyer
claim_idCLM-TKN-8F21Tokenized; stable across tables
pro_number_tkn / bol_number_tknPRO-TKN-44A0 / BOL-TKN-19C3Join to shipment and exception events
modeLTLLTL, FTL, intermodal, parcel
claim_typedamageshortage, damage, loss, delay
os_d_notation"2 cartons crushed, subject to inspection"From delivery receipt; empty means clean receipt
released_value_declaredtrueDrives limitation of liability
claimed_amount_band2,500-5,000 USDValue banded, not exact
documents[]claim_letter.pdf, bol.pdf, pod.tiff, invoice.pdf, photos/*.jpgOriginal files plus OCR text
claim_received_at / acknowledged_at / disposition_atISO 8601 datesShifted consistently if required
dispositioncompromisedpaid, declined, compromised, withdrawn
declination_reason_codenullMapped to supplier taxonomy
settlement_pct_of_claim0.6Coded outcome for regression

Rights and privacy issues in freight claim files

Freight claim files carry more commercial sensitivity than personal data, and both need handling before delivery. Shipper and consignee identities, product descriptions, invoice values and carrier names can reveal customer lists, pricing and carrier performance, so expect tokenization of parties and value banding. Personal data appears in signatures on delivery receipts, driver names and phone numbers, and in damage photos that capture drivers or dock workers.

Photos deserve a specific review. Photographs alone fall outside BIPA's definition of a biometric identifier, but scans of face geometry fall inside it [5], so ask how faces in dock and trailer photos are handled and whether any face processing was run on them. If you train generative models offered in California, confirm the documentation you will need to post about training data before you sign [6].

How to specify and license freight claim files

A good request describes the claim population, the packet contents, the outcome coding and the permitted uses. Use this checklist when you write it.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Population: modes (LTL, FTL, intermodal), claim types, date range, and whether the holder is a shipper, broker, 3PL or carrier.
  • Packet completeness: share of claims with BOL, POD, invoice and photos attached; how missing documents are flagged.
  • Outcomes: disposition values, reason taxonomy, compromise coding, and treatment of reopened claims.
  • Keys and timestamps: tokenized PRO, BOL and load IDs; received, acknowledged and disposition dates.
  • De-identification: party tokenization, value banding, signature and face redaction method, and sample checks.
  • Rights: confirmation that the holder may license documents it received from counterparties, plus allowed uses (SFT, evaluation, extraction training) and term.

Expect to source these files privately; vendors building logistics AI argue the open web is not a viable training source for this domain [3]. For neighboring claim-file patterns in other industries, compare claim denial prediction training data and the industry-specific operational data hub. Owner pages for freight brokerages, third-party logistics and logistics describe each segment.

How SourceX handles freight claims data requests

SourceX sources operational datasets from US companies on request and manages the licensing process; nothing is held in stock and a request does not guarantee a match. You describe the claim data you need on the SourceX buyers page, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, after personal details are removed or replaced and a sample is checked.

Request licensed freight claim files

If you are building claim triage, drafting or extraction models, describe the claim packets, outcomes and uses you need. SourceX works through Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Describe your freight claims data request.

Sources

  1. Kognitos, "Freight Claim Evidence Package Automation". https://www.kognitos.com/use-case/freight-claim-evidence-package-automation
  2. TLI Magazine, "Transport Logistics International, Volume 12, Issue 4 (p. 16)". https://magazine.tlimagazine.com/transport-logistics-international-volume-12-issue-4/0766956001732009956/p16
  3. The Loadstar, "Raft platform volumes surpass AI learning milestones". https://theloadstar.com/ls_press_release/raft-platform-volumes-surpass-ai-learning-milestones/
  4. Imagetotable.ai, "The Complete Guide to Shipping & Freight Document Extraction". https://imagetotable.ai/blog/complete-guide-shipping-freight-document-extraction
  5. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  6. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data