Skip to content

Document AI data

Freight Document Extraction Data: Bills of Lading, PODs and Rate Confirmations

Quick answer

A useful freight document extraction dataset pairs real bills of lading, proofs of delivery and rate confirmations from many shippers, carriers and brokers with field-level labels, ideally taken from the TMS shipment record each document belongs to. Public form benchmarks do not cover freight. The hard cases are template diversity, driver-phone photos, and handwritten POD exception notes about shortages and damage. Buy for those cases, insist on cross-document links, and require identifiers of drivers and consignees to be removed before delivery.

By SourceX Editorial · Updated

Why public datasets will not train a freight extraction model

As of October 2026 there is no widely used public, labeled corpus of US freight paperwork at useful scale, so teams either license real documents from operators or annotate their own. Searches for this need mostly return IDP vendor guides and product pages, such as [1] and [2], rather than datasets. The closest public benchmarks are general forms: FUNSD, for example, has 199 noisy scanned forms from marketing, advertising and science domains [4]. That is useful for pretraining a key-value linker, but it contains no PRO numbers, NMFC freight classes, seal numbers or signed delivery receipts.

Synthetic BOLs fill some gaps but fail on the failure modes that matter in production: carbon-copy bleed, stamp overlap, skewed photos taken on a dock, and handwriting written across printed fields. The trade-offs are covered in synthetic vs real documents for document AI. For the broader task map, start at the document AI datasets hub.

Which fields a freight label schema should cover

The schema should start from what federal rules and the carrier's billing process force onto the paper, then add the operational references your model must reconcile. Motor carrier bills of lading under 49 CFR Part 373 identify the consignor and consignee, origin and destination, package count, freight description, and weight where it affects rating. Hazmat shipping papers under 49 CFR 172.202 add the UN identification number, proper shipping name, hazard class and packing group. Everything else, from PRO and PO numbers to accessorials, is commercial convention that varies by template; confirm the current regulatory text on eCFR before encoding it as a validation rule.

Rate confirmations share a charge structure with freight invoices, and vendors split freight invoice fields into document-level, charge-level and shipment-level data, which is a sensible label hierarchy for both [3]. See key-value extraction labels and invoice line-item extraction data for annotation conventions that carry over.

Illustrative example: invented to show structure; it does not describe an available dataset.

DocumentCore fields to labelHard cases to over-sample
Bill of lading (straight, VICS-style or carrier form)Shipper, consignee, bill-to, pickup and delivery addresses, BOL number, PRO, PO and reference numbers, pieces, handling units, weight, NMFC item and class, hazmat line (UN number, proper shipping name, class, packing group), seal numbers, special instructionsMultiple commodity lines, hazmat mixed with non-hazmat, handwritten weight corrections, third-party billing
Proof of delivery (signed BOL or delivery receipt)Delivery date and time, receiver printed name and signature presence, piece count received, exception flag, exception type (shortage, overage, damage, refused), exception free text, stampsHandwritten "2 cartons crushed" notes, "subject to count" clauses, signatures over printed text, phone photos at angles
Rate confirmation (broker to carrier)Broker and carrier names, MC/DOT numbers, load number, pickup and delivery windows, equipment type, linehaul rate, fuel, accessorials (detention, lumper, TONU), payment termsAmended rate cons, multi-stop loads, accessorials buried in notes

How to get labels without annotating every page

The cheapest high-quality labels come from joining each document to the system-of-record entry it produced. A TMS shipment record already holds the carrier, PRO, pieces, weight, rate and accessorials that someone keyed from the paperwork; a freight claim or OS&D record holds the outcome of the POD exception. Pairing them yields weak labels at scale and lets you measure where keying staff corrected the paper. The method and its pitfalls (keyed values that differ from what is printed, late adjustments, partial loads) are covered in documents paired with ERP records.

Cross-document agreement is where freight extraction actually breaks: a cross-border shipment packet can run five to seven documents, and making the fields agree across them is the hard part [1]. Vendors frame the task as classify, extract, then flag mismatches across the packet [2]. Ask for packets with stable shipment keys so you can train and evaluate the reconciliation step, not only single-page extraction. If your production pipeline has human reviewers, their correction logs are a second label source; see IDP human correction data.

Handwritten POD exceptions: the target that pays

POD exception notes are a small slice of a freight corpus and among the most valuable, because they drive freight claims, carrier chargebacks and customer disputes. A model that reads printed BOL fields at high accuracy but misses "short 1 pallet" written across the signature line fails the use case. Ask suppliers to estimate the share of PODs with any annotation, then over-sample those documents and label the note text, the exception class and the affected line.

Evaluate this slice separately. Report field-level accuracy for printed fields, exception detection recall, and exception classification accuracy, with a held-out set drawn from carriers and shippers absent from training. The setup is described in document extraction evaluation ground truth.

Diversity and capture conditions to specify in a request

Template and capture diversity matter more than raw page count. Specify them as distributions, not adjectives.

Illustrative example: invented to show structure; it does not describe an available dataset.

request: freight-document-extraction
documents: [bill_of_lading, proof_of_delivery, rate_confirmation]
modes: [LTL, FTL]
diversity:
  min_distinct_shippers: 50
  min_distinct_carrier_templates: 30
  min_distinct_broker_ratecon_templates: 15
capture:
  flatbed_scan_share: 0.4
  mobile_photo_share: 0.4      # driver app and dock photos
  fax_or_multi_generation_copy_share: 0.2
must_include:
  - pod_with_handwritten_exception: ">= 15% of PODs"
  - hazmat_bol_lines: true
  - amended_rate_confirmations: true
labels:
  source: tms_shipment_record_join
  key: shipment_id
  fields: [pro_number, bol_number, pieces, weight_lb, nmfc_class, linehaul_usd, accessorials]
privacy:
  remove: [driver_name, driver_phone, receiver_signature_image, consignee_contact]
delivery: [pdf_or_tiff_originals, page_images, labels_jsonl, data_card]

Pair this with a data card that records upstream sources, collection and annotation methods and intended use [5]. Mixed-language paperwork from cross-border lanes belongs in a separate spec; see multilingual business document data. Commercial invoices, packing lists and certificates of origin sit in trade document extraction data.

Privacy and confidentiality checks before anything ships

Freight paperwork carries personal data and commercially sensitive terms, so both need handling before delivery. Driver names and phone numbers, receiver signatures, residential consignee addresses and contact names should be removed or replaced, with the method recorded and a sample checked. Shipper and customer identities, lane rates and accessorial terms can also be confidential under the supplier's own contracts, so confirm which names the supplier may release and whether rates must be masked or bucketed.

At SourceX, personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery.

Where these documents come from

Real freight paperwork sits with brokers, 3PLs, shippers and carriers that have scanned BOLs, PODs and rate confirmations for billing and claims for years. The dataset-level view of linked BOL, POD and invoice chains lives on the supply chain and logistics datasets page; buyer context for operators is on freight brokerages and third-party logistics.

SourceX sources operational datasets from US companies on request; nothing is held in stock and a request does not guarantee a match. Buyers describe the data, not the businesses, and every release is approved by the supplying company. You can describe your freight document requirements to SourceX using the request structure above.

Request freight document extraction data

SourceX looks for US businesses that hold the documents you describe, then runs Find, Assess, Agree, Transact and Manage, with nothing contracted until a supplier agrees. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Describe the freight documents you need.

Sources

  1. Imagetotable.ai, "The Complete Guide to Shipping & Freight Document Extraction". https://imagetotable.ai/blog/complete-guide-shipping-freight-document-extraction
  2. InvoiceDataExtraction.com, "Freight Document Extraction: Logistics Parser Guide". https://invoicedataextraction.com/blog/freight-document-extraction
  3. Parsio, "How to Extract Data from Freight Invoices Automatically". https://parsio.io/blog/how-to-extract-data-from-freight-invoices-automatically/
  4. Jaume, Ekenel, Thiran (arXiv), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  5. Pushkarna, Zaldivar, Kjartansson (Google Research, FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data