Skip to content

Document AI data

Commercial Invoice, Packing List and Certificate of Origin Data for Trade Document AI

Quick answer

A usable trade document extraction dataset is a set of real import packets (commercial invoices, packing lists and certificates of origin, ideally with the bill of lading) where each document carries field labels tied to what was actually filed with customs. Public datasets rarely cover this, so teams building entry-preparation or compliance models usually license operational records from importers, brokers or forwarders, with counterparty pricing masked and every document tagged by language, issuer country and template.

By SourceX Editorial · Updated

What "trade document" data means for customs automation

For a customs or trade-compliance team, trade document data means the paperwork that supports an import entry, not financial trade confirmations. The term is ambiguous: one extraction paper labels derivatives trade confirmations with fields such as trade date, price and volume [6], which is useless for entry preparation. Write requests around the document types you need (commercial invoice, packing list, certificate or certification of origin, bill of lading, arrival notice) and the tasks you are training.

Public options are thin. Benchmark work notes that business documents are scarce in open datasets because their content is legally protected or commercially sensitive [4], and the public multi-task sets that do exist, such as BuDDIE's 1,665 documents, carry non-commercial licenses [5]. See which public document AI datasets allow commercial use before building on an academic set.

Fields to label on each document type

The label schema should follow the regulatory data elements, because those are what your model ultimately has to produce. Under 19 CFR 141.86, an invoice of imported merchandise must state the port of entry, the date and place of sale, the seller and buyer, and must name a responsible employee of the exporter. Section 142.6 adds an adequate merchandise description, quantities, values and the eight-digit HTSUS subheading.

Certificates of origin are less standardized than buyers expect. For USMCA, the agreement's Annex 5-A sets nine minimum data elements: who certifies, the certifier, exporter, producer, importer, description and HS classification of the good, origin criterion, blanket period, and authorized signature and date. No prescribed form exists, and the certification can sit on the invoice itself or another document, in writing or electronically, under 19 CFR 182.12. Your classifier therefore has to find origin statements embedded in invoices, not only standalone forms.

Three fields deserve explicit normalization rules:

  • Incoterms: label the rule and the named place separately (for example "FCA" and "Shenzhen"), and record the Incoterms version where printed, since ICC revises the rules periodically and, as of October 2026, Incoterms 2020 is the current set [1].
  • Tariff codes: record the code exactly as printed, its digit length and a normalized HTSUS form. Foreign suppliers often print 6-digit HS codes, while US entries use HTSUS subheadings and statistical suffixes [2].
  • Values and currency: keep unit price, extended value, currency code and any discounts, freight or assists as distinct fields so valuation logic can be tested.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "packet_id": "PKT-00412",
  "documents": [
    {"doc_id": "INV-1", "doc_type": "commercial_invoice", "language": "zh-Hans+en", "issuer_country": "CN", "pages": 2, "template_id": "seller-tpl-07"},
    {"doc_id": "PL-1", "doc_type": "packing_list", "language": "en", "issuer_country": "CN", "pages": 1},
    {"doc_id": "COO-1", "doc_type": "origin_statement_on_invoice", "agreement": "none", "location": "INV-1 p2"}
  ],
  "fields": {
    "INV-1.seller_name": {"value": "[MASKED_SUPPLIER_01]", "bbox": [84, 120, 410, 142], "page": 1},
    "INV-1.incoterm_rule": {"value": "FOB", "named_place": "Ningbo", "version_printed": null},
    "INV-1.line_items[0].hs_code_printed": {"value": "8504.40", "digits": 6},
    "INV-1.line_items[0].htsus_as_filed": {"value": "8504.40.9580", "source": "entry_record"},
    "INV-1.line_items[0].quantity": {"value": 1200, "uom": "PCS"},
    "INV-1.line_items[0].unit_price": {"value": "[MASKED_PRICE]", "currency": "USD"},
    "INV-1.country_of_origin": {"value": "CN"}
  },
  "consistency_checks": [
    {"rule": "invoice_qty_eq_packing_qty", "status": "fail", "detail": "INV 1200 PCS vs PL 1150 PCS"},
    {"rule": "gross_weight_pl_eq_bol", "status": "pass"}
  ],
  "label_provenance": "keyed_by_broker_then_matched_to_filed_entry"
}

Cross-document consistency is its own labeled task

Discrepancy detection between invoice, packing list and bill of lading is a separate supervision target, and it is often where compliance agents earn their keep. Market practice describes a cross-border shipment as a packet of five to seven documents whose quantities, weights, parties and descriptions must agree [7]. A model that extracts each document perfectly can still miss that the invoice says 1,200 pieces and the packing list says 1,150.

Ask for packets, not loose documents, and for consistency labels at the rule level (quantity, gross and net weight, carton count, consignee, marks and numbers). The most valuable signal is the broker's resolution: which document was corrected, and whether the entry was later amended through a CBP post-summary correction [3]. Those correction trails are covered in more depth on document extraction correction logs from production IDP.

Where good labels come from

The strongest labels for trade documents are values matched to what was actually filed, not fresh annotation. A broker's entry data provides a system-of-record answer for description, quantity, value, origin and tariff classification, the same pattern described in documents paired with system-of-record entries. Fresh human annotation still matters for bounding boxes and for fields that never reach the entry, such as marks and numbers.

Classification is a related but distinct problem. At least one public HTS model trains on CBP rulings text from CROSS rather than on shipping documents [8], which teaches the tariff schedule but not how a supplier in Vietnam describes a product on an invoice. If your goal is invoice-to-HTSUS prediction, request the printed description, the filed classification and the classifier's notes together.

Multilingual and template variety from foreign suppliers

Trade documents arrive in many languages, scripts and layouts, so coverage metadata matters as much as volume. Record language (including mixed-language documents), issuer country, template or seller identifier, scan versus native PDF, and page count for every file. USMCA certifications may be in English, Spanish or French, and Asian-origin invoices frequently mix Chinese, Japanese or Korean with English.

Split train and test sets by seller template, not by page, or your evaluation will reward memorized layouts. For script-specific OCR concerns see multilingual and mixed-script business document data, and for single-document invoice depth see invoice line-item extraction data.

Confidentiality, masking and licensing questions

Trade packets expose counterparties' commercial terms, so masking must be agreed before any sample moves. Supplier names, unit prices, bank details and contact names on invoices are confidential to the importer and its suppliers, and a licensed dataset should state what was masked, how, and whether masking is consistent across a packet so cross-document checks still work.

Buyer checklist before signing:

  • Who owns the documents, and does the importer or broker have the right to license supplier-issued paperwork for model training?
  • Are prices replaced with consistent surrogates (preserving arithmetic) or removed?
  • Are personal names and signatures on certificates removed or replaced, and how was that verified on a sample?
  • Which agreements and origin regimes are represented (USMCA, other FTAs, non-preferential)?
  • Is a dataset card provided covering sources, collection, annotation and intended use [9]?

Letter-of-credit document sets raise different issues; see trade finance document data. Domestic freight paperwork is covered by freight document extraction data, and the full task map sits on the document AI data hub. Industry context is on supply chain and logistics datasets and logistics buyers.

How SourceX approaches trade document requests

SourceX sources operational datasets from US companies on request, rather than holding stock, so a request for import packets does not guarantee a match. You describe the data you need, such as invoice, packing list and origin packets with filed entry values, and SourceX looks for US businesses that hold it; every release is approved by the supplying company. You can describe your trade document requirements to SourceX.

Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Request trade document extraction data

SourceX handles the commercial process from Find and Assess through Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Start a trade document data request with SourceX.

Sources

  1. International Chamber of Commerce, "Incoterms Rules FAQ". https://www.iccwbo.org/incoterms_faq/
  2. U.S. International Trade Commission, "Harmonized Tariff Schedule of the United States: Preface". https://usitc.gov/publications/docs/tata/hts/bychapter/1501preface.pdf
  3. U.S. Customs and Border Protection, "Post Summary Correction". https://cbp.gov/trade/programs-administration/entry-summary/post-summary-correction
  4. arXiv, "Business Document Information Extraction: Towards Practical Benchmarks" (2022). https://arxiv.org/pdf/2206.11229
  5. arXiv, "BuDDIE: A Business Document Dataset for Multi-task Information Extraction" (2024). https://arxiv.org/pdf/2404.04003
  6. arXiv, "Information Extraction from Visually Rich Documents with Font Style Embeddings" (2021). https://arxiv.org/pdf/2111.04045
  7. Imagetotable.ai, "The Complete Guide to Shipping & Freight Document Extraction". https://imagetotable.ai/blog/complete-guide-shipping-freight-document-extraction
  8. Hugging Face, "atlas-llama3.3-70b-hts-classification model card". https://huggingface.co/flexifyai/atlas-llama3.3-70b-hts-classification/blob/main/README.md
  9. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data