Document AI data
ACORD Form Data for Insurance Submission Intake Models
Quick answer
An ACORD form extraction dataset worth buying is a set of real, filled commercial submissions (ACORD 125 applications with their line-of-business sections such as 126 and 140, plus loss runs, schedules of values and broker emails) where every field is labeled with its value, location, form number and edition date. We found no public set of this kind as of October 2026, so teams license it from the carriers, MGAs and brokerages that hold submission archives.
By SourceX Editorial · Updated
Why there is no public ACORD training set
Public search for ACORD extraction data returns extraction products, not datasets, because the vendors behind them train on private collections of real forms [1][3][4]. Vendors list ACORD 25 certificates and ACORD 125 commercial applications among the most-used forms, and advertise models for ACORD 25, 125, 126, 130 and 140 [1][3]. Their accuracy figures rest on form sets you cannot inspect, so they are not a benchmark you can reproduce.
General form benchmarks do not fill the gap. FUNSD contains 199 noisy scanned forms from domains such as marketing and scientific reports, annotated for OCR, layout and entity linking [5]. It is useful for pretraining a key-value linker, but it has no insurance vocabulary, no multi-page applications and no checkbox-heavy coverage grids. For the broader landscape of public and licensed sets, see the document AI datasets hub.
What a usable submission packet contains
A commercial submission is a packet, not a form, and models trained on isolated ACORD pages fail on real broker inboxes. A typical packet mixes the ACORD 125 with line sections (126 general liability, 140 property, 127 business auto, 130 workers compensation), carrier supplemental applications, a schedule of values spreadsheet, three to five years of loss runs, driver and vehicle lists, and the broker's cover email with its attachments.
That mix drives three distinct tasks, and each needs its own labels:
- Page-stream segmentation and classification: where each document starts and what it is, covered in depth on page-stream segmentation data.
- Field extraction: named insured, mailing and premises addresses, FEIN, SIC or NAICS code, entity type, years in business, requested effective date, limits and deductibles, prior carrier and premium history, and loss summaries.
- Triage: appetite fit, completeness and missing items, which links to underwriting decision rationale data for agent workflows.
Loss runs are their own extraction problem with carrier-specific layouts; buy them alongside the ACORD forms using the guidance on loss run extraction data.
Edition dates and layout drift
Record the form number and edition date for every document, because extraction vendors support several ACORD revision years and templates trained on one edition break on another [2]. The edition stamp printed in the form footer is the cheapest feature for routing a page to the correct schema. Ask suppliers for the edition distribution before you agree to anything, since an archive dominated by one revision will overstate accuracy.
Rendering format matters as much as edition. Fillable PDFs exported from agency management systems carry clean text layers, while scanned, faxed or photographed forms carry skew, stamps and handwriting. Checkbox grids for coverages and entity type deserve dedicated labels; see checkbox and selection mark data.
Label schema buyers should specify
Ask for field-level labels that tie each value to its page, bounding box, form number and edition, with a separate flag for values a human corrected. Labels derived from the policy admin or rating system ("what the underwriter actually keyed") are stronger than fresh annotation, because they reflect production truth; the pattern is covered on documents paired with system-of-record entries.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example key | Value type | Label notes |
|---|---|---|---|
| Form identity | form_number, edition_date | string, date | From footer; one per page |
| Named insured | named_insured | string | Primary plus additional named insureds as a list |
| Tax ID | fein | string | Masked or tokenized before delivery |
| Classification | naics_code, sic_code | string | Keep as written; flag invalid codes |
| Premises | locations[] | array of objects | Address, occupancy, building and contents values |
| Limits | limits{} | object | Per coverage, with deductibles |
| Prior coverage | prior_carriers[] | array | Carrier, policy period, premium |
| Loss summary | losses[] | array | Date, type, paid, reserved, status |
| Coverage selections | coverage_checkboxes{} | boolean map | Checked, unchecked, ambiguous |
| Provenance | label_source | enum | system_of_record, annotator, corrected |
Use a held-out set split by broker and by submission date, not by page, so the same insured never appears in both train and test. Field-level scoring methods are covered on document extraction evaluation sets.
Rights, confidentiality and personal data in submissions
Treat the form template and the filled data as separate rights questions: the layouts belong to the forms publisher (ACORD), while use of the answers is governed by the agreements among the insured, the broker and the carrier. Confirm both before you license, and ask whether the supplier's agreements with brokers allow use of submissions beyond underwriting.
Most submission content is commercially confidential business data such as revenue, payroll, loss history and pricing. Personal data still appears: sole proprietors who enter a Social Security number in the FEIN field, driver lists with license numbers and dates of birth, and owner contact details. Gramm-Leach-Bliley privacy rules cover nonpublic personal information about consumers held by financial institutions [6]; insurers comply under state insurance regulation, and most commercial submission data concerns businesses rather than consumers, so confirm which records are in scope. California defines deidentified data with specific conditions the holder must meet [7].
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Buyer request checklist for ACORD submission data
A precise request shortens assessment because suppliers can check their archives against it.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Lines of business and forms: for example ACORD 125 with 126 and 140, plus supplementals.
- Edition coverage: which revision years, and the share of each.
- Channel and format: broker email with attachments, portal upload, fax, or scan; native PDF vs image.
- Packet completeness: loss runs, schedules of values, driver lists included or not.
- Labels: field-level values with bounding boxes, source of truth, correction flags.
- Outcomes: quoted, declined or bound, if triage models are in scope.
- De-identification: which fields are masked or replaced, and how a sample is checked.
- Allowed uses: training, fine-tuning, evaluation, and term, written into the license.
How SourceX sources ACORD submission packets
SourceX sources operational datasets from US companies on request, including document and finance workflows, and manages licensing and ongoing purchases. Nothing is held in stock: you describe the forms and labels you need, SourceX looks for US businesses that hold them, and a request does not guarantee a match. You can start that description on the SourceX buyers page.
The process runs Find, Assess (the data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.
Related SourceX pages cover underwriting files for AI training, insurance claims datasets and buyers in insurance brokerages.
Request ACORD form extraction data
SourceX sources filled submission documents from US businesses on request, under a license that defines records, uses, term and delivery, with every release approved by the supplying company. Describe the forms, editions and labels you need on the SourceX buyers page.
Frequently asked questions
Can I train on blank ACORD forms instead of filled ones?
Blank templates help with layout and template matching but not with extraction. Real submissions carry handwriting, overflow onto supplemental pages, inconsistent addresses and unchecked boxes that blank or synthetic forms rarely reproduce.
How many editions should an evaluation set cover?
Cover every edition your production inbox actually receives, weighted to match it. Measure accuracy per edition so a regression on an older revision is not hidden by the average.
Should broker emails be included?
Include them if your model must segment attachments or read appetite signals from the cover note. Emails carry broker names and contact details, so specify how they are de-identified.
Sources
- Base64.ai, "Extract Data From ACORD Forms". https://Base64.Ai/features/data-extraction-api/acord
- Sensible, "Extract ACORD forms to structured JSON". https://www.sensible.so/doc-types/acord-forms
- Docsumo, "Get 99%+ accurate data from ACORD forms with AI". https://docsumo.com/automated-acord-form-processing
- UiPath Marketplace, "Acord Form & Unique Document Data Extraction". https://marketplace.uipath.com/listings/form-data-extraction-by-using-document-understanding
- Jaume, Ekenel, Thiran, "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
- Consumer Financial Protection Bureau, "CFPB Laws and Regulations: GLBA Privacy" (2016). https://files.consumerfinance.gov/f/documents/102016_cfpb_GLBAExamManualUpdate.pdf
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.