Document AI data
Document Dataset Requirements Spec: What to Put in a Request for Document AI Data
Quick answer
A usable document dataset request specifies seven things a supplier can check against its own files: the document classes and their mix, issuer and template diversity, capture conditions (born-digital vs scanned, DPI, color, degradation), page and packet structure, the annotation layers and file formats you need, the privacy method and whether it must preserve layout, and the acceptance tests you will run on a sample. Add permitted uses and counterparty-content handling, and suppliers can say yes, no or partially with evidence.
By SourceX Editorial · Updated
This page covers only the document-specific fields. For the cross-modality request structure, start with how to write a data request for suppliers; for adjacent modalities, compare the multimodal specification template. Everything below assumes your model task is already chosen, as mapped on the document AI data hub.
Why document requests fail without document-specific fields
Generic data requests fail for documents because "50,000 invoices" says nothing about how many vendors, templates, scan conditions or label layers those invoices carry. A set of 50,000 invoices from twelve vendor templates teaches template memorization, not extraction. Public benchmarks show the same gap from the other side: FUNSD is useful precisely because it states its scope (199 noisy scanned forms with entity and link annotations) and its non-commercial research terms [4].
Dataset documentation frameworks give you a proven outline to reverse into a request. Datasheets for Datasets organizes questions under motivation, composition, collection process, preprocessing and labeling, uses, distribution and maintenance [1]; Microsoft's Aether DataDoc and Google's Data Cards turn similar questions into templates that cover upstream sources, annotation methods and intended use [2][3]. Ask the supplier to answer your spec in that structure, and the answer doubles as the datasheet you will need for internal model governance.
Document mix: classes, issuers and templates
Define the class mix as a table of document types with target shares and a tolerance, and count issuers and templates separately from documents. Classification work needs a written taxonomy with tie-break rules, because the RVL-CDIP analysis by Larson et al. found label errors, ambiguous documents that fit several classes, and train/test overlap in a benchmark that many teams still use [6]. If your production stream includes "other" documents, request them explicitly as negatives; see document classification training data for taxonomy design.
Specify these fields for each class:
- Issuer count and concentration cap. For example, at least 300 distinct issuers, with no single issuer above 3% of the class.
- Template or layout variant count. Suppliers rarely track templates, so accept a clustering-based estimate and say which method you will use to verify it.
- Date range. Form revisions change field positions; a W-9 from 2018 and one from 2024 are different layouts.
- Languages and scripts. Include mixed-language pages and numeric formats (1,000.00 vs 1.000,00); see multilingual business document data.
- Page counts and packet structure. Single pages, multi-page documents, or batch-scanned packets that need splitting, covered in page-stream segmentation data.
Capture conditions: born-digital, scanned and photographed
State the capture mix as percentages, because a model trained on clean born-digital PDFs fails on fax-quality scans, and the reverse wastes capacity. Born-digital PDFs carry a text layer; scanned TIFF or JPEG pages need OCR; phone photos add perspective, shadow and blur. Ask suppliers to report per file: source type, resolution in DPI, color mode (bitonal, grayscale, color), compression, page rotation, and whether a text layer exists.
Hybrid formats need their own line. Factur-X and ZUGFeRD invoices embed structured XML inside a PDF/A-3 file next to the rendered invoice [8], which gives you free field labels if the supplier keeps the embedded file intact. Decide whether you want originals as produced, or normalized renderings, and forbid silent re-rasterization. For degraded capture specifically, see real-world document capture conditions.
Annotation layers, formats and agreement thresholds
Name each annotation layer, its format and its guideline version, and set a measurable agreement threshold per layer. Typical layers are page class, layout regions, reading order, table structure (cells, spans, headers), key-value fields with bounding boxes, entity linking, and document-level labels such as split points. DocLayNet is a useful reference because it publishes its layout class definitions and measures inter-annotator agreement on a subset of doubly annotated pages [5]; ask for the same evidence from a supplier.
Formats should be machine-checkable: COCO JSON for regions, PAGE-XML or hOCR for OCR output with coordinates, and a JSON schema for fields that records value, normalized value, page index and box. If labels come from systems of record (ERP or AP entries) rather than human annotators, say so; that approach is covered in documents paired with system-of-record entries. For a single schema covering all layers, use the document annotation schema.
Privacy handling that preserves layout
Require the supplier to state the de-identification method, where it was applied (text layer, image pixels, embedded XML, metadata), and whether replacements preserve string length and format. Black boxes over names break extraction training; format-preserving surrogates (a fake IBAN with a valid checksum, a synthetic address of similar length) keep the layout learnable. Ask for PDF metadata, XMP, embedded attachments and OCR text layers to be scrubbed along with pixels, because redaction that covers the image but leaves the text layer is a common leak.
Name the legal standard that applies. Health documents from HIPAA covered entities and their business associates fall under HIPAA, where de-identification means either Safe Harbor removal of 18 identifier types or an Expert Determination [9][10]; consumer data under the CCPA has its own definition of deidentified information and conditions on the holder [11]. Even then, no method is perfect, so the spec should also prohibit re-identification attempts and require a recorded method you can audit.
Rights and permitted uses for business documents
Write the permitted uses into the request so suppliers can rule themselves in or out before sampling. List training, evaluation, fine-tuning, synthetic data generation from the documents, and retention of derived labels after the term ends. Business documents carry counterparty content (a supplier's invoice sits in a buyer's AP archive), so ask who holds rights to that content and whether the release was approved by the company that holds the files.
Public sets rarely settle this: FUNSD is distributed for non-commercial research [4], which is why commercial teams compare options in public document datasets and commercial use.
Acceptance tests on samples and deliveries
Acceptance should be a sample-based audit with numeric thresholds and a rework clause, agreed before delivery. Label errors are common even in widely cited test sets and can reorder model rankings [7], so treat the supplier's own QA report as input, not proof. Run the same audit on the pre-purchase sample and on each delivery batch, as described in how to run a data pilot with a supplier.
A practical audit checks four things: condition metadata against the files (a claimed 300 DPI page that is really 150 DPI upscaled), field-level label accuracy on a random stratified sample, residual personal data in pixels and text layers, and duplicates or near-duplicates across splits. Define what happens on failure: rework of the batch, replacement records, or rejection.
Document dataset request template
Use the template below as the body of a request; fill each field, and mark any field as "flexible" where you can trade off.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example entry | Why it matters |
|---|---|---|
| Model task | Field extraction from AP invoices and credit memos | Determines which layers matter |
| Class mix | Invoices 70% ±5, credit memos 15% ±5, statements 15% ±5 | Prevents a class missing from training |
| Volume | 20,000 documents, about 45,000 pages | Pages drive annotation cost |
| Issuer diversity | ≥ 400 issuers; no issuer > 3% | Limits template memorization |
| Date range | 2019-2025 | Captures form revisions |
| Languages | English 85%, Spanish 15%; US and EU number formats | Normalization coverage |
| Capture mix | Born-digital 50%, flatbed scan 35%, fax 10%, phone photo 5% | Matches production stream |
| Image spec | Originals as produced; scans ≥ 200 DPI; report DPI, color, compression per page | Detects upscaling and lossy recompression |
| Packet structure | 10% delivered as unsplit multi-document PDFs with split labels | Trains page-stream segmentation |
| Annotation layers | Page class; fields with boxes (JSON schema v1.2); line items as table cells | Defines deliverable |
| Guideline version | Buyer guideline v3.1, attached | Stops label drift |
| Agreement threshold | Field-level exact match ≥ 95% on double-annotated 5% subset | Measurable quality |
| Privacy method | Format-preserving surrogates in pixels, text layer and metadata; method documented | Keeps layout learnable |
| Permitted uses | Training, evaluation, derived labels retained | Rules suppliers in or out |
| Sample | 200 stratified documents before contract | Basis for acceptance tests |
| Acceptance | Audit 2% per batch; < 2% field errors; zero residual direct identifiers; rework within agreed period | Defines remedy |
How SourceX handles document requests
SourceX sources operational datasets, including documents and finance and legal workflows, from US companies on request; nothing is held in stock, and a request does not guarantee a match. Buyers describe the data rather than the businesses, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, with personal details removed or replaced and the method recorded. You can send a document data request to SourceX using the fields above.
Request business documents for your document AI program
SourceX sources documents and other operational records from US companies on request, rights-reviews each dataset, and manages licensing from assessment through ongoing purchases. Describe the classes, conditions and labels you need, and SourceX looks for businesses that hold them. Start your document data request.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Frequently asked questions
How large should the pre-purchase sample be?
Large enough to stratify across every class and capture condition in the spec, with at least a few dozen documents in the rarest cell. A sample drawn only from the cleanest born-digital files hides the conditions most likely to fail.
Should I send my annotation guideline with the request?
Yes, if you need labels delivered. A versioned guideline with examples and edge-case rules lets suppliers estimate effort and lets you reject labels that follow a different convention.
Can synthetic documents fill gaps in the mix?
They can cover rare templates, but generated documents tend to miss real capture noise and issuer quirks; see synthetic vs real documents before substituting.
Sources
- Gebru et al. (arXiv:1803.09010), "Datasheets for Datasets" (2018). https://arxiv.org/abs/1803.09010v7
- Microsoft Research, "Aether Data Documentation Template" (2022). https://www.microsoft.com/en-us/research/uploads/prod/2022/07/aether-datadoc-082522.pdf?lang=fr_ca
- Pushkarna, Zaldivar, Kjartansson (Google Research; FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- Jaume, Ekenel, Thiran (arXiv:1905.13538), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
- Pfitzmann et al. (arXiv:2206.01062), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
- Larson et al. (arXiv:2306.12550), "On Evaluation of Document Classification using RVL-CDIP" (2023). https://arxiv.org/html/2306.12550v1
- Northcutt, Athalye, Mueller (arXiv:2103.14749), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- eConnect, "Hybrid invoice formats: PDF with embedded XML". https://accp.econnect.eu/en/docs/learn/document-formats/basics/hybrid-invoices
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- California Legislative Information, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.