Skip to content

Document AI data

Synthetic vs Real Documents for Document AI: Where Generated Data Breaks

Quick answer

Synthetic document data is strong for OCR text-rendering pre-training, fixed-layout forms and bootstrapping extraction schemas, because a generator emits perfect boxes, transcriptions and field labels at near-zero marginal cost. It breaks where production variance lives: the long tail of vendor templates, handwriting, fax and phone-camera degradation, stamps, multi-page packets and the statistical quirks of real field values. Train on synthetic where it is cheap and accurate, license real documents for the tail, and always measure on a held-out set of real documents.

By SourceX Editorial · Updated

What synthetic document generators actually produce

Synthetic document generators produce rendered pages plus exact ground truth, but only for the distribution their author encoded. Most fall into three families: text renderers that paint lines of fonts over paper textures (the approach behind SynthDoG, released with the Donut model), schema-driven generators, and layout generators. AWS Labs' SEED pipeline, for example, turns a JSON schema into a fictional PDF with matching ground truth [3], and recent research builds graph-based synthetic layouts to diversify the arrangement of page components [1].

Each family inherits the generator's assumptions. A schema-driven invoice generator knows the fields you listed (invoice_number, due_date, line_items[].amount) and the ranges you sampled from; it does not know that a regional distributor prints totals in a rotated box beside a carbon-copy watermark. Vendors describe synthetic document data as the answer when real documents are scarce or sensitive [4], which is true for coverage of known structure and false for coverage of unknown structure.

The useful question is therefore not "synthetic or real" but "which variance does my model need to see that my generator cannot express."

Where synthetic documents perform well

Synthetic documents perform well when the target variance is enumerable and the label is a deterministic function of the rendering. Four cases hold up in practice.

  • OCR and text-recognition pre-training. Rendering millions of word and line crops across fonts, sizes, kerning and backgrounds teaches glyph recognition with free, exact transcriptions.
  • Fixed government and industry forms. W-9s, ACORD certificates, CMS-1500 claims and bills of lading with a stable template can be filled with sampled values and rendered faithfully.
  • Schema bootstrapping. Before any real data arrives, a generator lets you test the extraction head, output JSON schema and post-processing on documents whose ground truth is known [3].
  • Privacy-safe debugging and demos. Fictional pages let engineers inspect failures without handling personal data or restricted records.

Synthetic data also helps rare-class augmentation, such as generating extra examples of a seldom-seen form version, provided you verify the rendered output matches real versions of that form.

Where generated documents break: five failure modes

Generated documents break where the production distribution is open-ended, physically produced or behaviorally correlated. These are the failure modes to test for before you scale a synthetic-heavy recipe.

1. Template diversity. Accounts-payable teams see thousands of supplier layouts, and the long tail is where extraction fails. A generator with 40 templates teaches 40 layouts well; models trained on narrow or automatically labeled layout sources generalize poorly to diverse layouts, which is the motivation behind human-annotated sets such as DocLayNet [6]. See document layout analysis datasets for the layout side of this gap.

2. Handwriting. Synthetic handwriting fonts are regular in slant, pressure and baseline; real annotations, signatures, struck-through corrections and checkbox ticks are not. Models trained on handwriting fonts often read neat cursive and fail on a field technician's block capitals. Our page on handwriting recognition data from real business forms covers what to request.

3. Capture degradation. Augmentation libraries add blur, JPEG artifacts and noise independently, but real degradation is compound: a photocopy of a fax of a stamped page, skewed, with a shadow from a phone held over a desk. FUNSD was built around noisy, low-resolution scanned forms, which behave differently from clean renders; it contains 199 annotated forms [5]. See real-world document capture conditions.

4. Field-value realism. Sampled values are independent; real values are correlated. Real invoices carry tax amounts that do not quite reconcile, PO numbers in three formats from the same supplier, dates in mixed locales and line items that wrap across pages. Extraction models learn shortcuts on clean synthetic data that fail on these inconsistencies.

5. Document mix and packet structure. A real mortgage or claims packet interleaves classes, duplicates pages and inserts cover sheets. Classification and page-stream segmentation models need the real class priors and ordering, which a generator rarely captures; a Fraunhofer study trained a classifier on synthetic multi-page real-estate documents and evaluated it on real OCR-extracted documents, the right test design for this question [2].

Decision table: generate, license or mix by task

The right source depends on whether the task's variance is enumerable. Use this table as a starting position and then confirm it with the measurement plan below.

Illustrative example: invented to show structure; it does not describe an available dataset.

TaskSynthetic aloneReal documents needed forSuggested starting mix
OCR line recognition (printed)Strong for pre-trainingScanner noise, rare fonts, domain vocabularyMostly synthetic pre-train, real fine-tune
Fixed-form key-value extractionAdequate if template is stableHandwritten entries, form revisions, poor scansSynthetic plus a few hundred real pages per form
Invoice and receipt extractionWeak beyond seen templatesSupplier layout tail, correlated values, multi-page tablesReal-led, synthetic for rare fields
Table structure recognitionPartial: clean grids onlyBorderless, spanning and page-broken tablesReal-led
Document classification and splittingWeak on class priorsReal mix, near-duplicate classes, packet orderReal-led
Handwriting recognitionWeakWriter variance, overlap with printed textReal-led
Evaluation and acceptance testingNot sufficientEverything you will reportReal only

For tables and splitting specifically, see table structure recognition data and page-stream segmentation data.

How to measure the synthetic-to-real gap before you spend

Measure the gap with a three-arm experiment that holds the test set fixed and real. A synthetic-heavy model can look excellent on synthetic validation data and still fail on production traffic, so the comparison only means something on documents the generator never touched; see using a licensed real-data holdout to validate synthetic data.

Illustrative example: invented to show structure; it does not describe an available dataset.

experiment: invoice_extraction_mix_study
test_set:
  source: real, held out, never used for generation or tuning
  size_pages: 1500
  stratify_by: [supplier_layout_cluster, capture_type, page_count, handwriting_present]
  dedup: perceptual hash + text near-duplicate check vs all training arms
arms:
  A_synthetic_only:   {synthetic_pages: 200000, real_pages: 0}
  B_real_only:        {synthetic_pages: 0, real_pages: 5000}
  C_mixed:            {synthetic_pages: 200000, real_pages: 5000, schedule: pretrain_synth_then_finetune_real}
  D_real_scaling:     {real_pages: [500, 1000, 2500, 5000]}
metrics:
  - field_level_f1 per field (total_amount, invoice_date, vendor_name, line_items)
  - exact_match per document
  - error breakdown by stratum (unseen layout, fax, phone photo, handwritten)
decision_rule: >
  If arm C beats arm B by less than the confidence interval on unseen-layout and
  degraded strata, synthetic is not buying tail coverage; spend on real pages.
  If arm D has not plateaued at 5000, more real data is the cheapest gain.

Two details decide whether the result is trustworthy. Deduplicate the test set against every training arm, because near-duplicates inflate scores and are common in collected corpora [9]; documents from the same supplier template count as near-duplicates for layout purposes. Report per-stratum results, since an aggregate F1 hides the tail that motivated the experiment. For field-level ground truth design, see document extraction evaluation sets and when synthetic evaluation data misleads.

Using real documents to seed better synthetic data

Real documents make synthetic generation better when you use them as a specification, not just as training rows. A licensed sample of real invoices tells you the layout clusters to template, the empirical distribution of each field (formats, lengths, locales, null rates), the co-occurrence of fields and the capture conditions to simulate.

Typical seeding steps include clustering real pages by layout embedding to pick templates, fitting value distributions per field and replaying real degradation profiles rather than random augmentation. This is a distinct use from training on the documents, so confirm that the license covers derived templates, statistics and generator tuning, not only model training. Ask the same question about any documents paired with system-of-record entries you intend to use as label sources.

Provenance and terms for generated documents

Generated documents carry the terms of whatever produced them, so record the generator before the data enters training. If an LLM wrote field values or a model rendered page images, the provider's output terms may restrict downstream use; Google's Gemma Terms of Use (version dated 24 March 2025), for example, define Model Derivatives to include models trained to perform similarly to Gemma, including through distillation or synthetic data generated from Gemma outputs [7]. Licensing metadata is already unreliable for collected datasets, with audits finding license omission above 70% on popular hosting sites [8], and synthetic sets built from them inherit that uncertainty.

A minimal provenance record lists the generator and version, any seed documents and their license, model names and terms for text or image generation, prompts or templates, random seeds and the date generated. See provenance records for synthetic training data for a fuller schema.

When to buy real documents

Buy real documents when the measurement shows the tail strata are not improving with more synthetic data, or when the deliverable is an evaluation set you will report to customers or reviewers. In requests, describe the target distribution precisely: document classes, layout diversity (distinct issuers or templates), capture types, handwriting share, page counts, the label layer you need and the uses the license must allow. The document dataset requirements spec gives a template, and the Document AI data hub maps tasks to label layers.

For the broader trade-off across data types, see licensed vs synthetic vs scraped AI training data, combining licensed and synthetic data and the synthetic data glossary entry. If you need real operational documents from US companies to close the gap, you can describe the documents to SourceX; sourcing is on request, and a request does not guarantee a match.

Sourcing real business documents to close the synthetic gap

SourceX sources operational datasets, including documents and finance and legal workflows, from US companies on request, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced before delivery. Describe the real documents you need.

Sources

  1. arXiv, "Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts" (2024). https://arxiv.org/pdf/2412.03590
  2. Fraunhofer-Publica, "Fraunhofer study on classifying multi-page real-estate documents with synthetic training data". https://publica.fraunhofer.de/handle/publica/516664
  3. AWS Labs, "Synthetically Engineered Evaluation Data (SEED)". https://awslabs.github.io/synthetically_engineered_evaluation_data/
  4. LlamaIndex, "Synthetic data for document training". https://www.llamaindex.ai/glossary/synthetic-data-for-document-training
  5. arXiv (Jaume, Ekenel, Thiran), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  6. arXiv (Pfitzmann et al.), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
  7. Google (copy hosted by Wind River license list), "Gemma Terms of Use (version dated 2025-03-24)" (2025). https://open.windriver.com/info/uni-license-list/licenses/gemma-tou-2025-03-24.html
  8. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  9. arXiv (Lee et al.); ACL 2022, "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data