Skip to content

Document AI data

OCR Ground Truth Data: Transcriptions Aligned to Real Business Scans

Quick answer

OCR ground truth is a set of page images paired with human-verified transcriptions, aligned at a stated granularity (word or line boxes, plus reading order) and delivered in a format your pipeline reads, such as ALTO, PAGE XML, hOCR or COCO-style JSON. For business scans, specify the alignment level, a written transcription guideline, a character error rate (CER) acceptance threshold measured on double-keyed pages, coverage of real capture defects, and a license that covers both the images and the transcriptions.

By SourceX Editorial · Updated

Most public OCR ground truth comes from historical printing and digital-humanities projects, not from invoices, remittances, claim forms or faxed purchase orders. That gap is why teams fine-tuning text recognition or a vision-language model (VLM) for OCR on business documents usually end up commissioning or licensing data. This guide sits in the document AI data hub; for the general concept, see the OCR glossary entry.

What "ground truth" has to mean for business scans

Ground truth for OCR is only useful when the transcription, the geometry and the transcription rules are all specified, because a text string without coordinates can train an end-to-end model but cannot score a detector or diagnose line-level failures. Decide the granularity before you request anything.

GranularityWhat each record holdsTrains or measuresTypical failure if missing
Page textFull-page transcription, no geometryVLM OCR supervised fine-tuning, page-level CERCannot locate which region failed
Block or regionPolygon per text block, text, block typeLayout-aware OCR, reading-order modelsColumns and tables merge in wrong order
LinePolygon or box per line, baseline, line textLine recognizers (CRNN/CTC, transformer line models)Skewed or curved lines clipped at crop time
WordBox per word, word text, optional confidenceText detection, word recognition, key-value linkingDetection recall cannot be scored
CharacterBox per glyphGlyph classifiers, rare-script workRarely worth the cost for Latin business text

For most business-scan programs, request line polygons with baselines plus word boxes, a reading-order index on each line, and a page-level text file derived from those lines. Line-level pairs are the workhorse for recognizer training; word boxes let you score detection separately and support key-value extraction labels later without re-annotating.

Choosing a ground-truth format your pipeline will read

The format decision is mostly about which tools you already run: ALTO for library-style METS packages, PAGE XML for ground-truth and HTR toolchains, hOCR for HTML-based OCR output, and COCO-style JSON for detection training. Ask suppliers to deliver one canonical format and a converter-tested secondary format rather than hand-maintained parallel copies.

  • ALTO XML. Maintained by the Library of Congress, ALTO records the position, size and style of text blocks, lines and words (String elements) and is commonly packaged alongside METS. Coordinates may be stored in units other than pixels, so confirm the MeasurementUnit value before you train on it.
  • PAGE XML. Developed by the PRImA lab at the University of Salford, PAGE is built around ground truth and is used as an import and export format by Transkribus and eScriptorium. It carries region and line polygons, baselines and reading order, which suits line-recognizer training.
  • hOCR. An HTML microformat in which elements with classes such as ocr_page, ocr_line and ocrx_word carry a bbox (left, top, right, bottom in pixels) and properties like x_wconf in the title attribute. Tesseract emits it, which makes it convenient for comparing a baseline engine against ground truth.
  • COCO-style JSON. Sections for images, annotations and categories, with absolute-pixel [x, y, width, height] boxes [4]. DocLayNet ships layout boxes this way [3]; detection frameworks expect it, but it has no native reading order, so add a field for it.

Common conversion failures to test on a sample: ALTO unit conversion errors that shift every box, PAGE polygons flattened to rectangles that clip rotated lines, hOCR bbox on non-rectangular regions, and lost Unicode normalization (NFC versus NFD) when text passes through HTML entities.

Writing the transcription guideline before anyone keys a page

A written transcription guideline is the single biggest driver of usable CER numbers, because two correct transcribers can disagree on hyphens, ligatures, currency symbols and struck-through text. GT4HistOCR, a widely used set of 313,173 line-image and transcription pairs under CC-BY 4.0, warns that its subcorpora followed different transcription guidelines [1], which is exactly the inconsistency to prevent in a commissioned set.

Cover at least these rules: diplomatic versus normalized transcription (keep "lnv0ice" typos as printed or not); handling of hyphenation at line ends; whitespace collapsing; how to mark illegible spans (for example a fixed token such as [?]); whether stamps, handwritten annotations and printed text over a stamp are transcribed and tagged separately; and the Unicode form for digits, currency and dashes. The ASR literature shows the same effect: reference style differences distort error rates and can overstate real errors [5].

CER and WER as acceptance metrics

Acceptance should be a CER threshold on a held-out, double-keyed sample, with WER reported alongside, because CER measures transcription quality at the level the recognizer works while WER reflects what downstream extraction sees. CER is the Levenshtein edit distance (substitutions, deletions, insertions) between hypothesis and reference divided by the number of reference characters; WER is the same computation over tokens.

A practical acceptance protocol:

  1. Draw a stratified sample per batch (by document type and capture condition), not a random slice dominated by clean pages.
  2. Have two transcribers key the sample independently, then an adjudicator resolves disagreements to produce the reference.
  3. Score the delivered transcriptions against the adjudicated reference, both with and without the guideline's normalization, and report both numbers.
  4. Track inter-transcriber CER as your ceiling. DocLayNet measured agreement by annotating a subset two or three times [3]; the same idea tells you whether a model "error" is really guideline ambiguity.
  5. Reject or rework a batch when CER exceeds the agreed threshold on any stratum, not just the overall mean.

Keep the scored test pages frozen and out of every training split; for building that frozen set from real records, see golden evaluation datasets and document extraction evaluation ground truth.

Born-digital PDFs versus scanned pages

A born-digital PDF's text layer is effectively free ground truth for that page image when you render it yourself, but it says nothing about how your model handles the scanned, faxed or photographed copies that dominate real intake. Use rendered PDFs for volume and font variety, then reserve human-verified transcription budget for genuine scans.

Two traps: text layers in PDFs that were already OCR'd (often PDF/A from a scanner) contain the original engine's errors, so check the PDF Producer metadata and reject "searchable scan" layers as ground truth; and rendering resolution should match production (for example 200 or 300 DPI) or the model learns crisp glyphs it will never see. For the image-to-markdown task specifically, see PDF parsing training data, and for where generated pages fall short, synthetic vs real documents.

Coverage checklist for real business capture conditions

Coverage is defined by the defects your production traffic contains, so write the mix into the request as quotas per condition rather than accepting whatever pages are easiest to transcribe. Public sets such as FUNSD (199 noisy scanned forms annotated for text detection and OCR [2]) are useful for benchmarking but far too small to cover business variance.

  • Print types: laser, inkjet, dot-matrix carbon copies, thermal receipts with fade.
  • Capture: flatbed at 200 to 300 DPI, low-DPI batch scans, Group 3/Group 4 fax compression, phone photos with perspective and shadow. See degraded document images.
  • Overlays: rubber stamps over printed text, signatures crossing lines, highlighter, punch holes, staples and skew.
  • Mixed content: handwriting next to print on the same form, covered in handwriting recognition data.
  • Typography: small-font footers, condensed table fonts, all-caps headers, MICR lines, barcodes with human-readable text.
  • Languages and scripts, including accented names in otherwise English documents; see multilingual business document data.

A request template for OCR ground truth

Describe the data, alignment and acceptance terms precisely; that is what lets a supplier judge whether its records fit and what lets you test the first delivery.

Illustrative example: invented to show structure; it does not describe an available dataset.

request: OCR ground truth for US business scans
document_types: [vendor invoices, remittance advices, bills of lading, faxed purchase orders]
page_volume_target: 20000 page images
capture_mix: {flatbed_300dpi: 40%, batch_200dpi: 30%, fax_g4: 15%, phone_photo: 15%}
alignment: {line: polygon + baseline + reading_order, word: bbox + text}
formats: {canonical: PAGE XML, secondary: hOCR, image: lossless TIFF or PNG as captured}
guideline: diplomatic transcription; illegible = "[?]"; stamps tagged separately; NFC Unicode
acceptance: {sample: stratified 2% double-keyed + adjudicated, metric: CER and WER, reject_if: any stratum above agreed CER}
privacy: names, emails, phones, account numbers removed or replaced in image and text; method recorded
license: images and transcriptions; training and evaluation of OCR and VLM models; term and delivery stated

And one illustrative line record, as a JSON Lines row a training loader might consume:

Illustrative example: invented to show structure; it does not describe an available dataset.

{"page_id":"inv-000412","line_id":"l17","reading_order":17,"polygon":[[212,1840],[1630,1836],[1631,1882],[213,1886]],"baseline":[[214,1874],[1629,1870]],"text":"Net 30 - Remit to: [REDACTED_ADDR]","capture":"fax_g4","guideline_version":"1.2","adjudicated":true}

De-identification without breaking alignment

Redaction on OCR ground truth has to change the image and the transcription together, or the pairs no longer agree and the model learns to hallucinate placeholder tokens. Ask whether personal details are masked in pixels, replaced with rendered surrogate text, or both, and whether boxes for redacted spans are kept with a class label.

Masking a name with a black bar leaves the transcription with a token like [REDACTED_NAME] that never appears in production; surrogate rendering in a matching font keeps the training signal but needs a documented method. Either way, exclude redacted spans from CER scoring or score them separately so they do not inflate error rates.

Rights: license the images and the transcriptions

Transcriptions are derived from the scans, so the agreement should grant use of both the page images and the ground-truth files, for the specific uses you plan (OCR training, evaluation, VLM fine-tuning) and the term and delivery you need. Public dataset licenses are often missing or wrong: an audit of 1,800+ text datasets found license omission above 70% and license error rates above 50% on popular hosting sites [6], so verify provenance rather than trusting a repository tag. For the build-or-buy question, compare annotating your own documents vs licensing labeled data.

SourceX sources operational datasets, including documents and finance and legal workflow records, from US companies on request, and every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. If you need business scans for an OCR program, you can describe the pages and alignment you need; see also scanned forms and handwritten documents and evaluation sets built from real business work.

Sourcing OCR ground truth for business documents

SourceX looks for US businesses that hold the scans you describe, assesses the data and licensing permissions, and agrees pricing and allowed uses in a license; nothing is contracted until the supplying company agrees, and a request does not guarantee a match. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Start a buyer request for OCR ground truth data.

Sources

  1. Zenodo, "GT4HistOCR: Ground Truth for training OCR engines on historical documents". https://zenodo.org/records/1344132
  2. Jaume, Ekenel, Thiran (arXiv:1905.13538), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  3. IBM Research (arXiv:2206.01062), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
  4. CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
  5. arXiv:2412.07937, "Style-agnostic evaluation of ASR using multiple reference transcripts" (2024). https://arxiv.org/pdf/2412.07937
  6. Longpre et al. (arXiv:2310.16787), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data