Skip to content

Document AI data

Document Extraction Correction Logs: Human Validation Data from Production IDP

Quick answer

Human-in-the-loop document extraction data is the log a production intelligent document processing (IDP) system leaves behind when reviewers fix what the model got wrong: the page image, the model's extracted value and confidence, the reviewer's corrected value, and what action they took. Buyers use these logs to fine-tune on hard cases, calibrate confidence thresholds and study real error modes. The value depends on joins back to page images, honest sampling and clean rights to the vendor's model outputs.

By SourceX Editorial · Updated

What a correction log contains, and why the join to the page image matters

A usable correction log ties every human edit to a specific field, a specific bounding region and a specific page image; without that join it is only a list of strings. Most IDP review queues work the same way: the extractor scores each field, values below a confidence threshold go to a reviewer, and everything else passes straight through [1][3]. What the reviewer does in that queue (accept, edit, reject, reassign a label) is the signal you are buying.

Ask for these layers, in this order of importance:

  • Page image or rendered PDF page, at the resolution the model saw, with a stable page_id and document hash.
  • Model output per field: field key, extracted value, normalized value, bounding box, OCR tokens and the model's confidence score.
  • Human outcome per field: corrected value, reviewer action code, and whether the correction changed the text, the field assignment or the box.
  • Context: model or template version, document type, queue name, timestamps, and the confidence threshold in force at the time.

Without the model version, you cannot tell whether an error pattern belongs to today's model or one retired two releases ago. Without the threshold, you cannot reconstruct which fields a reviewer never saw.

A field-level record schema buyers can request

The record below is the minimum structure that supports training, calibration and error analysis from the same export. Ask suppliers to map their native export (often JSON per document, or a database table of field events) to something close to it before you sign.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "doc_hash": "sha256:9f2c...e41a",
  "page_id": "inv-000812-p1",
  "doc_type": "vendor_invoice",
  "model_version": "extractor-2025.11",
  "threshold_in_force": 0.85,
  "field": "invoice_total",
  "bbox": [1412, 2203, 1688, 2251],
  "extracted_value": "1,284.50",
  "confidence": 0.93,
  "routed_to_review": false,
  "audit_sampled": true,
  "corrected_value": "7,284.50",
  "reviewer_action": "edit_value",
  "error_class": "ocr_char_confusion",
  "reviewer_role": "AP_REVIEWER_L2",
  "reviewed_at": "2026-03-14T15:22:07Z"
}

Note the two flags. routed_to_review: false with audit_sampled: true marks a high-confidence error caught only by a random audit, which is the most valuable row type for calibration work. Note also that reviewer_role replaces any named person.

Correction logs are a biased sample: pair them with random audits

Corrections cluster on the documents the model already found hard, so a log made only of reviewed items over-represents low-confidence fields and under-represents confident errors. Vendor guides frame HITL as automating the clear cases and sending exceptions to people [2], which means passed-through fields are not checked unless someone audits them. If you train or evaluate only on reviewed items, you learn the routing policy as much as the documents.

The remedy is a random audit stream: a fixed percentage of auto-accepted documents pulled for full human review regardless of confidence. Ask whether the supplier ran one, at what rate, and whether audit rows are flagged separately from queue rows. A log with no audit stream can still train a correction-aware model, but it cannot tell you the true field-level error rate.

Also check reviewer quality. Human labels are not ground truth by default; an audit of popular benchmark test sets estimated an average label error rate of at least 3.3% [5]. Where a supplier has double-reviewed a subset, request the overlap so you can compute agreement with a coefficient such as Krippendorff's alpha [8], per field type.

Error classes worth tagging before you train

Field-level errors in production IDP fall into a small set of classes, and a supplier who can tag them (or let you tag a sample) saves most of the error-analysis work. These are the classes that show up across invoices, forms and statements:

Error classWhat the reviewer changedTypical causePrimary use
OCR character confusionDigits or letters (1/7, 0/O, 5/S)Low resolution, fax or skewOCR post-correction models
Wrong field assignmentCorrect text moved to another keySimilar labels, multi-column layoutsKey-value linking and layout models
Missed fieldValue added where model returned nullUnusual template, handwritingRecall training, template coverage
Boundary errorText trimmed or extendedWrapped lines, merged cellsSpan and box regression
Normalization errorDate, currency or unit reformattedLocale formats, ambiguous datesPost-processing rules, structured output
Line-item misalignmentRows split, merged or reorderedMulti-page tablesTable extraction

Degraded capture drives much of the first row; see degraded document images for sourcing scans, faxes and phone photos with known capture conditions. Line-item errors deserve their own sample design, covered in invoice line-item extraction data.

Using corrections for calibration and active learning

The highest-value analysis is finding fields where the model was confident and wrong, because those errors bypass the review queue entirely. Bin extracted fields by confidence, compute the corrected-value rate per bin per field type, and compare that curve to the threshold in force. If invoice_total at 0.90–0.95 confidence is corrected more often than vendor_name at 0.70, a single global threshold is the wrong policy.

For active learning, vendors describe feeding reviewer corrections back as new training examples [1]. Before you copy that loop with purchased data, separate three uses that need different splits:

  1. Fine-tuning on hard cases: reviewed corrections, deduplicated by template so one vendor's invoice layout does not dominate.
  2. Calibration: audit rows only, since they are the only unbiased estimate of error at high confidence.
  3. Evaluation: a held-out set split by document source and date, never by random row, so templates do not leak across splits. For evaluation-grade field truth, see document extraction evaluation ground truth.

Public sets do not substitute here. FUNSD, for example, contains 199 annotated scanned forms [6]; it is useful for form understanding research but carries no model outputs, confidence scores or reviewer actions. Check the license of any public set before commercial use, as covered in public document datasets and commercial use.

Rights checks specific to IDP correction logs

Correction logs contain three layers that can each carry restrictions: the business documents themselves, the reviewers' work product, and the extraction vendor's model outputs. The first two are usually governed by the data holder's own contracts and policies; the third is governed by the IDP or OCR vendor's terms of service, which may restrict using outputs to train a competing model. Ask the supplier to show which extraction product generated the values and to confirm that the relevant terms permit the intended use.

Licensing gaps are common across AI data generally. One audit of more than 1,800 datasets found license information omitted in over 70% of cases on popular hosting sites [7], so do not infer permission from the absence of a restriction. Documents also carry third-party personal data: payee names, account numbers, addresses and signatures on invoices, claims and statements.

Use this pre-signature checklist:

  • Which IDP or OCR product produced the extracted values, and do its terms allow output reuse for model training?
  • Who owns the reviewers' corrections (employees, a BPO, or an annotation vendor offering HITL review for Document AI [4]), and do their contracts assign the work product?
  • Are reviewer names and IDs replaced with role codes before delivery?
  • How are personal details in page images handled: redaction boxes, synthetic replacement or exclusion, and is the method recorded?
  • Are health or substance-use records present that would trigger HIPAA or 42 CFR Part 2 rules?
  • Is the export documented in a machine-readable card, for example using Croissant-RAI fields for labeling and life cycle [9]?

Where correction logs fit among other document labels

Correction logs are one of three ways to get field truth from real operations, and each has a different bias. System-of-record pairing uses the value that landed in the ERP or claims system as the label, which covers every document but can be wrong when the posted value was itself adjusted; see documents paired with system-of-record entries. Purpose-built annotation is cleanest but most expensive per page; see key-value extraction labels.

Correction logs sit between them: cheap because the work was already done, rich in error signal, but skewed toward hard cases. Broader QA scores and draft-to-final edits across text tasks are covered on human feedback QA scores and corrections, and overrides of automated business decisions (not field extraction) are covered in human override and correction logs. If you need fresh expert labeling instead, see expert annotations and labels. The full set of document data guides is on the Document AI data hub.

How SourceX approaches correction-log requests

SourceX sources operational datasets, including documents and finance and legal workflows, from US companies on request; it does not hold inventory, and a request does not guarantee a match. Buyers describe the data they need (document types, fields, whether audit rows and model versions are required) and SourceX looks for US businesses that hold it, with every release approved by the supplying company. You can start a buyer request once you know which record layers you need.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows after an executed agreement and supplier approval.

Request human-validated document extraction data

SourceX sources operational document datasets from US companies on request and manages licensing, from finding a willing supplier through agreement and ongoing purchases. Nothing is contracted until a supplier agrees, and pricing and allowed uses are set per deal in the license. Describe the extraction correction data you need.

Frequently asked questions

Is OCR correction data the same as extraction correction data?

No. OCR correction fixes characters within a token, while extraction correction also covers field assignment, missing fields, spans and normalization. A good log records which kind of change the reviewer made so you can train separate OCR post-correction and key-value models.

Can I use correction logs as an evaluation set?

Only the random audit portion gives an unbiased error estimate. Queue-reviewed rows over-represent hard cases, so an evaluation built from them will understate production accuracy and can distort comparisons between model versions.

How many corrections are enough?

There is no fixed number. Count corrections per field type and per template family rather than in total, because a large log dominated by one vendor's invoice layout teaches less than a smaller log spread across many layouts.

Sources

  1. Mindee, "The role of human-in-the-loop (HITL) in document automation". https://www.mindee.com/blog/what-is-human-in-the-loop-automation
  2. Airparser, "Human-in-the-loop document extraction". https://airparser.com/blog/human-in-the-loop-document-extraction/
  3. Unstract, "AI + Human-in-the-loop = 99% Accurate Data Extraction (Here's How)". https://unstract.com/?p=15965
  4. iMerit, "Human-in-the-loop for Document AI". https://imerit.ai/domains/financial-services/human-in-the-loop-for-document-ai/
  5. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. Jaume, Ekenel, Thiran, "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  7. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  8. Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  9. Jain et al. (MLCommons), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data