Skip to content

Fine-tuning and post-training data

Structured-output fine-tuning data: JSON and field-extraction pairs

Quick answer

A structured output fine-tuning dataset pairs an input, usually a business document, with the schema the model must follow and the verified JSON record it should return. Strong targets often come from values a business already keyed into a system of record, such as posted invoice fields in an ERP. Use them only after checking each value against the document, adding examples where fields are absent, and replacing personal data identically in the input and the target.

By SourceX Editorial · Updated

This page covers supervised fine-tuning pairs for schema-following extraction. General document data is covered under training data for document understanding, general SFT sourcing in how to source supervised fine-tuning data, and other post-training data in the fine-tuning data guide.

Valid JSON is the easy part; correct values are what you are buying

A model learns to emit well-formed JSON from a small, consistent set of examples, and decoding constraints can enforce syntax at inference time. Neither tells it which number on a crowded invoice is the total, so what you are paying for is correct targets over realistic, messy documents.

LIMA's authors fine-tuned a 65B-parameter model on 1,000 curated pairs and concluded that limited, high-quality instruction data is enough to teach output format, since almost all knowledge comes from pre-training [1]. Constrained decoding, which blocks tokens a grammar or JSON Schema does not allow, prevents most syntax errors. But a constrained model that reads "Amount due: 1,420.00" beside "Subtotal: 1,240.00" can still return the wrong number inside a valid object.

Production failures are semantic: a remit-to address returned as the bill-to address, a purchase order number invented when none is printed, day and month swapped, a line item counted twice across a page break. Sizing is covered in how much data you need to fine-tune an LLM.

Four sources of target values, compared

Targets come from a system of record, human keying, a stronger model, or documents rendered from structured data; each fails differently, and strong datasets usually combine two or more.

System-of-record valuesDouble-keyed annotationDistilled model outputsSynthetic rendered documents
SourceValues posted in an ERP, claims, loan or AP system, joined to the documentTwo keyers per field; a third adjudicatesA stronger model extracts; validators filterTemplates filled with structured values, rendered to PDF or image
StrengthExists at scale; tied to an outcome (paid, adjudicated, funded)Matches the page; agreement is measurableCheap; any schemaExact document-target alignment
Typical failureNormalized, derived, corrected or mis-keyed valuesCost; fatigue on long line-item tablesTeacher errors become targets; teacher terms may bar trainingClean layouts unlike real scans
RequireReconciliation to the document; a re-key auditAgreement statistics; adjudication rulesTeacher, version, date and terms per recordRealistic degradation; a real held-out test set

Pairing mechanics are in documents paired with system-of-record entries and key-value extraction labels; model-made targets in building distillation datasets and due diligence for synthetic fine-tuning data.

What one document-to-JSON training pair contains

A usable pair holds the document as the model will see it, the schema version, the target object, per-field evidence and origin metadata. Without evidence and origin fields, you cannot audit targets or filter weak ones later.

  • Document. The original PDF, TIFF or PNG plus what the model receives: page images, or OCR text with word coordinates. Record the OCR engine and version, because a different engine at inference changes the input.
  • Schema. A versioned JSON Schema or equivalent with types, enumerations, required fields and nullability, rendered into every prompt so the model follows the schema it is given instead of memorizing one.
  • Field evidence. Page and bounding box or text span per value, or an explicit not-present status. FUNSD, a public set of 199 noisy scanned forms, annotates entities and the links between them [2]; evidence plays that role for a target object.
  • Origin metadata. Where each value came from, the record's state when extracted, and the de-identification applied.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "id": "apinv-000412",
  "document": {"file": "docs/apinv-000412.pdf", "sha256": "9f2c...e41a", "pages": 2,
               "ocr": {"file": "ocr/apinv-000412.json", "engine": "ocr-engine-x", "engine_version": "5.1"}},
  "schema_id": "ap_invoice",
  "schema_version": "3.2",
  "target": {
    "vendor_name": "Halden Fluid Supply LLC",
    "invoice_number": "0004417",
    "invoice_date": "2025-03-04",
    "po_number": null,
    "currency": "USD",
    "subtotal": "1180.00",
    "tax": "60.00",
    "total": "1240.00",
    "remit_to_account": "6203158840",
    "line_items": [
      {"description": "Hydraulic hose 3/8 in, 50 ft", "quantity": "4", "unit_price": "245.00", "amount": "980.00"},
      {"description": "Freight", "quantity": "1", "unit_price": "200.00", "amount": "200.00"}
    ]
  },
  "field_evidence": {
    "invoice_number": {"page": 1, "bbox": [412, 88, 498, 104], "source_text": "INV# 0004417"},
    "invoice_date": {"page": 1, "bbox": [412, 108, 490, 124], "source_text": "03/04/25"},
    "po_number": {"status": "not_present"},
    "total": {"page": 2, "bbox": [455, 610, 520, 628], "source_text": "1,240.00"}
  },
  "meta": {
    "target_origin": "erp_posted_value_reconciled",
    "erp_state": "final_after_corrections",
    "fields_taken_from_document_not_erp": ["vendor_name"],
    "rekey_audit": "sampled_agree",
    "deid": "format_preserving_surrogates_in_document_and_target",
    "template_cluster": "vt-118",
    "split": "train"
  }
}

The example makes decisions a supplier should not make silently: amounts as strings to keep precision, the invoice number's leading zeros kept, the printed "03/04/25" resolved to YYYY-MM-DD with the original kept as evidence, and a null purchase order because none is printed. Line amounts sum to the subtotal and subtotal plus tax equals the total, which a validator can check on every pair. For dataset-level documentation, Croissant describes metadata, files and record-level structure in a JSON-LD vocabulary built on schema.org [3].

When the system of record disagrees with the document

System-of-record values are verified in a business sense, because someone approved, paid or adjudicated them, but they record what the business decided, not always what the document says. Before a value becomes a target, check whether it is printed, normalized, derived, later corrected or mis-keyed.

Even curated labels are imperfect: Northcutt and colleagues estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets [4].

FieldHow the stored value diverges from the pageRule for the spec
Party namesVendor-master name or ID, not the letterhead nameChoose printed or master; treat master matching as a separate task
DatesPosting or received date stored; printed formats ambiguousMap each date to one source column; keep printed text as evidence
AmountsConverted, rounded or discounted; decimal commas on foreign documentsTarget the document currency and amount
IdentifiersLeading zeros stripped; PO number looked up, not printedStore as strings; null when not printed
Derived fieldsGL account, cost center or tax code chosen by the clerkExclude, or label as classification outputs
Line itemsCollapsed, or split for cost allocationUse only where posted lines map one-to-one to printed lines
CorrectionsCredit memos, reversals and re-postingsTake values after a settle period; record changed fields

Derived fields matter most for hallucination: a target value the document never shows teaches the model to invent plausible ones. In ERPs such as SAP or NetSuite, a posted vendor bill carries both values copied from the document and coding the clerk added, so separate them field by field.

The practical check is a re-key audit: independent keyers extract a random sample without seeing system values, and per-field disagreement estimates target noise. Fixing mislabeled pairs at scale is covered in finding label errors with confident learning.

Coverage that makes schema adherence generalize

Schema adherence carries over to new documents when training shows absent fields, unfamiliar layouts and more than one schema; a set where every field is always filled teaches the model to fill every field.

  • Absent and unreadable fields. Define null (not present) and a separate illegible status in the schema, and set a minimum share of null-bearing documents per optional field; without them, models fill fields from nearby text.
  • Template spread and splits. Cap documents per sender template. Same-template documents are near-duplicates; for language-model training corpora, Lee and colleagues showed that deduplication made models emit memorized text about ten times less often [5]. Split train, validation and test by template and sender, not by document.
  • Capture conditions. Native PDFs, scans, phone photos, faxes, stamps and handwriting in production proportions (scanned forms and handwritten documents).
  • Schema and format variants. Renamed fields, reordered properties, added optional fields, and YAML, XML or tool-call output where needed (function-calling fine-tuning data). Mix in general instruction data to limit catastrophic forgetting.

Field-level acceptance metrics

Accept a structured-output dataset on field-level measurements from a random sample, not on a claim that targets come from a system of record. Report each metric per field and document type, because a strong average hides a weak field.

MetricHow it is computedWhat it catches
Schema validityShare of targets valid against the stated schema versionMalformed or out-of-version targets
Evidence coverageShare of non-null values with page and location evidenceDerived or looked-up values passed off as extraction
Re-key agreementMatch between targets and blind re-keying, after normalizationKeying errors, stale values, wrong column mapping
Null accuracyAgreement on not-present fieldsSystems that default missing values to zero or a placeholder
Arithmetic consistencyLine amounts, tax and total reconcile within a toleranceSplit or duplicated line items; wrong totals
Template concentrationLargest share of documents from one templateLayout memorization

For two-keyer agreement on enumerated fields such as currency or document type, a chance-corrected coefficient such as Krippendorff's alpha, where 1 is perfect reliability and 0 is chance-level agreement, says more than percent agreement [6]. ISO/IEC 5259-2 offers a data quality model with measurable characteristics and reporting guidance to draw on for the acceptance clause [7]. Run a sample ablation before full volume (evaluating a fine-tuning dataset before you buy it) and keep test documents out of training (document extraction evaluation sets).

If the model will be part of a high-risk AI system under the EU AI Act, Article 10(3) requires training, validation and testing data to be relevant, sufficiently representative and, to the best extent possible, free of errors and complete in view of the intended purpose [8]. As of October 2026 the Act has been amended by Regulation (EU) 2026/1744, which secondary sources report moved Annex III high-risk obligations to 2 December 2027 [9].

Personal and confidential fields: surrogates that keep their format

Extraction targets are dense with personal and confidential values: names, addresses, bank and card numbers, tax IDs, policy and member numbers. Replace them with format-preserving surrogates applied identically in the document text or image and in the target, or the pair teaches the model to output values absent from the page.

Masking an account number as "XXXX" in the target while the scan still shows the digits teaches the model to emit a mask instead of reading the page; redacting only the image leaves a target value the model cannot see. A surrogate of the same shape, such as a nine-digit routing number with a valid check digit, keeps the task intact (masking vs surrogate replacement).

SFT trains a model to reproduce target strings, which raises the stakes: Carlini and colleagues extracted hundreds of verbatim training sequences, including contact details, from GPT-2 [10]. Detectors miss things; the open-source Presidio project states there is no guarantee it finds all sensitive information [11]. Require the detection method, a manual sample check, and a scan of page images as well as text (PII scans before fine-tuning).

Sector rules follow the documents:

  • Health claim forms and remittances held by a HIPAA covered entity are protected health information; the de-identification standard is met by Safe Harbor, which removes 18 listed identifiers and requires no actual knowledge that the rest could identify the person, or by Expert Determination [12].
  • California consumer data is deidentified under the CCPA only if the holder, among other conditions, contractually binds recipients to the deidentification requirements, so expect those terms in the license [13] (CCPA deidentified data obligations).
  • Loan files and bank statements: under Regulation P, a recipient of nonpublic personal information from a nonaffiliated financial institution faces reuse and redisclosure limits, whether or not it is a financial institution itself [14].

For datasets sourced through SourceX, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded per dataset, and a sample is checked after processing. No method is perfect, so run your own scan. Health records require HIPAA de-identification by Safe Harbor or Expert Determination before SourceX considers them for a license.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Rights questions for document and record pairs

A document-to-JSON dataset combines documents and a system-of-record extract, so the license must cover both, including third-party content on the documents. Ask:

  • May the holder license documents received from counterparties, and are any under confidentiality terms?
  • Does the grant cover extracts, field evidence and targets, not only the files?
  • Does permitted use include fine-tuning and use of the resulting weights (fine-tuning-only data licenses)?

SourceX sources operational datasets from US companies, including documents and finance workflows. Each dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and is delivered under a license defining the included records, allowed uses, term and delivery. Datasets are sourced on request, not held in stock, so a request does not guarantee a match; you can describe the document types and schema you need. Related pages: document understanding models, invoices and receipts and purchase orders.

Request checklist for a structured-output dataset

Fix these items in the request so proposals are comparable:

  1. Document types, capture conditions, date range, volume per type and a cap per template.
  2. Versioned schema files with null semantics and normalization rules.
  3. Per field: the system-of-record source column, "human keyed", or "excluded as derived".
  4. A settle rule for corrections, reversals and credit memos.
  5. Per-field evidence and a target-origin field on every record.
  6. A minimum share of documents with absent optional fields.
  7. Re-key audit sample size and per-field agreement.
  8. De-identification applied identically to documents and targets.
  9. Splits grouped by template and sender.
  10. UTF-8 JSON Lines, one pair per line (.jsonl) [15], plus documents referenced by path and SHA-256 hash, schema files and a datasheet.
  11. License scope covering documents, extracts, targets and model use.

For request structure in general, see how to write a data request for suppliers; for models that enter data into a system instead of returning JSON, see document-to-system entry pairs for back-office agents.

Need document-to-field pairs from real business systems?

Describe the document types, the schema you need filled, the system-of-record fields that should serve as targets, and the license scope. SourceX looks for US businesses that hold matching documents and records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Specify your extraction dataset.

Sources

  1. Zhou et al., Meta AI and collaborators (arXiv:2305.11206; NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  2. Jaume, Ekenel, Thiran (arXiv:1905.13538), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  3. Akhtar et al., MLCommons Croissant working group (arXiv:2403.19546), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  4. Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  5. Lee et al. (arXiv:2107.06499; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  6. Klaus Krippendorff, University of Pennsylvania Annenberg School for Communication, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  7. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  8. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  9. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2026/1744 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
  10. Carlini et al. (USENIX Security 2021), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  11. Microsoft presidio project (pkg.go.dev), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  12. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  13. California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  14. Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  15. jsonlines.org, "JSON Lines". https://jsonlines.org/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data