Skip to content

Fine-tuning and post-training data

Report-generation fine-tuning data: source inputs paired with expert reports

Quick answer

A report-generation dataset for fine-tuning is a set of pairs: the full input bundle an expert actually worked from (measurements, findings, photos, logs, tickets, prior reports) and the final report that expert signed. Reports alone teach a model house style; only the paired inputs teach it which statements are supported. Buy pairs with stable record IDs, versioned drafts where available, boilerplate flagged, identifiers masked consistently across inputs and outputs, and a license that names fine-tuning as an allowed use.

By SourceX Editorial · Updated

Why report-only corpora fail for grounded generation

Report-only corpora train fluent narrative but give the model nothing to ground claims in, so it learns to invent plausible findings. Our working hypothesis, consistent with how data-to-text systems are built, is that a model fine-tuned on reports without inputs reproduces section headings, hedging phrases and typical numbers, then fills them from priors rather than evidence. Published report-generation work leans on institutional archives and structured inputs: one radiology project built its dataset from 5,192 consecutive low-dose CT screening reports from a single institution [1], and analytics-report systems such as Satyrn generate text from computed results rather than free recall [3].

Expert reference reports are also scarce in public form. A 2026 deep-research report-generation paper evaluated against just 40 professional reports from Our World in Data as its gold standard [2]. That scarcity is why operational archives, where reports are produced daily against real inputs, are the realistic source for supervised fine-tuning (SFT) at useful scale. For the general SFT sourcing workflow, see how to source supervised fine-tuning data.

What a complete input bundle contains

A complete bundle is everything the author had open when writing, captured as of the report's sign-off time, not as of export day. Inputs that changed after sign-off (a corrected measurement, a closed ticket) create pairs where the report contradicts its own evidence, which is a silent label-noise source.

Typical bundle components by report family:

  • Field inspection reports (property, equipment, construction punch lists): checklist responses, readings with units, photo references with captions, prior inspection report, asset register row. See SourceX's page on licensing inspection reports.
  • Project status reports: task tracker exports (Jira, Asana), milestone tables, budget-to-actual lines, risk register, prior status report. See project plans and status reports.
  • Analyst and finance reports: ledger extracts, KPI tables, variance calculations, management commentary, prior period report.
  • Support and incident reports: ticket thread, logs or monitoring alerts, timeline, root-cause notes, postmortem template.
  • Clinical and medical-writing reports: source tables and study data; these carry HIPAA obligations and are covered on clinical study reports and source tables.

Ask for the prior report in each bundle. Many real reports are deltas ("no change since 14 March inspection"), and a model cannot learn that behavior without seeing the previous document.

Example pair record for report-generation SFT

A usable record keeps inputs, target and provenance in separate fields so you can rebuild prompts, mask loss and trace any sentence back to evidence.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "pair_id": "insp-000412",
  "report_family": "hvac_preventive_maintenance",
  "signed_at": "2025-11-03T16:20:00Z",
  "author_role": "licensed_technician",
  "inputs": {
    "checklist": [{"item": "condenser_coil", "status": "dirty", "note": "fin damage NE corner"}],
    "readings": [{"name": "supply_air_temp", "value": 58.2, "unit": "F", "ts": "2025-11-03T15:02:00Z"}],
    "photos": [{"ref": "IMG_0193", "caption_by_author": "coil fin damage"}],
    "prior_report_ref": "insp-000377",
    "site_ref": "[SITE_7]", "client_ref": "[CLIENT_2]"
  },
  "target_report": {
    "format": "markdown",
    "sections": ["Summary", "Findings", "Readings", "Recommendations"],
    "text": "..."
  },
  "draft_versions": [{"version": 1, "author": "technician", "text": "..."}, {"version": 2, "author": "reviewer", "text": "..."}],
  "boilerplate_spans": [[0, 212]],
  "statement_links": [{"sentence_idx": 3, "evidence": ["readings[0]"]}],
  "masking": {"method": "consistent_pseudonym", "applied_to": ["inputs", "target_report"]}
}

Three fields matter most for buyers. statement_links (even on a sample) lets you measure faithfulness. draft_versions turns reviewer edits into preference pairs. boilerplate_spans lets you down-weight template text without rewriting reports. If you also need strict schema output, compare with structured-output fine-tuning pairs.

Template boilerplate and deduplication

Template-heavy archives need boilerplate deduplication before training, or the model will overfit disclaimers and section scaffolding. Lee et al. found long repeated substrings across common corpora, with one sentence repeated over 60,000 times in C4, and reported that over 1% of unprompted output from models trained on such data was copied verbatim [4]. Report archives show the same pattern at small scale: legal disclaimers, method statements and "scope of inspection" paragraphs repeat in nearly every document.

Practical handling:

  • Detect repeated spans with exact substring matching (suffix arrays) or MinHash on sentence shingles, then mark them rather than delete them.
  • Exclude marked spans from the loss, or keep one canonical copy per template version, so the model still learns to emit required language.
  • Cluster near-duplicate reports (same site, same week, copy-paste narratives) and split train and evaluation by site or client, not by row, to avoid leakage.
  • Track template versions; a mid-archive template change otherwise looks like author disagreement.

Faithfulness checks: every statement traceable to an input

Faithfulness means each factual sentence in the target report can be traced to a field in the input bundle; unsupported sentences should be flagged before training, not after deployment. This is our recommended acceptance test, and it should run on a sample before you sign. Word-level hallucination corpora such as RAGTruth show the annotation pattern: label spans of generated text that conflict with or are absent from the provided context [5]. Apply the same lens to human-written targets, because experts legitimately add knowledge not in the bundle (site history, phone calls, tacit judgment).

Sort sentences in a 100 to 200-pair audit sample into three buckets:

BucketExampleWhat to do
Supported by input"Supply air measured 58.2 F" with a matching readingKeep
Expert inference"Fin damage likely from hail" with no input mentionKeep, but tag; decide whether the model should hedge
Unsupported or contradicted"Filters replaced" with no work-order lineDrop the pair or repair the bundle; ask whether inputs are missing

A high unsupported rate usually means the bundle is incomplete (missing attachments, emails, verbal notes), not that experts invent facts. Fix the extraction, not the reports. For data that teaches abstention and grounded answers, see fine-tuning data that reduces hallucinations.

Drafts, reviewer edits and preference data

Draft-to-final histories are the most valuable byproduct of report archives because they give you preference signal for free. InstructGPT-style pipelines use SFT on human demonstrations, then learn from human rankings of outputs [6]; Direct Preference Optimization fits a policy straight from preference pairs without a separate reward model [7]. A technician draft plus a reviewer-approved final is a natural (rejected, chosen) pair, provided the reviewer's edits are substantive rather than formatting.

Ask suppliers whether their document system (SharePoint version history, Google Docs revisions, a report-builder's approval workflow) retains drafts, and whether reviewer comments survive export. Filter pairs where the diff is only whitespace or a date stamp.

Masking client, site and personal identifiers

Identifiers must be masked consistently across inputs and target, or the model learns to copy names that no longer appear in the prompt. Reports name clients, sites, addresses, technicians, patients and account numbers in headers, signatures and free text; inputs repeat them in different formats (asset tags, photo EXIF, file paths).

Ask for consistent pseudonyms ([SITE_7] in both the checklist and the report body), not random per-field replacement, and confirm that photo metadata and embedded document properties were stripped. Health reports that contain protected health information need HIPAA de-identification, via Safe Harbor (removing 18 identifier types of the patient and of relatives, employers and household members) or Expert Determination [8]. For detector choice and scanning targets as well as inputs, see scanning a training corpus for PII before fine-tuning.

Rights, provenance and documentation to request

Reports are commercial work product, so the license must cover both the inputs and the expert-authored text for fine-tuning. Public dataset licensing is unreliable: the Data Provenance Initiative found license omission above 70% and license error rates above 50% on popular hosting sites [9]. For operational data, confirm the supplier owns the reports (not a client who commissioned them), and that client contracts do not restrict reuse of deliverables.

Request a datasheet in the Data Cards style covering upstream systems, collection window, author roles, review process, masking method and known gaps [10]. Our data provenance guide and AI training data licensing guide cover the clauses and records in more depth.

Buyer checklist for report-generation pairs

Use this checklist in your first call with any supplier.

  1. Is every report delivered with its full input bundle as of sign-off, including the prior report?
  2. Are pair IDs stable, and can inputs and targets be joined without manual matching?
  3. Are drafts and reviewer edits available, and which system holds them?
  4. Are boilerplate spans or template versions identified?
  5. What fraction of a sample's sentences trace to inputs, and who audited it?
  6. Are client, site and personal identifiers pseudonymized consistently across inputs and reports?
  7. Does the license cover fine-tuning on both inputs and expert-written text, and for what term?
  8. Can train and evaluation splits be made by site, client or author?

Run the same sample through the steps in evaluating a fine-tuning dataset before you buy it, and see the fine-tuning and post-training data hub for adjacent formats such as summarization and long-context pairs.

How SourceX approaches report-generation pairs

SourceX sources operational datasets from US companies on request, including documents, engineering records, support histories and finance and legal workflows, and manages licensing and ongoing purchases. Data is not held in stock, and a request does not guarantee a match; buyers describe the data they need, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. You can describe the bundle and report structure you need on the SourceX buyer page.

Source report-generation training pairs through SourceX

SourceX finds US businesses that hold the input-and-report data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered. Personal details are removed or replaced before delivery, with the method recorded and a sample checked. Start by describing your report family and bundle at sourcex.si/buyers.

Frequently asked questions

How many pairs do I need for report-generation SFT?

There is no fixed number; it depends on the base model, report length and how many report families you cover. Run a small pilot fine-tune on a few hundred audited pairs per family and measure faithfulness on a held-out site before scaling purchases.

Can I use synthetic inputs with real reports?

Reconstructing inputs from a report reverses the causal direction: the model learns that every input field appears in the report, which is false for real bundles. If you do generate inputs, label them and keep a real-pair evaluation set. See due diligence for purchased synthetic fine-tuning data.

Should photos be included if I only train a text model?

Keep photo references and the author's captions even for text-only models, since reports cite them. If you later add a vision encoder, the same pairs become multimodal training data; see multimodal training data.

Sources

  1. arXiv, "Development and Validation of a Large Language Model for Generating Fully-Structured Radiology Reports" (2024). https://arxiv.org/pdf/2409.18319
  2. arXiv, "CogGen: A Cognitively Inspired Recursive Framework for Deep Research Report Generation" (2026). https://arxiv.org/pdf/2604.17072
  3. arXiv, "Satyrn: A Platform for Analytics Augmented Generation" (2024). https://arxiv.org/pdf/2406.12069
  4. arXiv (ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  5. arXiv (ACL 2024), "RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models" (2024). https://arxiv.org/pdf/2401.00396
  6. arXiv (OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  7. arXiv (Stanford), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  8. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  9. Longpre et al., Nature Machine Intelligence 6 (2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI". https://www.nature.com/articles/s42256-024-00878-8
  10. arXiv (Google Research, FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data