Skip to content

Fine-tuning and post-training data

Long-context fine-tuning data: sourcing long documents and tasks

Quick answer

Long-context fine-tuning data is a set of naturally long, internally coherent inputs (contracts with amendments, multi-year case files, engineering design histories, long project threads) paired with tasks that cannot be answered without reading across the whole input. Buy for document structure, not just token count. Concatenated short texts mainly teach a model to tolerate long inputs; real document families with cross-references are more likely to teach it to use them.

By SourceX Editorial · Updated

Why naturally long documents beat concatenated short ones

Naturally long documents carry dependencies that span tens of thousands of tokens, and those dependencies are what long-context training is supposed to teach. A 90-page master services agreement defines terms in section 1 that change the meaning of an indemnity clause in section 14 and are overridden by Amendment 3. A packed sequence of fifty unrelated support tickets has the same length but no reason for the model to look back.

Published work points the same way. The LongAlign authors built long instruction data by writing tasks over diverse long-context sources, and paired it with packing and sorted batching so long SFT runs stay efficient [1]. Gao et al. reported that the source of long data matters for continued training, naming code repositories and books as strong long sources, and that long data should be mixed with high-quality short data [4]. The working hypothesis for buyers: concatenation is a cheap filler, and real document families are the scarce asset.

Evaluation research explains why this matters. Models often lose information placed in the middle of long inputs [2], and simple needle-in-a-haystack retrieval overstates effective context length compared with tracing and aggregation tasks [3]. A model trained only on synthetic needles can pass the needle test and still fail to reconcile two versions of a contract.

Which business records are long, coherent and licensable

The best candidates are operational records that already exist as linked families inside a company's systems. Each is long because the work was long, not because someone padded it.

  • Contracts with their amendment chains. Base agreement, statements of work, change orders and amendments, ideally with redline history. See contract redline datasets for legal AI training.
  • Case and claim files. Intake, correspondence, adjuster or examiner notes, supporting documents and the final determination, in date order.
  • Engineering histories. Design docs, RFCs, linked tickets, code review threads and postmortems for one system across months.
  • Multi-week project and support threads. Escalated support cases with dozens of turns and attachments, or a deal cycle from first call notes to signed order form.
  • Manuals and technical reports. Equipment manuals, SOP binders and audit reports with internal section references and appendices.
  • Finance and legal workflows. Close packages, audit workpapers, diligence data rooms and matter files.

Broader enterprise collections are covered in enterprise document datasets for AI training. If what you need is multi-step action trajectories rather than long reading inputs, that is a different purchase: see training data for long-horizon task agents.

Task types that force the model to read the whole input

A long input only trains long-context use if the target output depends on several distant spans. Specify task families in the order, and ask the supplier or annotation team to record which spans each answer depends on.

Task familyExample prompt over a document familyEvidence the answer needsFailure mode it targets
Cross-reference resolution"Under the current terms, what cap applies to data-breach indemnity?"Definitions section, indemnity clause, later amendmentAnswering from the first matching clause
Version and change tracking"List every change to payment terms between v1 and Amendment 4."All versions in sequenceMissing middle versions [2]
Whole-file summarization"Summarize the claim history and the reasons for the final decision."Spread across the full fileSummaries that only reflect the opening and closing pages
Contradiction and consistency"Where do the field notes conflict with the final report?"Two or more separated passagesHallucinated agreement
Aggregation"Total the approved change-order amounts across all SOWs."Many small factsCounting errors that grow with length [3]
Timeline reconstruction"Order the escalation events and name the owner at each step."Dated entries across the threadWrong ordering when dates are far apart

Single-span lookup questions are still useful as a minority of the mix, but a dataset dominated by them mostly re-teaches retrieval. Pair this work with summarization fine-tuning data from business records when whole-file summaries are a target, and with retrieval-augmented fine-tuning data when your production system will chunk and retrieve instead of reading the full window.

How to write the length and structure specification

A usable order states length in your tokenizer, not pages or characters, and defines how documents link into families. Page counts vary wildly once tables, exhibits and scanned images are converted to text.

Put these fields in the request:

  • Token-length bands. Minimum and target bands measured with your tokenizer (for example 16k-32k, 32k-64k, 64k-128k), and the share of examples you want in each band. LongAlign's LongBench-Chat evaluation used queries of roughly 10k to 100k tokens, which is a reasonable reference range for instruction-style long data [1].
  • Document-family key. A stable family_id linking every member of a contract chain, case file or project, plus member_seq and effective_date so order is recoverable.
  • Internal reference density. A rough count of cross-references (section citations, "as defined in", ticket links) per family, or a sample that lets you estimate it.
  • Text fidelity. Whether exhibits, tables and scanned pages were converted, by what OCR or parser, and whether layout was preserved as Markdown or plain text.
  • Short-data companion set. Whether short instruction data from the same domain is included, since mixing long and short data is part of published recipes [4] and helps limit forgetting (see data mixtures for fine-tuning).

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "family_id": "fam_7c21",
  "family_type": "contract_chain",
  "members": [
    {"member_seq": 1, "doc_type": "MSA", "effective_date": "2021-03-01", "tokens": 31840},
    {"member_seq": 2, "doc_type": "SOW", "effective_date": "2021-04-15", "tokens": 9120},
    {"member_seq": 3, "doc_type": "amendment", "effective_date": "2023-01-10", "tokens": 4210}
  ],
  "total_tokens": 45170,
  "tokenizer": "buyer-specified",
  "task": {
    "type": "version_tracking",
    "prompt": "List every change to the limitation-of-liability clause across all members.",
    "answer": "...",
    "evidence_spans": [
      {"member_seq": 1, "char_start": 81233, "char_end": 82410},
      {"member_seq": 3, "char_start": 2210, "char_end": 3105}
    ]
  },
  "deidentification": {"method": "entity replacement, consistent per family", "sample_checked": true}
}

Recording evidence_spans lets you measure where answers sit in the window and build a position-balanced mix, which directly addresses the middle-position weakness [2].

Real, synthetic and hybrid long-context data

Synthetic long-context data is fast to generate and fine for teaching retrieval mechanics; real document families supply the structure synthetic pipelines struggle to fake. Vendor guidance describes generating problem-specific synthetic datasets for context extension and refining needle-style evaluation [5], and LongAlign itself used Self-Instruct to write tasks over real long sources [1].

A practical hybrid: source real long documents, generate candidate questions with a strong model, then have domain reviewers verify answers and evidence spans. Ask any supplier to label which fields are model-generated, which model produced them, and whether a human checked them. The checks in due diligence for purchased synthetic fine-tuning data apply to the generated task layer even when the documents are real.

Privacy and confidentiality risk grows with document length

Longer records concentrate more identifiers and more context in one example, so re-identification risk and confidential-information exposure rise with length (a hypothesis worth testing on samples). A one-page ticket might name a customer. A two-year case file names the customer, family members, employer, addresses, account numbers and a detailed timeline that can identify someone even after names are removed.

Three controls matter more for long data than short:

  • Consistent replacement across the family. If "Acme Corp" becomes "Company A" in the MSA but "Org 17" in Amendment 3, cross-reference tasks break and the training signal is lost. Ask for per-family consistent pseudonyms.
  • Quasi-identifiers and third parties. HIPAA Safe Harbor lists 18 identifiers, including those of relatives, employers and household members, and Expert Determination is the alternative method [6]. Long health or claims files touch these categories far more often than short records.
  • Third-party confidential content. Contract chains and data rooms contain counterparties' pricing and terms. Screen for these as covered in third-party confidential information screening, and run a detector pass on both inputs and targets as in scanning a training corpus for PII before fine-tuning.

Rights and provenance checks before you buy

Long documents usually have more than one rights holder, so provenance review takes longer per example than for short text. A contract chain involves both parties; a case file may include medical records, expert reports and third-party correspondence.

Ask for per-family records of the originating system, the entity that holds the records, consents or contractual basis, and the allowed uses. Do not assume that a public dataset's label is accurate: an audit of more than 1,800 text datasets found license information was frequently missing or wrong on hosting sites [7]. The broader framework is in the data provenance guide and the AI training data licensing guide.

How SourceX sources long business documents

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the commercial process through licensing and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match. You describe the data, not the businesses; SourceX looks for US companies that hold it, and every release is approved by the supplying company.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Health records require HIPAA de-identification by Safe Harbor or Expert Determination. To scope a long-context request, start at SourceX for AI data buyers, and see the fine-tuning and post-training data hub for related guides.

Request long-context fine-tuning data

SourceX sources operational datasets, including documents, engineering records and legal workflows, from US companies on request and serves AI teams wherever they are based. The process runs Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Describe the document families, length bands and tasks you need at sourcex.si/buyers.

Frequently asked questions

How long should each long-context fine-tuning example be?

Match the bands to the context length you want the model to use well, measured in your tokenizer, and keep some examples near the top of the target window. Evaluate effective length with tasks beyond simple retrieval, because needle tests overstate it [3].

Can I extend context with short instruction data only?

Gao et al. reported that short instruction data can carry much of the SFT stage once continued training on long data is done [4]. Long, task-bearing examples still help when your target behaviors (version tracking, whole-file summaries) only appear at length.

Should evaluation documents come from the same families as training data?

No. Split by familyid, not by example, or an amendment seen in training will leak answers into the test set. See how to evaluate a fine-tuning dataset before you buy it.

Sources

  1. Bai et al. (arXiv), "LongAlign: A Recipe for Long Context Alignment of Large Language Models" (2024). https://arxiv.org/abs/2401.18058v1
  2. Liu et al. (arXiv), "Lost in the Middle: How Language Models Use Long Contexts" (2023). https://arxiv.org/abs/2307.03172v3
  3. Hsieh et al. (arXiv), "RULER: What's the Real Context Size of Your Long-Context Language Models?" (2024). https://arxiv.org/abs/2404.06654v2
  4. Gao et al. (arXiv), "How to Train Long-Context Language Models (Effectively)" (2024). https://arxiv.org/pdf/2410.02660v4
  5. Anyscale, "Fine-tuning LLMs for longer context and better RAG systems". https://anyscale.com/blog/fine-tuning-llms-for-longer-context-and-better-rag-systems
  6. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  7. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data