Skip to content

Evaluation and benchmarking datasets

Long-context evaluation beyond needle-in-a-haystack: real document sets

Quick answer

A useful long-context evaluation dataset tests whether a model can find, connect and reason over evidence buried in real, messy document collections, not whether it can spot an out-of-place sentence in filler text. Needle-in-a-haystack (NIAH) scores can saturate near 100% for recent models, while models that look strong on short inputs degrade at length on questions that share few words with the evidence [1][2]. Build sets from authentic long documents, with near-duplicate distractors, multiple gold passages, controlled evidence positions and length buckets, and score evidence access separately from the final answer.

By SourceX Editorial · Updated

Why needle-in-a-haystack scores saturate

NIAH saturates because it measures literal retrieval of a single inserted fact that is semantically foreign to its surroundings, a task modern attention handles easily. The authors of the FILM-7B paper argued that near-perfect NIAH results can overstate how well a model uses its full context, and proposed probing with varied context styles and retrieval patterns instead [1]. Synthetic benchmarks such as RULER [6] and application-centric suites such as HELMET [7] were built for the same reason: once tasks require multi-hop tracing, aggregation or full-context reasoning, scores fall in ways a NIAH heatmap never shows.

The failure is structural. Filler text such as essays or repeated sentences gives the model no competing candidates, no version drift and no reason to read anything but the anomaly. For a research team, the takeaway is blunt: a NIAH heatmap is a smoke test, not an evaluation.

Literal-match shortcuts and how to remove them

The most common hidden flaw in long-context tests is lexical overlap between the question and the evidence, which lets the model solve the task with string matching. NoLiMa rewrote needles and questions so they share almost no words, forcing a latent association (for example, a question about a city answered by a sentence naming a landmark) instead of literal matching [2]. Under that design, models that look strong at short lengths lose much of their accuracy as context grows [2].

Real documents make this easier to engineer than synthetic filler, because business text already expresses the same fact in many surface forms. A contract amendment refers to "the Effective Date" while the question asks when obligations began; a support ticket says "the unit bricked after the firmware push" while the question asks about a failed update. When you author questions, apply three rules:

  • Paraphrase every question against its gold span and reject any whose content words overlap the span above a fixed threshold (BM25 or token-Jaccard).
  • Require at least one inference step: entity aliasing, unit conversion, date arithmetic or a cross-reference between documents.
  • Run a lexical-only baseline (BM25 top-1 span) over the full context; if it answers a question correctly, rewrite or drop it.

Distractors, multiple gold documents and evidence position

Meaningful long-context items contain several gold passages and realistic distractors that look relevant but are wrong, and they vary where the evidence sits. The "lost in the middle" effect [8], in which accuracy depends on whether evidence sits near the start, middle or end of the input, is widely reported, so depth must be a controlled variable rather than an accident of document order. Recent work on realistic setups pairs multiple gold documents with near-duplicate distractors and reports evidence access separately from answer correctness [3].

Real collections supply distractors for free if you source them deliberately. Superseded contract drafts, prior quarters of the same financial report, duplicate tickets with different resolutions and earlier revisions of a runbook are near-duplicates that differ in exactly the detail the question targets. Synthetic distractors rarely reproduce that kind of version drift, which is what production retrieval and long-context agents actually face. See synthetic vs real documents for broader failure modes of generated text.

Length is its own variable. Evaluate every item at several context lengths with the gold evidence held constant, so you can separate "could not find it" from "found it but reasoned worse with more input"; reporting evidence access separately makes that split visible [3].

What real long documents to source

The right corpus contains naturally long, internally cross-referenced documents whose answers depend on details spread across pages. Useful document families include:

Document familyTypical length driverNatural distractorsLong-context skill tested
Commercial contracts with amendments and SOWsDefinitions, schedules, exhibitsSuperseded drafts, sibling contractsCross-reference, override resolution
Engineering design docs, RFCs and incident postmortemsAppendices, timelines, logsEarlier revisions, related incidentsMulti-hop causal tracing
Support case histories (full threads)Long back-and-forth, attachmentsDuplicate tickets, different fixesTemporal reasoning, final-state tracking
Policy manuals and SOPsNested sections, exceptionsPrior policy versionsException handling, aggregation
Finance and audit workpapersTables, footnotes, reconciliationsPrior-period workpapersNumeric aggregation across sections
Litigation and review setsEmail chains, productionsNear-duplicate threadsMulti-document synthesis

Prefer native formats (DOCX with tracked changes, PDF with real layout, Markdown or HTML exports from wikis, JSON exports of ticket threads) and keep a text-normalized copy with stable character offsets. Offsets are what let you place, move and verify evidence later. For layout-heavy material, coordinate with multi-page long-document data and document extraction evaluation ground truth, which cover field-level parsing rather than reasoning over length.

SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, on request rather than from stock; a request does not guarantee a match. Buyers can describe the long-document collections they need through the SourceX buyer intake.

Designing items: length buckets, positions and scoring

A defensible long-context set is a grid: each question appears at fixed length buckets and evidence depths, with gold spans recorded by offset so placement is reproducible. Making length an explicit experimental variable, rather than an accident of the source document, is what turns a pile of long documents into a benchmark you can compare across model versions.

Practical design choices:

  • Length buckets: for example 8K, 32K, 64K, 128K tokens, measured with the target model's tokenizer and recorded per item.
  • Evidence depth: place gold spans at controlled relative positions (0.1, 0.3, 0.5, 0.7, 0.9) by reordering whole documents, never by splicing sentences into foreign text.
  • Padding policy: extend context with other real documents from the same organization and period, so padding is topically plausible.
  • Scoring layers: score evidence citation (did the model cite the gold span IDs), answer correctness (exact match, normalized numeric match or rubric), and abstention when the answer is absent.
  • Negative items: include questions whose answer exists only in a superseded distractor, so the correct output is the current value or "not stated."

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "item_id": "lc-contracts-0412",
  "question": "Which party bears freight costs for shipments after the second amendment?",
  "question_type": "override_resolution",
  "lexical_overlap_jaccard": 0.06,
  "context_length_tokens": 65536,
  "tokenizer": "target-model-tokenizer-v1",
  "documents": [
    {"doc_id": "MSA-2022", "role": "gold", "char_start": 0, "char_end": 48210},
    {"doc_id": "AMD-2-2024", "role": "gold", "char_start": 48211, "char_end": 61980},
    {"doc_id": "AMD-2-DRAFT", "role": "near_duplicate_distractor", "char_start": 61981, "char_end": 75102},
    {"doc_id": "MSA-SIBLING-VENDOR", "role": "topical_padding", "char_start": 75103, "char_end": 241790}
  ],
  "gold_spans": [
    {"doc_id": "AMD-2-2024", "char_start": 52004, "char_end": 52311, "relative_depth": 0.24},
    {"doc_id": "MSA-2022", "char_start": 30122, "char_end": 30408, "relative_depth": 0.14}
  ],
  "answer": "Buyer",
  "answer_type": "entity",
  "distractor_answer": "Supplier",
  "requires_hops": 2,
  "pii_transform": "names and account numbers replaced with consistent pseudonyms",
  "reviewer_signoff": ["annotator_a", "adjudicator_b"]
}

Fields like distractor_answer and relative_depth let you report how often the model takes the superseded value and how accuracy changes with depth, which is far more diagnostic than a single score.

Quality control for long-context gold labels

Gold labels in long-context sets fail more often than in short QA because annotators cannot hold 100K tokens in mind. Even established short-form test sets carry an estimated average label error rate of at least 3.3% [5], and long-context items are harder to adjudicate. Build in these checks:

  1. Dual annotation with span adjudication: two annotators locate gold spans independently; disagreements go to a third reviewer.
  2. Exhaustive-search audit: for a sample, a reviewer searches the full context for any other span that also answers the question, which catches unlabeled second gold passages.
  3. Shortcut probes: run question-only (no context) and BM25-only baselines; items either can solve are rewritten.
  4. Position-invariance check: move evidence across depth buckets and confirm the reference answer does not change.
  5. Contamination screen: confirm documents are not public web text; see contamination-resistant evaluation design.

Document the set with a data card covering source organizations by type, collection period, preparation steps, annotation protocol and known limitations [4]. Pair it with the guidance on private eval sets vs public benchmarks when deciding what to keep held out.

Rights, privacy and handling for real long documents

Real long documents are useful precisely because they contain operational detail, so rights and privacy review must happen before any item is authored. Confirm the supplying organization owns the documents and that consents and contracts permit use for model evaluation, and decide how personal data is treated. Pseudonymize consistently across a document set: if a customer name becomes "Acme-17" in the contract, it must be "Acme-17" in every amendment and email, or cross-document questions break.

At SourceX, every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Health records require HIPAA de-identification by Safe Harbor or Expert Determination. Delivery runs through private, access-controlled workflows, never email attachments, and only after an executed agreement and supplier approval.

Buyer checklist for a long-context evaluation request

A precise request describes the documents and the test, not the companies. Use this as a starting brief:

  • Document families and formats: for example, master agreements with amendments, full support threads, incident postmortems; DOCX, PDF, JSON exports.
  • Length profile: minimum native document length and the context buckets you will construct.
  • Version depth: whether superseded drafts, prior periods or duplicate cases must be present as distractors.
  • Cross-reference density: whether documents must reference each other (amendment to MSA, ticket to postmortem).
  • Privacy treatment: required pseudonymization consistency across documents.
  • Allowed uses: evaluation only, or also fine-tuning; whether results can be published.
  • Diligence materials: source, rights, preparation and allowed use for each dataset.

For the wider landscape of private test data, start at the evaluation datasets hub, and see summarization evaluation with expert references and RAG faithfulness labels for adjacent tasks. SourceX's owner pages on evaluation sets from real business work, RAG evaluation from company documents and licensing internal documentation cover those use cases directly.

Sourcing real long documents for long-context evaluation

SourceX sources operational datasets, including documents, engineering records and support histories, from US companies on request and manages the commercial process from assessment through licensing. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Describe the long-document collections your evaluation needs at https://sourcex.si/buyers.

Frequently asked questions

Is NIAH still worth running?

Yes, as a cheap regression check on positional encoding and context-window plumbing. It should not be the headline metric, because near-perfect NIAH scores coexist with large drops on low-overlap tasks [1][2].

How many items does a long-context set need?

Size the set by cells, not totals: each combination of length bucket, depth bucket and question type needs enough items for a stable estimate. A grid of 4 lengths, 5 depths and 4 question types already has 80 cells, which is why reusing each item across positions matters.

Can we reuse RAG evaluation triples for long-context testing?

Partly. RAG triples give you questions and gold passages, but long-context tests also need full surrounding documents, realistic distractors and controlled placement. Treat them as separate artifacts that can share source documents.

Sources

  1. arXiv, "Make Your LLM Fully Utilize the Context" (2024). https://arxiv.org/pdf/2404.16811
  2. arXiv (Modarressi et al.), "NoLiMa: Long-Context Evaluation Beyond Literal Matching" (2025). https://arxiv.org/abs/2502.05167
  3. arXiv (Lin et al.), "Beyond the Needle's Illusion: Decoupled Evaluation of Evidence Access and Use" (2026). https://arxiv.org/abs/2601.20276
  4. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  5. arXiv (Northcutt, Athalye, Mueller; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. arXiv (Hsieh et al.), "RULER: What's the Real Context Size of Your Long-Context Language Models?" (2024). https://arxiv.org/abs/2404.06654
  7. arXiv (Yen et al.), "HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly" (2024). https://arxiv.org/abs/2410.02694
  8. arXiv (Liu et al.), "Lost in the Middle: How Language Models Use Long Contexts" (2023). https://arxiv.org/abs/2307.03172

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data