Skip to content

Evaluation and benchmarking datasets

Keeping a private eval set private: access controls, canaries and API exposure

Quick answer

To keep an evaluation set private, treat it like a credential, not a document. Store items in one access-controlled location, split who can read items from who can only read scores, run scoring inside your own environment wherever possible, and send items to hosted model APIs only under retention settings you have confirmed in writing. Mark files with a unique canary, contract vendors to no redistribution, deletion and audit, and keep a retirement plan for items that do leak.

By SourceX Editorial · Updated

Where private eval items actually leak

Private eval items rarely leak through a breach; they leak through routine workflow. A held-out set loses its value the moment its items enter a training corpus, because contamination inflates scores without any real capability gain [4]. Public benchmarks show the end state: OpenAI stopped reporting SWE-bench Verified in February 2026 because score gains increasingly reflected training-time exposure [10].

The exposure paths for a private set are more mundane than scraping. Map each one before you write a policy:

  • Publication. Example items pasted into papers, model cards, blog posts, slide decks or GitHub issues. One worked example in a public README is enough to seed a crawler.
  • Contractors and annotators. Gold-label writers and adjudicators keep local copies, paste items into chat tools, or reuse them in later projects for other clients.
  • Eval and data vendors. A curator that holds your items and also sells training data has an incentive problem, which Bansal and Maini describe as a conflict of interest in private evaluation [3].
  • Hosted model APIs. Every item you send to a third-party endpoint becomes a prompt in that provider's logs, subject to its retention and use terms [6].
  • Internal sprawl. Copies in notebooks, CI artifacts, experiment trackers (prompt and completion tables in MLflow or Weights & Biases), shared drives and Slack threads.

Access tiers: who sees items and who sees only scores

The most effective control is separating item access from score access, so most people who need eval results never touch eval content. Bansal and Maini's analysis of private curators supports the same principle in reverse: the fewer parties that hold both the test items and a stake in the scores, the more trustworthy the result [3].

A three-tier model works for most teams. Keep the item-reader group small and named, issue scores through a service, and log every read of raw items.

Illustrative example: invented to show structure; it does not describe an available dataset.

TierWhoCan seeCannot seeTypical control
T1 CustodianEval lead, 1-3 named maintainersItems, gold labels, rubrics, canary IDs, item historyn/aSeparate cloud project or bucket, MFA, just-in-time access, object-level access logs
T2 GraderAdjudicators, rubric reviewersItems assigned to them, one batch at a timeFull set, gold labels for other batchesAnnotation tool with per-task assignment, no export, watermarked views
T3 ConsumerModel teams, product, leadershipAggregate and per-slice scores, failure categories, paraphrased exemplarsRaw items, gold answersScoring service returns metrics only; no item text in dashboards or tracker runs

Two failure modes break this model. The first is debug output: a scorer that logs full prompts and completions to an experiment tracker silently promotes T3 users to T1. The second is the "exemplar" slide, where a real item is shown in a review deck; write paraphrased exemplars for that purpose and keep them in a separate file. For the delivery-side view of the same controls on licensed data, see access controls for licensed training data after delivery.

Run scoring where the items live

Bring the model to the eval set, not the eval set to the model, whenever you can. Running evaluation inside the buyer's own environment keeps test data off third-party systems, an approach at least one eval-data vendor recommends [8]. For open-weight models, this means serving the checkpoint in your own VPC (vLLM, TGI or a managed endpoint in your account) and scoring there.

For closed models you can only reach through an API, items have to leave your boundary, so the question becomes what happens to them on arrival. Research on private benchmarking proposes evaluating models without revealing test data to the model owner at all, using trusted execution or cryptographic protocols [5]. Those designs are not yet routine procurement options, so most teams rely on contractual and configuration controls instead.

Scoring jobs should also be hardened. Pin the harness version, write outputs to the custodian bucket, and strip item text from anything that flows into shared CI logs.

Eval set leakage via hosted model APIs

Sending an eval item to a hosted API is a disclosure, so treat each provider's data controls as part of your eval protocol. As of October 2026, OpenAI's enterprise privacy terms, as recorded by a policy tracker, state that abuse-monitoring logs are generated for all API usage and retained for up to 30 days by default [6]. Zero Data Retention and Modified Abuse Monitoring require prior approval, are selected per organization or project, and do not cover every endpoint; some ineligible capabilities keep application state even with ZDR enabled [6].

The practical consequences for an eval lead are specific:

  • Use a dedicated project or organization for eval traffic, with the strongest retention setting the provider has approved for you, and confirm the setting is inherited rather than set to none.
  • Avoid stateful endpoints (stored conversations, assistants with file storage, vector stores, batch inputs kept as files) for eval runs unless you have confirmed their retention.
  • Do not upload eval files into fine-tuning, file or retrieval features "just to test"; uploaded files follow different lifecycle rules from transient prompts.
  • Record the configuration used for each run (provider, project ID, retention mode, endpoint, date) so you can answer later whether a given item was exposed.

Contractual promises matter too. FTC technology staff have written that model-as-a-service companies may face liability if they break commitments not to use customer data for undisclosed purposes such as training [7]. That gives written no-training terms weight, but it does not make retention zero; logs held for abuse review are still copies outside your control.

When you compare providers with the same set, rotate a fresh subset into each bake-off rather than replaying the full set to every vendor. The tradeoffs are covered in running a model bake-off with your own eval set.

Canary strings: what they detect and what they do not

A canary string is a unique marker embedded in eval files so that corpus builders can filter them out and so you can test later whether a model has seen them. BIG-bench put a canary GUID in every task definition file for exactly this purpose [1]. It is cheap insurance, but it is a filter request, not a lock.

Canaries have known limits. Community researchers reported that a production model could reproduce the BIG-bench canary, which suggests marked files still reached training data [2]. A canary also travels only with the file; once an item is copied into a prompt, a spreadsheet or a paper, the marker is usually gone.

Use canaries in layers:

  • Set-level GUID in every file header, README and dataset card, with a plain-language "do not train" notice.
  • Per-recipient canaries: give each vendor or contractor a copy carrying its own GUID and, where the format allows, a few unique decoy items. If a decoy surfaces, you know which copy leaked.
  • Periodic probes: ask candidate models to complete the GUID or decoy items, and log the results with the model version. A positive result is evidence; a negative result proves little.

Detection after the fact is a separate discipline. The SourceX guide to contamination checks for licensed eval data covers n-gram overlap, perplexity and completion tests.

License and vendor terms that back up the controls

Technical controls fail quietly, so the contract has to say what happens when they do. Whether you built the set in-house with contractors, commissioned it, or licensed it, the paper should cover the same points. For licensed sets, see terms to negotiate in evaluation-only data licenses; for vendor conflicts, see independence and verification for third-party eval vendors.

Illustrative example: invented to show structure; it does not describe an available dataset.

Eval data handling clause checklist

  • No redistribution, publication or inclusion of items, labels or close paraphrases in any dataset, model or product.
  • No use of items to train, tune, select or validate any model, including the vendor's own.
  • Named personnel or roles with access; subcontractors only with written approval.
  • Storage location, encryption and access logging requirements; no personal devices or consumer chat tools.
  • Deletion of all copies at project end, with a written certificate listing systems purged.
  • Audit right covering access logs and storage locations.
  • Prompt notice of suspected exposure, with a defined window.
  • Per-recipient canary and decoy items acknowledged in the agreement.

Leak response: retire, replace and re-baseline

Assume some items will leak and decide in advance how you will retire them. A leak-response plan turns a contamination scare into a routine maintenance task.

  1. Version every item with a stable ID, creation date and the list of recipients and providers it has been exposed to.
  2. Define triggers: a decoy or canary surfaces, an item appears in public, a vendor reports an incident, or a model's score on one slice jumps without explanation.
  3. Quarantine affected items and anything derived from them, then re-score past runs with and without them to size the effect.
  4. Replace from a reserve of unexposed items written to the same specification, then re-baseline the models you track.
  5. Report the change in eval release notes so score history is read correctly.

Keeping part of a benchmark permanently private is the approach SWE-Bench Pro takes, holding back a subset rather than publishing everything [9]. The same reserve logic underpins eval set refresh cadence and saturation planning, and the design choices that make leaks less damaging are covered in designing contamination-resistant evaluation sets.

Sourcing eval data with handling controls built in

Leak controls start before the data arrives. Eval items built from real business records are valuable precisely because they have not been published, so the sourcing path should not create copies you cannot track. If you are sourcing held-out data from operating companies, describe the eval data you need to SourceX: SourceX sources operational datasets from US companies on request, has every dataset rights-reviewed, and delivers through private, access-controlled workflows only after an executed agreement and supplier approval, never as email attachments. For the broader landscape, start at the LLM evaluation datasets hub or the AI data hub.

Find held-out evaluation data for your models

SourceX sources operational datasets such as support histories, engineering records and documents from US companies on request; categories are not inventory, and a request does not guarantee a match. Each dataset is rights-reviewed, has personal details removed or replaced before delivery, and is delivered under a license that defines records, uses, term and delivery. Describe the evaluation data you need.

Frequently asked questions

Does a no-training clause from an API provider make my eval set safe?

No. A no-training commitment limits how the provider may use your prompts, but abuse-monitoring logs may still be retained for a period unless you have an approved reduced-retention setting [6]. Treat the clause as one layer alongside retention configuration and item rotation.

Should I publish any items from a private eval set?

Publish paraphrased exemplars written for the purpose, never live items. If you need public comparability, release a separate public split and keep the scored split private, as some coding benchmarks now do [9].

How many decoy items should each recipient copy include?

Enough to be distinctive but too few to distort scores; a handful per copy is usually sufficient. Keep decoys out of scoring and record which copy carries which decoys.

Sources

  1. arXiv (BIG-bench authors), "Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models" (2022). https://arxiv.org/pdf/2206.04615
  2. AI Alignment Forum, "BIG-Bench Canary Contamination in GPT-4". https://alignmentforum.org/posts/kSmHMoaLKGcGgyWzs/big-bench-canary-contamination-in-gpt-4
  3. arXiv (Bansal and Maini; ICLR 2025), "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
  4. arXiv, "Benchmark data contamination study (arXiv:2410.09247)" (2024). https://arxiv.org/pdf/2410.09247.pdf
  5. Papers with Code, "Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs". https://astro.paperswithcode.com/paper/private-benchmarking-to-prevent-contamination
  6. Conduct Atlas, "OpenAI Enterprise Privacy (provision record: API data retention and Zero Data Retention option)". https://conductatlas.com/platform/openai/openai-enterprise-privacy/
  7. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  8. AIxBlock, "LLM Evaluation Datasets: Held-Out Sets for Production". https://www.aixblock.io/blogs/llm-evaluation-datasets
  9. arXiv (Scale AI), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
  10. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data