Evaluation and benchmarking datasets
Evaluating models on data you can't take: supplier-hosted and enclave evaluation
Quick answer
You can evaluate a model on test data without keeping it, usually by moving the model to the data instead. There are three common arrangements: the data owner runs your model in its own environment, both parties use a neutral attested enclave, or you host a copy briefly under audited deletion. In every case the contract has to define exactly which outputs may leave (scores, per-slice metrics, error categories), how you verify the run was honest, and how your model weights and prompts are protected.
By SourceX Editorial · Updated
Why test data increasingly stays with its owner
Data owners keep test data in place because a released test set stops being a test set. Public benchmarks leak into pretraining corpora: in February 2026, OpenAI stopped reporting SWE-bench Verified because score gains increasingly reflected training-time exposure rather than capability [4]. SWE-Bench Pro responded by keeping a commercial subset built from private codebases and publishing only results on it [2].
Owners of operational data have a second reason: the records are commercially sensitive or regulated, and the owner may approve evaluation use while refusing to release copies. For you, the trade is simple. You get a test that is far less likely to be contaminated, and you give up the ability to inspect every item. The rest of this guide is about making that trade on good terms. For the broader case for held-out data, see private evaluation sets vs public benchmarks, and for the full map of test-data options, the evaluation datasets hub.
Three arrangements and what each one trusts
Each arrangement moves trust to a different party, so choose by what you can and cannot expose. Market practice also includes the reverse direction, where the buyer's ML team runs evaluations inside its own environment [3]; that is the third row below and the weakest isolation.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Arrangement | Where the model runs | What the buyer exposes | What the data owner exposes | Main failure mode | Fits when |
|---|---|---|---|---|---|
| Supplier-hosted run | Owner's infrastructure, against your API endpoint or a container with weights | Prompts, outputs and possibly weights to the owner | Nothing beyond agreed results | Owner misreports, cherry-picks or reruns until scores look a certain way | You serve via API anyway and the owner is not a competitor |
| Neutral attested enclave | A TEE such as an AWS Nitro Enclave, with keys released only to a measured image | Model inside the enclave only | Test data inside the enclave only | Harness bug or over-broad output channel leaks items | Both sides distrust each other, or weights must stay private |
| Buyer-hosted with deletion | Your environment, under access controls and a deletion attestation | Nothing | Full copy for a limited window | Copies, caches and logs survive deletion; test becomes contaminated | Low sensitivity, short window, strong audit rights |
A data clean room is a related but different tool: it is designed for joint computation and training on records neither party sees in full, covered in data clean rooms for AI training. Sharing protocols such as Delta Sharing, which exposes tables over REST APIs in open formats like Parquet, give the recipient read access to the data [5]. That is a transfer, not compute-to-data, and should be treated as one in the license.
How an enclave proves the evaluation ran as agreed
An enclave earns trust through remote attestation: a signed measurement of the exact code image, checked before any decryption key is released. In AWS Nitro Enclaves [8], for example, the enclave obtains an attestation document, and AWS KMS verifies it and encrypts its response to the enclave's public key, so only that enclave can decrypt the test set or the weights. Key policies can be written against the enclave's measurements (PCR values), and KMS denies the request when the running image does not match.
In practice, both parties should review and agree the evaluation harness, build it reproducibly, and pin its measurement hash in the key policy for the test-data key and, separately, the model key. Any harness change produces a new measurement and requires both parties to re-approve. The attestation proves which code ran; it does not prove the code is correct, so the harness itself still needs review. Typical leakage paths to check are stdout and debug logging, error messages that echo inputs, and result files with free-text fields.
What leaves the environment, and in what form
The output contract is the core of any no-transfer deal, because every allowed output is a possible leakage channel and every forbidden one is a blind spot. SWE-Bench Pro's commercial subset shows the minimum: aggregate results are released while the private code stays private [2]. That is rarely enough for a buyer making a model decision.
Agree in advance on three tiers. Tier 1 is aggregate scores with confidence intervals and item counts. Tier 2 is per-slice metrics on slices defined before the run, because models that score well overall often fail on specific subsets that are hard to discover after the fact [7]. Tier 3 is error categories from a fixed taxonomy, plus a small number of redacted or paraphrased failure exemplars the owner approves item by item. If the test items contain personal data, the exemplar process should follow the same rules as de-identifying evaluation data without breaking tests.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"run_id": "eval-2026-10-hosted-014",
"harness_measurement": "sha384:<PCR0 of agreed image>",
"model_ref": "buyer-model@ckpt-0912 (weights sealed to enclave key)",
"test_set_version": "owner-ticket-resolution-v3",
"items_scored": "<count>",
"metrics": { "resolution_accuracy": "<value>", "ci95": ["<low>", "<high>"] },
"slices": { "product_line": {"<slice>": "<value>"}, "ticket_age_days": {">30": "<value>"} },
"error_categories": { "wrong_policy_cited": "<count>", "missed_escalation": "<count>" },
"exemplars_released": "<n>, owner-approved, redacted",
"forbidden_outputs": ["raw inputs", "model outputs per item", "item IDs"]
}
Per-item model outputs deserve special care. They let you debug, but enough of them can reconstruct the test set, and an adaptive buyer who submits many models can probe item content through score differences. Cap the number of submissions per period and log every run.
Verifying results when you cannot see the items
You verify a hidden test through structure, not inspection. Bansal and Maini point out that a private curator who also sells training data has a conflict of interest, and that hidden evaluations lack the transparency that lets outsiders check them [1]. Your controls should assume that risk exists even with an honest owner.
Useful checks:
- Canary items you supply. Add a small set of items whose correct answers you know; if the reported score on them is wrong, the harness or scoring is wrong.
- Reference model. Run an open-weights model whose behavior you understand on every evaluation; large unexplained shifts flag a changed test set or harness.
- Frozen versions. Pin a test-set hash and harness measurement per run so scores across checkpoints are comparable.
- Label audit rights. Test sets carry label errors; an audit of 10 widely used test sets estimated an average error rate of at least 3.3% [6]. Negotiate a third-party or joint review of a sample of disputed items.
- Independent rerun. For high-stakes decisions, have a neutral party rerun the same measured harness and compare outputs.
These checks also help you judge the supplier itself; the supplier quality guide covers the wider questions.
Protecting your model when it travels to the data
Compute-to-data reverses the usual exposure: the data stays put, and your model, prompts and system instructions travel. In a supplier-hosted run against your API, the owner sees every prompt and output and could use them to distill or profile your model. Shipping weights in a container is worse unless they are sealed to an attested enclave key.
Contract for this explicitly: no retention of model outputs beyond the agreed result fields, no training on your outputs, rate limits on any endpoint you expose, and a dedicated, revocable key or endpoint per evaluation. Where weights must leave your control, insist on the enclave pattern and verify the measurement yourself rather than accepting a screenshot.
License terms specific to no-transfer evaluation
A no-transfer arrangement still needs a license, and its terms differ from a delivered eval set. Beyond the general points in evaluation-only data license terms, cover:
- The output contract (tiers, fields, forbidden outputs) as a schedule.
- Run caps, scheduling and who pays for compute.
- Harness approval and change control tied to measurements.
- Your right to publish or cite results, and in what form.
- Test-set refresh: how often items rotate and how version changes are reported, which matters for contamination-resistant evaluation design.
- Audit and dispute process for scoring and label errors.
- Treatment of your model artifacts, outputs and logs.
If you later want the data itself, that is a different deal; the contamination checks for licensed evaluation data apply once records move. When an owner will not release even a sample to scope the work, see sample access options for sensitive data. If you need operational records from US businesses as test data, you can describe the evaluation data to SourceX; each release still requires the supplying company's approval.
Find evaluation data that fits these constraints
SourceX sources operational datasets from US companies on request, including support histories, engineering records and finance and legal workflows; requests do not guarantee a match, and nothing is contracted until the supplying company agrees. Every dataset is rights-reviewed and licensed with defined records, uses, term and delivery. Describe the evaluation data you need and the constraints you can accept on the SourceX buyer page.
Frequently asked questions
Is a supplier-hosted evaluation enough for a production go/no-go?
Only with slice metrics, canaries and a reference model. Aggregate scores alone cannot show where the model fails, and you cannot check them against items you never see.
Can federated evaluation replace an enclave?
It distributes scoring across owners but still runs your model on their infrastructure, so it has the same output-contract and model-exposure questions as a supplier-hosted run, multiplied by the number of sites.
What if the owner wants to see my model's outputs?
Limit them to the agreed fields and make retention and reuse restrictions explicit. If per-item outputs are essential for the owner's own review, delete them on a fixed schedule and record the deletion.
Sources
- arXiv (Bansal and Maini), "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
- arXiv (Scale AI), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/html/2509.16941v1
- AIxBlock, "LLM Evaluation Datasets: Held-Out Sets for Production". https://www.aixblock.io/blogs/llm-evaluation-datasets
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
- arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (Eyuboglu et al.), "Domino: Discovering Systematic Errors with Cross-Modal Embeddings" (2022). https://arxiv.org/abs/2203.14960v2
- Amazon Web Services, "Cryptographic attestation support in AWS KMS" (2026). https://docs.aws.amazon.com/kms/latest/developerguide/cryptographic-attestation.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.