Evaluation and benchmarking datasets
Evaluation-only data licenses: terms buyers should negotiate
Quick answer
An evaluation-only data license lets your team use licensed records to test, score and compare AI models while prohibiting any training use. Whether it works in practice depends on five terms: a definition of the evaluation activities you actually perform, a no-training covenant that also reaches paraphrases and other derived material, rules for sending items to hosted model APIs, what results you may publish, and what you must delete, keep and prove when the license ends.
By SourceX Editorial · Updated
This guide does not cover trial licenses for testing a dataset before you buy it; those are in evaluation licenses and NDAs for dataset samples. The short supplier-side answer is at can I license data for evaluation only?, and general clause structure is in the AI data licensing agreement guide.
Three different agreements get called evaluation licenses
Confirm which kind of "evaluation" a draft grants, because the word covers testing your models, testing the data before purchase, and non-commercial research, and each forbids different things.
| Agreement | What is evaluated | Typical grant | Gap for a model-evaluation buyer |
|---|---|---|---|
| Evaluation-only production license (this page) | Your models, or models you are choosing between | Scoring and analysis for a fixed term; training prohibited | Fits only if the definitions below match your workflow |
| Trial or sample license | The dataset, before purchase | Short and revocable. NVIDIA's published sample-data evaluation license is limited, non-exclusive, revocable, non-transferable, non-sublicensable and solely for evaluating and testing NVIDIA technologies [1] | Purpose may be tied to the licensor's own products; access can end at will |
| Non-commercial research terms | Research findings | MS MARCO's datasets are intended for non-commercial research only and are provided without extending any license or other IP rights [2] | Does not cover evaluating a commercial model or release decision |
| Training license | Training, usually with evaluation | Broader grant, priced accordingly | Pays for rights you may not need; see the AI training data licensing guide |
An evaluation-only license is the narrower grant you buy when training rights are unavailable, unaffordable or unwanted, for instance because training on the items would destroy their value as held-out data.
Name the benchmark split, version and item count you receive
Some commercial benchmarks are partly open and partly licensed, so the license must identify the split, version and item count you receive. Patronus AI's documentation says only an open-source subset of FinanceBench is available in its datasets library and that the full benchmark is licensed from Patronus [3]. The FinanceBench paper describes 10,231 questions with answers and evidence strings, and its authors reported model results on a 150-case sample [4].
Published scores on an open sample are therefore not a baseline for the licensed full set, and the open sample is the part most likely to have entered training corpora, so treat it as a development split. SWE-Bench Pro uses three access policies in one benchmark: an open public set, a private held-out set kept for future overfitting checks, and a commercial set built from acquired startup codebases whose results are published while the code stays private [5].
Do not rely on a hosting site's license field. The Data Provenance Initiative reported license omission rates above 70% and license error rates above 50% on popular dataset hosting sites [6]. Attach the governing license text to the delivery manifest for each split and version. Licenses for open benchmarks are covered in checking public benchmark licenses for commercial use.
Define evaluation by the activities your team performs
List the permitted activities in the grant one by one. An undefined "evaluation purposes" clause leaves judge calibration, prompt iteration, checkpoint selection and few-shot use to argument later.
| Activity | What the licensor worries about | Position to negotiate |
|---|---|---|
| Scoring outputs against gold labels; human error analysis | Access only | Core permitted use |
| Regression runs in CI across model versions | Copies spread into pipelines and logs | Permit in named repositories with set log retention |
| Prompting an LLM judge with items, labels or rubrics | Items reach the judge model | Permit for self-hosted or approved API judges |
| Fine-tuning a judge, reward model or verifier on items | Training in substance | Excluded unless priced as a separate grant |
| Iterating prompts or agent scaffolds against items | Items shape a shipped artifact; the set overfits | Permit on a designated development split; test split for scoring only |
| Items as in-context (few-shot) examples | Items end up in production prompts | Permit inside evaluation runs; prohibit in deployed prompts |
| Checkpoint or hyperparameter selection from scores | Scores steer training choices | Name it; aggregate scores are a lesser exposure than item content |
| Testing third-party models in a vendor bake-off | Items reach other model providers | Permit under the recipient rules below |
| Running evaluations for your own customers | Effectively a sublicense | Needs an express right and flow-down terms (rights that flow down to lab customers) |
Write the no-training covenant to reach derived material
The covenant should bar every route by which an item or its label can shape a model, not only "training", because rephrased test items inflate scores while evading verbatim checks. Yang and colleagues showed a 13B model trained on rephrased MMLU test items scored 85.9 on MMLU while n-gram overlap checks failed to flag it [7].
Cover these uses expressly: pre-training and continued pre-training; supervised fine-tuning, including adapters; preference pairs, reward models and verifiers; distillation that uses model outputs on items as targets; synthetic data generated from items, labels or rubrics; paraphrases, translations and templated variants; and production retrieval indices built from items. Licensors have their own reason to insist: Nasr and colleagues recovered thousands of training examples from aligned production models [8], so trained-on items can resurface in outputs. Leaked items also lose their value: in a February 2026 post, OpenAI said it had stopped reporting SWE-bench Verified because score gains increasingly reflected training-time exposure [9].
Pair the covenant with three buyer protections:
- Residuals. General learnings such as failure categories, capability gaps and rubric ideas may inform independently created training data, provided no item, label or close paraphrase is reproduced. Without this, any training work in the same domain looks like a breach.
- A licensor covenant. The licensor does not license the same items for training to anyone during the term and discloses other evaluation licensees. Each licensee is a leak path, and research on private evaluation curators describes the conflicts that arise when curators also do business with model developers [10].
- Agreed proof of exclusion. Item content hashes, a filter of your training corpora against them, and the overlap checks in contamination checks for licensed evaluation data become the evidence of compliance.
Decide who may receive items: APIs, contractors and enclaves
Evaluating a closed model means sending items to its provider, so the license must name each class of permitted recipient and the conditions it must meet.
- Hosted model APIs. Permit providers whose written terms exclude training on inputs and set a retention limit, and log every exposure by item and model. In a January 2024 staff post, the FTC said model-as-a-service companies that break promises not to use customer data for training or updating models may be liable under the laws it enforces [11]; rely on those written terms, not default settings.
- Contractors and evaluation vendors. Named, bound by equivalent terms, with you liable for their conduct.
- Affiliates. State whether they count as internal; internal-use-only licenses show where that line causes disputes.
- No transfer at all. Supplier-hosted or enclave runs where only scores leave are covered in evaluating models on data you can't take. Access-based sharing also lets a licensor end access at term: in Delta Sharing, the provider's sharing server manages recipient access [12].
If you source records through SourceX, delivery happens through private, access-controlled workflows, never email attachments, and nothing is delivered until an agreement is executed and the supplier approves the terms. SourceX itself does not train AI models.
Fix what may leave the evaluation environment
Agree the disclosure level before the first run, because publication rights are hard to add after a model launch. The options are: nothing; aggregate scores; per-slice scores with a named dataset and method; or a few licensor-approved example items. Full item release is incompatible with an evaluation-only grant.
SWE-Bench Pro's commercial set is a precedent for scores without items [5]. Private evaluation also gives up the transparency that lets outsiders check a public benchmark [10], so ask for the right to describe methodology, slice definitions and item counts, and to let a reviewer inspect items under NDA. Replace open-ended consent with a fixed-deadline pre-publication review, and settle whether the licensor may be named. Detail is in publishing benchmark results on licensed evaluation data.
Separate what you delete, what you keep and how you prove it
Write end-of-term obligations as three lists, so deletion does not destroy the regression history and compliance records you still need.
| Artifact | At end of term | Why |
|---|---|---|
| Item text, inputs and attachments, including harness copies and caches | Delete and certify | Licensed content |
| Gold labels and licensor rubrics | Delete | Licensed content |
| Per-item model outputs and judge rationales | Negotiate: delete, or keep under confidentiality | They contain item content but explain regressions |
| Embeddings or indices built from items | Delete | Derived copies |
| Backups | Expire on normal rotation; remain confidential | Immediate purge is often impractical |
| Aggregate and per-slice scores, methodology | Keep without time limit | Your results |
| Item IDs and content hashes | Keep | Prove training exclusion |
| Access and exposure logs | Keep for the audit period | Audit evidence |
Regulation can require records about test data. Under the EU AI Act, Article 10 applies quality and governance requirements to the testing data sets of high-risk AI systems, and where a system uses no training techniques, Article 10(6), as amended by Regulation (EU) 2026/1744, applies paragraphs 2, 3 and 4 and Article 4a(1) only to its testing data [13]. As of October 2026, after amendment by Regulation (EU) 2026/1744, these high-risk obligations reportedly apply from 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems [14]. If the model under test may sit in such a system, keep the right to retain provenance and preparation records after items are deleted.
Limit audit to records (access logs, exposure logs and hash-exclusion reports) reviewed by an independent auditor under confidentiality, at a capped frequency, with no licensor access to your weights or corpora. Expect a licensor to seek deletion of any model trained in breach; regulators have imposed a comparable remedy, as in the FTC's 2021 Everalbum order requiring deletion of models and algorithms developed from users' photos and videos [15]. Define "derived material" narrowly so legitimate outputs and aggregate results fall outside it; derivative and successor model rights covers the drafting.
Price the grant against a training license, not by volume
An evaluation-only grant removes the most valuable right, but secrecy, publication, recipients and term can push the price back up. Evaluation-specific price drivers:
- Secrecy. A cap on other licensees, or the licensor covenant above, preserves held-out value and costs more than a grant the licensor can resell freely.
- Publication rights. Scores you can cite in a model card or sales material are worth more than internal-only results.
- Recipients and scope. Third-party APIs, contractors and evaluation for customers widen exposure.
- Term and refresh. A term spanning several model releases, or periodic new items, adds value.
- Upgrade option. A pre-agreed formula to convert retired items to training rights avoids renegotiating later; compare fine-tuning-only data licenses.
SourceX does not publish prices; terms depend on scope, volume, history, rights and exclusivity, and are agreed per deal in writing. When you describe the evaluation uses you need licensed, list the activities from the table above; whether a supplier agrees to that scope is settled in the license, which defines the records included, permitted uses, term and delivery.
Illustrative clause set for an evaluation-only grant
Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.
1. "Licensed Items": the [N] items of dataset version [v], Test Split and
Development Split, identified by SHA-256 hash in Schedule A.
2. "Permitted Evaluation": (a) running Models on Licensed Items and scoring
outputs against Gold Labels; (b) human error analysis; (c) prompting Judge
Models with Licensed Items, Gold Labels and Rubrics; (d) prompt and scaffold
iteration on the Development Split only; (e) choosing among Model
checkpoints using Aggregate Results.
3. No Training: Licensee shall not use Licensed Items, Gold Labels, Rubrics or
Derived Item Content to pre-train, fine-tune, align, distill or otherwise
update the parameters of any model, including judge, reward and embedding
models, or to generate synthetic training data. "Derived Item Content"
includes paraphrases, translations, templated variants and model outputs
that reproduce a Licensed Item in substance.
4. Residuals: General knowledge gained from Permitted Evaluation may be used
for any purpose, provided no Licensed Item or Derived Item Content is
reproduced.
5. Permitted Recipients: employees; contractors listed in Schedule B; Model
Providers whose written terms prohibit training on inputs and limit
retention to [X days]. Every exposure is logged by item hash and model.
6. Disclosure: Licensee may publish Aggregate Results, slice definitions and
methodology, citing the dataset as "[agreed attribution]". Example items
need Licensor approval, deemed given if no objection within [10 business
days].
7. Licensor Covenant: during the Term, Licensor will not license the Licensed
Items for training use to any third party and will notify Licensee of each
additional evaluation licensee.
8. End of Term: within [30 days] Licensee deletes Licensed Items, Gold Labels,
per-item outputs and item embeddings, except Retained Records (Aggregate
Results, item hashes, exposure logs, provenance records). Backups expire on
normal rotation and remain confidential. An officer certifies deletion.
9. Audit: once per year, on [30 days'] notice, by an independent auditor under
confidentiality, limited to access logs, exposure logs and exclusion reports.
10. Upgrade Option: Licensee may convert retired Licensed Items to a training
license at [pricing formula].
Review checklist: who owns which clause
Split the review so each reviewer owns the clauses they can judge before the first draft goes back to the licensor.
- Evaluation lead: activity list, development and test splits, models and providers to be tested, per-item outputs you must keep.
- Counsel: grant wording, derived-material definition, residuals, licensor covenant, remedies, audit scope, and the licensor's right to license the records for this purpose (warranties to require).
- Security: storage location, named access list, exposure logging, API terms, deletion method and certificate (keeping a private eval set private).
- Privacy: personal data in items and outputs, and the de-identification method applied.
- Procurement: price basis, term and renewal, refresh deliveries, upgrade option; fallback positions are in the data license negotiation checklist.
Terms that quietly break these deals: "evaluation purposes" with no activity list, deletion that reaches aggregate results, a training ban without residuals, no permission for the APIs you test, and an unnamed split. The evaluation and benchmarking datasets hub covers the choice between licensed, commissioned and public test data.
Know which evaluation uses you need licensed?
Describe the records you need, the evaluation activities, recipients and disclosure rights your plan requires, and how long you need them. SourceX looks for US companies that hold such records, checks the data and the supplier's licensing permissions, and manages the license and delivery. Submit your licensing requirements.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Sources
- NVIDIA (published license terms, market practice), "NVIDIA Sample Data License for Evaluation (2026.01.19)" (2026). https://developer.download.nvidia.com/licenses/nvidia-sample-data-license-for-evaluation-2026.01.19.pdf
- Microsoft (MS MARCO project site), "MS MARCO datasets". https://microsoft.github.io/msmarco/Datasets.html
- Patronus AI (vendor documentation, market practice), "FinanceBench (Patronus Datasets documentation)". https://docs.patronus.ai/docs/financebench-1
- Islam et al. (Patronus AI, Contextual AI, Stanford), "FinanceBench: A New Benchmark for Financial Question Answering," arXiv:2311.11944 (2023). https://arxiv.org/abs/2311.11944v1
- arXiv:2509.16941, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/html/2509.16941v1
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI," arXiv:2310.16787; Nature Machine Intelligence 6 (2023-2024). https://arxiv.org/abs/2310.16787
- Yang et al. (LMSYS), "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples," arXiv:2311.04850 (2023). https://arxiv.org/pdf/2311.04850v1
- Nasr et al., "Scalable Extraction of Training Data from Aligned, Production Language Models," ICLR (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- arXiv:2503.04756, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- Databricks (vendor launch post, market practice), "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- Official Journal of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 amending Regulation (EU) 2024/1689 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
- Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.