Skip to content

Evaluation and benchmarking datasets

Evaluation-only data licenses: terms buyers should negotiate

Quick answer

An evaluation-only data license lets your team use licensed records to test, score and compare AI models while prohibiting any training use. Whether it works in practice depends on five terms: a definition of the evaluation activities you actually perform, a no-training covenant that also reaches paraphrases and other derived material, rules for sending items to hosted model APIs, what results you may publish, and what you must delete, keep and prove when the license ends.

By SourceX Editorial · Updated

This guide does not cover trial licenses for testing a dataset before you buy it; those are in evaluation licenses and NDAs for dataset samples. The short supplier-side answer is at can I license data for evaluation only?, and general clause structure is in the AI data licensing agreement guide.

Three different agreements get called evaluation licenses

Confirm which kind of "evaluation" a draft grants, because the word covers testing your models, testing the data before purchase, and non-commercial research, and each forbids different things.

AgreementWhat is evaluatedTypical grantGap for a model-evaluation buyer
Evaluation-only production license (this page)Your models, or models you are choosing betweenScoring and analysis for a fixed term; training prohibitedFits only if the definitions below match your workflow
Trial or sample licenseThe dataset, before purchaseShort and revocable. NVIDIA's published sample-data evaluation license is limited, non-exclusive, revocable, non-transferable, non-sublicensable and solely for evaluating and testing NVIDIA technologies [1]Purpose may be tied to the licensor's own products; access can end at will
Non-commercial research termsResearch findingsMS MARCO's datasets are intended for non-commercial research only and are provided without extending any license or other IP rights [2]Does not cover evaluating a commercial model or release decision
Training licenseTraining, usually with evaluationBroader grant, priced accordinglyPays for rights you may not need; see the AI training data licensing guide

An evaluation-only license is the narrower grant you buy when training rights are unavailable, unaffordable or unwanted, for instance because training on the items would destroy their value as held-out data.

Name the benchmark split, version and item count you receive

Some commercial benchmarks are partly open and partly licensed, so the license must identify the split, version and item count you receive. Patronus AI's documentation says only an open-source subset of FinanceBench is available in its datasets library and that the full benchmark is licensed from Patronus [3]. The FinanceBench paper describes 10,231 questions with answers and evidence strings, and its authors reported model results on a 150-case sample [4].

Published scores on an open sample are therefore not a baseline for the licensed full set, and the open sample is the part most likely to have entered training corpora, so treat it as a development split. SWE-Bench Pro uses three access policies in one benchmark: an open public set, a private held-out set kept for future overfitting checks, and a commercial set built from acquired startup codebases whose results are published while the code stays private [5].

Do not rely on a hosting site's license field. The Data Provenance Initiative reported license omission rates above 70% and license error rates above 50% on popular dataset hosting sites [6]. Attach the governing license text to the delivery manifest for each split and version. Licenses for open benchmarks are covered in checking public benchmark licenses for commercial use.

Define evaluation by the activities your team performs

List the permitted activities in the grant one by one. An undefined "evaluation purposes" clause leaves judge calibration, prompt iteration, checkpoint selection and few-shot use to argument later.

ActivityWhat the licensor worries aboutPosition to negotiate
Scoring outputs against gold labels; human error analysisAccess onlyCore permitted use
Regression runs in CI across model versionsCopies spread into pipelines and logsPermit in named repositories with set log retention
Prompting an LLM judge with items, labels or rubricsItems reach the judge modelPermit for self-hosted or approved API judges
Fine-tuning a judge, reward model or verifier on itemsTraining in substanceExcluded unless priced as a separate grant
Iterating prompts or agent scaffolds against itemsItems shape a shipped artifact; the set overfitsPermit on a designated development split; test split for scoring only
Items as in-context (few-shot) examplesItems end up in production promptsPermit inside evaluation runs; prohibit in deployed prompts
Checkpoint or hyperparameter selection from scoresScores steer training choicesName it; aggregate scores are a lesser exposure than item content
Testing third-party models in a vendor bake-offItems reach other model providersPermit under the recipient rules below
Running evaluations for your own customersEffectively a sublicenseNeeds an express right and flow-down terms (rights that flow down to lab customers)

Write the no-training covenant to reach derived material

The covenant should bar every route by which an item or its label can shape a model, not only "training", because rephrased test items inflate scores while evading verbatim checks. Yang and colleagues showed a 13B model trained on rephrased MMLU test items scored 85.9 on MMLU while n-gram overlap checks failed to flag it [7].

Cover these uses expressly: pre-training and continued pre-training; supervised fine-tuning, including adapters; preference pairs, reward models and verifiers; distillation that uses model outputs on items as targets; synthetic data generated from items, labels or rubrics; paraphrases, translations and templated variants; and production retrieval indices built from items. Licensors have their own reason to insist: Nasr and colleagues recovered thousands of training examples from aligned production models [8], so trained-on items can resurface in outputs. Leaked items also lose their value: in a February 2026 post, OpenAI said it had stopped reporting SWE-bench Verified because score gains increasingly reflected training-time exposure [9].

Pair the covenant with three buyer protections:

  • Residuals. General learnings such as failure categories, capability gaps and rubric ideas may inform independently created training data, provided no item, label or close paraphrase is reproduced. Without this, any training work in the same domain looks like a breach.
  • A licensor covenant. The licensor does not license the same items for training to anyone during the term and discloses other evaluation licensees. Each licensee is a leak path, and research on private evaluation curators describes the conflicts that arise when curators also do business with model developers [10].
  • Agreed proof of exclusion. Item content hashes, a filter of your training corpora against them, and the overlap checks in contamination checks for licensed evaluation data become the evidence of compliance.

Decide who may receive items: APIs, contractors and enclaves

Evaluating a closed model means sending items to its provider, so the license must name each class of permitted recipient and the conditions it must meet.

  • Hosted model APIs. Permit providers whose written terms exclude training on inputs and set a retention limit, and log every exposure by item and model. In a January 2024 staff post, the FTC said model-as-a-service companies that break promises not to use customer data for training or updating models may be liable under the laws it enforces [11]; rely on those written terms, not default settings.
  • Contractors and evaluation vendors. Named, bound by equivalent terms, with you liable for their conduct.
  • Affiliates. State whether they count as internal; internal-use-only licenses show where that line causes disputes.
  • No transfer at all. Supplier-hosted or enclave runs where only scores leave are covered in evaluating models on data you can't take. Access-based sharing also lets a licensor end access at term: in Delta Sharing, the provider's sharing server manages recipient access [12].

If you source records through SourceX, delivery happens through private, access-controlled workflows, never email attachments, and nothing is delivered until an agreement is executed and the supplier approves the terms. SourceX itself does not train AI models.

Fix what may leave the evaluation environment

Agree the disclosure level before the first run, because publication rights are hard to add after a model launch. The options are: nothing; aggregate scores; per-slice scores with a named dataset and method; or a few licensor-approved example items. Full item release is incompatible with an evaluation-only grant.

SWE-Bench Pro's commercial set is a precedent for scores without items [5]. Private evaluation also gives up the transparency that lets outsiders check a public benchmark [10], so ask for the right to describe methodology, slice definitions and item counts, and to let a reviewer inspect items under NDA. Replace open-ended consent with a fixed-deadline pre-publication review, and settle whether the licensor may be named. Detail is in publishing benchmark results on licensed evaluation data.

Separate what you delete, what you keep and how you prove it

Write end-of-term obligations as three lists, so deletion does not destroy the regression history and compliance records you still need.

ArtifactAt end of termWhy
Item text, inputs and attachments, including harness copies and cachesDelete and certifyLicensed content
Gold labels and licensor rubricsDeleteLicensed content
Per-item model outputs and judge rationalesNegotiate: delete, or keep under confidentialityThey contain item content but explain regressions
Embeddings or indices built from itemsDeleteDerived copies
BackupsExpire on normal rotation; remain confidentialImmediate purge is often impractical
Aggregate and per-slice scores, methodologyKeep without time limitYour results
Item IDs and content hashesKeepProve training exclusion
Access and exposure logsKeep for the audit periodAudit evidence

Regulation can require records about test data. Under the EU AI Act, Article 10 applies quality and governance requirements to the testing data sets of high-risk AI systems, and where a system uses no training techniques, Article 10(6), as amended by Regulation (EU) 2026/1744, applies paragraphs 2, 3 and 4 and Article 4a(1) only to its testing data [13]. As of October 2026, after amendment by Regulation (EU) 2026/1744, these high-risk obligations reportedly apply from 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems [14]. If the model under test may sit in such a system, keep the right to retain provenance and preparation records after items are deleted.

Limit audit to records (access logs, exposure logs and hash-exclusion reports) reviewed by an independent auditor under confidentiality, at a capped frequency, with no licensor access to your weights or corpora. Expect a licensor to seek deletion of any model trained in breach; regulators have imposed a comparable remedy, as in the FTC's 2021 Everalbum order requiring deletion of models and algorithms developed from users' photos and videos [15]. Define "derived material" narrowly so legitimate outputs and aggregate results fall outside it; derivative and successor model rights covers the drafting.

Price the grant against a training license, not by volume

An evaluation-only grant removes the most valuable right, but secrecy, publication, recipients and term can push the price back up. Evaluation-specific price drivers:

  • Secrecy. A cap on other licensees, or the licensor covenant above, preserves held-out value and costs more than a grant the licensor can resell freely.
  • Publication rights. Scores you can cite in a model card or sales material are worth more than internal-only results.
  • Recipients and scope. Third-party APIs, contractors and evaluation for customers widen exposure.
  • Term and refresh. A term spanning several model releases, or periodic new items, adds value.
  • Upgrade option. A pre-agreed formula to convert retired items to training rights avoids renegotiating later; compare fine-tuning-only data licenses.

SourceX does not publish prices; terms depend on scope, volume, history, rights and exclusivity, and are agreed per deal in writing. When you describe the evaluation uses you need licensed, list the activities from the table above; whether a supplier agrees to that scope is settled in the license, which defines the records included, permitted uses, term and delivery.

Illustrative clause set for an evaluation-only grant

Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.

1. "Licensed Items": the [N] items of dataset version [v], Test Split and
   Development Split, identified by SHA-256 hash in Schedule A.
2. "Permitted Evaluation": (a) running Models on Licensed Items and scoring
   outputs against Gold Labels; (b) human error analysis; (c) prompting Judge
   Models with Licensed Items, Gold Labels and Rubrics; (d) prompt and scaffold
   iteration on the Development Split only; (e) choosing among Model
   checkpoints using Aggregate Results.
3. No Training: Licensee shall not use Licensed Items, Gold Labels, Rubrics or
   Derived Item Content to pre-train, fine-tune, align, distill or otherwise
   update the parameters of any model, including judge, reward and embedding
   models, or to generate synthetic training data. "Derived Item Content"
   includes paraphrases, translations, templated variants and model outputs
   that reproduce a Licensed Item in substance.
4. Residuals: General knowledge gained from Permitted Evaluation may be used
   for any purpose, provided no Licensed Item or Derived Item Content is
   reproduced.
5. Permitted Recipients: employees; contractors listed in Schedule B; Model
   Providers whose written terms prohibit training on inputs and limit
   retention to [X days]. Every exposure is logged by item hash and model.
6. Disclosure: Licensee may publish Aggregate Results, slice definitions and
   methodology, citing the dataset as "[agreed attribution]". Example items
   need Licensor approval, deemed given if no objection within [10 business
   days].
7. Licensor Covenant: during the Term, Licensor will not license the Licensed
   Items for training use to any third party and will notify Licensee of each
   additional evaluation licensee.
8. End of Term: within [30 days] Licensee deletes Licensed Items, Gold Labels,
   per-item outputs and item embeddings, except Retained Records (Aggregate
   Results, item hashes, exposure logs, provenance records). Backups expire on
   normal rotation and remain confidential. An officer certifies deletion.
9. Audit: once per year, on [30 days'] notice, by an independent auditor under
   confidentiality, limited to access logs, exposure logs and exclusion reports.
10. Upgrade Option: Licensee may convert retired Licensed Items to a training
   license at [pricing formula].

Review checklist: who owns which clause

Split the review so each reviewer owns the clauses they can judge before the first draft goes back to the licensor.

  • Evaluation lead: activity list, development and test splits, models and providers to be tested, per-item outputs you must keep.
  • Counsel: grant wording, derived-material definition, residuals, licensor covenant, remedies, audit scope, and the licensor's right to license the records for this purpose (warranties to require).
  • Security: storage location, named access list, exposure logging, API terms, deletion method and certificate (keeping a private eval set private).
  • Privacy: personal data in items and outputs, and the de-identification method applied.
  • Procurement: price basis, term and renewal, refresh deliveries, upgrade option; fallback positions are in the data license negotiation checklist.

Terms that quietly break these deals: "evaluation purposes" with no activity list, deletion that reaches aggregate results, a training ban without residuals, no permission for the APIs you test, and an unnamed split. The evaluation and benchmarking datasets hub covers the choice between licensed, commissioned and public test data.

Know which evaluation uses you need licensed?

Describe the records you need, the evaluation activities, recipients and disclosure rights your plan requires, and how long you need them. SourceX looks for US companies that hold such records, checks the data and the supplier's licensing permissions, and manages the license and delivery. Submit your licensing requirements.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Sources

  1. NVIDIA (published license terms, market practice), "NVIDIA Sample Data License for Evaluation (2026.01.19)" (2026). https://developer.download.nvidia.com/licenses/nvidia-sample-data-license-for-evaluation-2026.01.19.pdf
  2. Microsoft (MS MARCO project site), "MS MARCO datasets". https://microsoft.github.io/msmarco/Datasets.html
  3. Patronus AI (vendor documentation, market practice), "FinanceBench (Patronus Datasets documentation)". https://docs.patronus.ai/docs/financebench-1
  4. Islam et al. (Patronus AI, Contextual AI, Stanford), "FinanceBench: A New Benchmark for Financial Question Answering," arXiv:2311.11944 (2023). https://arxiv.org/abs/2311.11944v1
  5. arXiv:2509.16941, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/html/2509.16941v1
  6. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI," arXiv:2310.16787; Nature Machine Intelligence 6 (2023-2024). https://arxiv.org/abs/2310.16787
  7. Yang et al. (LMSYS), "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples," arXiv:2311.04850 (2023). https://arxiv.org/pdf/2311.04850v1
  8. Nasr et al., "Scalable Extraction of Training Data from Aligned, Production Language Models," ICLR (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
  9. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  10. arXiv:2503.04756, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
  11. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  12. Databricks (vendor launch post, market practice), "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
  13. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  14. Official Journal of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 amending Regulation (EU) 2024/1689 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
  15. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data