Skip to content

Evaluation and benchmarking datasets

Publishing benchmark results on licensed evaluation data

Quick answer

You can usually publish results from a licensed private eval set only to the extent the license says so, and most evaluation licenses are silent or restrictive about disclosure by default. Treat publication as its own negotiated grant with four tiers: aggregate scores, per-slice scores, example items, and full release. Settle supplier naming, reviewer access, model card wording and marketing claims in writing before the paper deadline, not after the camera-ready draft exists.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why disclosure is a separate grant from evaluation use

A right to run models against a dataset does not automatically include a right to tell anyone what happened. Evaluation-only licenses are typically drafted around access, internal use and confidentiality, so the score itself can fall under "confidential information" or "derived output" unless a clause carves it out. For the core grant, see evaluation-only data license terms and the owner answer on whether you can license data for evaluation only.

Software licensing has an old precedent for this conflict. "DeWitt clauses" in database and software EULAs bar publishing benchmark results without the vendor's approval, and their enforceability remains disputed [6]. A data supplier can impose the same restriction on an eval set, and a researcher who assumed academic freedom covers it learns otherwise at the worst moment.

The supplier's interests are concrete. Publishing items destroys the held-out value of the set, per-slice numbers can reveal the supplier's customer mix or failure patterns, and naming the source can expose a commercial relationship the supplier never announced. Your interest is equally concrete: reviewers, model card readers and procurement teams discount numbers they cannot trace.

The four disclosure tiers and what each exposes

Each step up the disclosure ladder exposes more of the licensed asset and needs a more explicit grant. SWE-Bench Pro is a useful public reference: it divides 1,865 problems from 41 repositories into public, held-out and commercial subsets, reports commercial-subset results while the underlying codebases stay private, and keeps held-out results unpublished [1][2]. As of October 2026, its public leaderboard reports results on the public split only [3].

TierWhat you publishWhat it can revealTypical condition to negotiate
1. Aggregate scoreOne headline metric per model (pass rate, accuracy, F1)Very little about items; the set's difficultyPermitted with a standard description and no supplier name unless agreed
2. Per-slice scoresResults by task type, domain, difficulty or languageSlice taxonomy, class balance, possibly the supplier's business mixPre-agreed slice list; minimum slice size; no slice that maps to a single customer or account
3. Example itemsA handful of prompts, inputs, reference outputs or rubric linesThe data itself, including any residual personal or commercial detailNamed items pre-approved by the supplier, re-redacted, and retired from scoring afterward
4. Full or partial releaseA public split, a dataset card, or the full setEverything; the set stops being held outSeparate publication license; usually only for a deliberately public split

Tier 2 is where most disputes start. A slice named "insurance claims escalations, Texas" may be harmless to you and commercially sensitive to the supplier, so agree the slice taxonomy when you write the evaluation dataset specification, not when you draft the results table.

Supplier identity, attribution and anonymity

Decide early whether the paper may name the source company, describe it generically, or say nothing beyond "a licensed proprietary dataset." Many suppliers allow a sector-level description ("support tickets from a US B2B software company") but not a name, a logo or a quote. If the supplier wants credit, the attribution string should be fixed in the license so co-authors do not improvise it.

A generic description has costs. Reviewers cannot check representativeness, and a later contamination question cannot be resolved by anyone outside your team. Pair anonymity with stronger documentation instead: a datasheet-style description of collection period, composition, preparation and intended use [8] that the supplier has approved for publication.

On a SourceX deal, every release is approved by the supplying company, and buyers describe the data they need rather than naming businesses. Agree the attribution wording as part of that approval rather than assuming it.

Answering reviewer reproducibility requests without releasing items

Reviewers increasingly treat unreleased eval data as a credibility problem, not a neutral choice. Work on private data curators argues that closed evaluation reduces transparency and can introduce curator bias and conflicts of interest [4], and test-set label errors averaging at least 3.3% across widely used benchmarks show why reviewers want to inspect labels [9]. Plan your response before submission.

Options that rarely require releasing items:

  • Evaluation-as-a-service: reviewers or third parties submit a model or endpoint, and you (or a neutral host) run it on the private set and return the score. Write this use into the license, including who may run it and for how long.
  • Escrowed access: a named reviewer or artifact-evaluation committee gets time-limited, logged read access under NDA. The supplier must approve the access path, and the license must permit it.
  • Public companion split: a small, supplier-approved public split drawn with the same pipeline, so others can check that your harness and scoring behave as described. This mirrors the public/held-out design in SWE-Bench Pro [1].
  • Harness and grader release: publish the evaluation code, prompts, rubric schema and scoring scripts without the data. This answers most "how was it scored" questions.
  • Statistical disclosure: item counts per slice, confidence intervals, inter-annotator agreement and known label-noise estimates, all of which reveal little about item content.

If you expect independent checking, read third-party eval vendor independence and data rights for independent evaluators before you pick a mechanism.

Contamination and the shelf life of a published number

Publishing anything about a private set shortens its useful life, so the disclosure plan and the contamination plan are one plan. Test data can leak into newer models' training corpora and make a benchmark obsolete; LiveBench responds by refreshing questions [10]. OpenAI stopped reporting SWE-bench Verified in February 2026, saying gains increasingly reflected training-time exposure [5].

Three practical controls belong in the publication clause. First, any example item you publish is retired from scoring and replaced. Second, published material carries a canary string so it can be detected and filtered from training corpora, an approach the BIG-bench benchmark uses [7]. Third, the paper states the set's version and freeze date, so a later score on a refreshed version is not compared with yours as if it were the same test. More on these controls is in keeping a private eval set from leaking and contamination-resistant evaluation design.

Model cards, papers and sales claims need different wording

The same score carries different risk in a paper, a model card and a sales deck, so negotiate each channel explicitly. A paper is read by peers who expect method detail. A model card persists and is quoted out of context. Marketing copy invites competitor scrutiny and, in some markets, regulatory attention to substantiation.

For model cards, agree a fixed description block: dataset name or generic descriptor, version, item count, slice list, metric definition, harness version and the statement that the set is licensed and not public. For marketing, many suppliers want approval of each claim, a ban on comparative claims against named competitors using their data, and the right to withdraw a claim if the set is later found to be contaminated.

Disclosure schedule to attach to an evaluation license

A short disclosure schedule turns a vague "publication subject to approval" into terms both sides can apply. Use the checklist below to draft it, then compare with your counsel's template.

Illustrative example: invented to show structure; it does not describe an available dataset.

disclosure_schedule:
  dataset_ref: "EVAL-SUPPORT-POLICY v1.2 (frozen 2026-08-31)"
  channels:
    peer_reviewed_paper: { tiers_allowed: [aggregate, per_slice], approval: "notice 20 business days before submission" }
    arxiv_preprint:      { tiers_allowed: [aggregate, per_slice], approval: "same as paper" }
    model_card:          { tiers_allowed: [aggregate], wording: "fixed description block, Annex B" }
    marketing:           { tiers_allowed: [aggregate], approval: "per claim, written", comparative_claims: false }
  per_slice_rules:
    approved_slices: [intent_type, channel, difficulty_band]
    min_items_per_reported_slice: 50
    forbidden_slices: [customer_account, named_product_line]
  example_items:
    max_items: 5
    selection: "supplier pre-approves item IDs"
    redaction: "re-check after selection; no names, emails, phones, account numbers"
    retire_from_scoring: true
    canary_string_required: true
  attribution:
    supplier_named: false
    approved_descriptor: "support conversations from a US B2B software company"
  reproducibility:
    eval_as_a_service: { allowed: true, operator: "licensee", requests_per_paper: 10 }
    reviewer_access: { allowed: true, form: "logged, time-limited, under NDA", max_days: 30 }
    public_companion_split: false
  survival: "approved publications remain permitted after license end"

Before signing, confirm each of these:

  1. The license grants a right to publish scores, not only to compute them, and says which channels it covers.
  2. Per-slice reporting has an approved slice list and a minimum slice size.
  3. Example items need named pre-approval, re-redaction and retirement from scoring.
  4. Supplier naming and the exact attribution string are fixed in writing.
  5. A reviewer reproducibility path exists and is permitted by the license.
  6. Model card wording is pre-agreed, and marketing claims have a separate approval route.
  7. Approval timelines fit your submission calendar, including rebuttal and camera-ready.
  8. Already-published results survive termination or expiry of the license.

How SourceX fits a licensed eval set you plan to publish on

SourceX sources operational datasets, such as support and sales histories, engineering records, documents, and finance and legal workflows, from US companies on request, and manages the licensing process; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, so agree publication terms before that license is signed. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery and a sample is checked, though no method is perfect, which is why example items still need re-review before you print them. You can describe your eval and publication plans when you submit a buyer request.

For wider context, start from the evaluation datasets hub, compare private eval sets and public benchmarks, or browse the AI data licensing guide.

Licensing an eval set you can publish on

SourceX sources operational data from US companies on request and manages the licensing process, with every release approved by the supplying company and terms set in the license. Nothing is contracted until a supplier agrees. To describe the data you need and how you plan to report results, start a buyer request.

Frequently asked questions

Can I publish an aggregate score if the license says nothing about publication?

Silence is not a grant. Confidentiality clauses often define outputs derived from confidential information broadly enough to cover a score. Get written confirmation, ideally as a disclosure schedule, before submitting.

Do I have to name the dataset supplier in my paper?

No, unless the license requires attribution. Most papers can use an approved generic descriptor plus a datasheet-style summary of collection and preparation. Make sure the descriptor cannot be used to identify the supplier.

What should I do if a reviewer asks for the test items?

Offer a mechanism the license already permits: evaluation-as-a-service, logged reviewer access under NDA, or the released harness and grader. Do not send items informally; that can breach the license and end the set's held-out value.

Does publishing a few example items contaminate the whole set?

It contaminates those items. Retire them from scoring, tag published copies with a canary string, and version the set, so later results are reported on items that were never public.

Sources

  1. arXiv (Scale AI), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/html/2509.16941v1
  2. Scale AI, "SWE-Bench Pro" (2025). https://scale.com/blog/swe-bench-pro
  3. Scale AI, "SWE-Bench Pro (Public) leaderboard". https://scale.com/leaderboard/swe_bench_pro_public
  4. arXiv, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
  5. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  6. David A. Wheeler, "The DeWitt clause's censorship should be illegal". https://dwheeler.com/essays/dewitt-clause.html
  7. arXiv (BIG-bench collaboration), "Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models" (2022). https://arxiv.org/pdf/2206.04615
  8. arXiv (Gebru et al.), "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
  9. arXiv (Northcutt, Athalye, Mueller), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  10. arXiv, "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data