Skip to content

Code and software engineering data

Evaluating a Code Dataset Sample Before You License It

Quick answer

Evaluate a code dataset sample by running it, not reading it. Before the sample arrives, agree a time-boxed protocol and written pass thresholds for build success rate, test-run rate, validated-task rate, issue-to-commit linkage rate, duplicate share against public code, residual secret findings and license-scan results. Then check that the sample was drawn from the same repositories and periods as the full delivery. A sample that passes objective checks earns a license negotiation; adjectives like "high quality" do not.

By SourceX Editorial · Updated

Why code samples need executable checks rather than inspection

A code sample proves value only when it builds, runs and links to the history you are paying for. Practitioners in community discussions report the same pattern across data categories: objective, checkable claims (valid format, duplicate rate, a measured pass rate) hold up in diligence, while suitability adjectives do not [1]. For code, the checkable claims are unusually cheap to verify, because a compiler, a test runner and a git log either succeed or fail.

Inspection misses the failure modes that matter most for coding agents. A repository can look clean in a viewer and still fail mvn package because a private Artifactory mirror was stripped, or pass pytest only because the suite is mostly skipped. Commit history can be squashed so issue links vanish. None of that shows up in a README; all of it shows up in a 48-hour harness run.

The metrics that decide a code data sample

Seven metrics cover most purchase decisions for training and evaluation code data. Define each one precisely in writing, because "build rate" means different things for a monorepo, a Gradle multi-project build and a set of Python packages.

  • Build success rate: share of repositories (or build targets) that compile or install in a pinned container from the supplied lockfiles (package-lock.json, poetry.lock, go.sum, Cargo.lock) with no network access beyond a public package mirror.
  • Test-run rate: share of repositories where the test command executes and reports results; separately record pass, fail, skip and error counts, since a 100% "run" with 80% skips is weak signal.
  • Validated-task rate: for issue-and-fix or agent-task data, share of task instances where designated tests fail before the reference patch and pass after it. This is the fail-to-pass construction SWE-bench uses to turn real GitHub issues and pull requests into checkable tasks [3].
  • Linkage rate: share of commits or pull requests that resolve to an issue, ticket or review thread (Jira key in the commit message, GitHub Fixes #123, Gerrit Change-Id).
  • Distribution profile: language mix by bytes and by file, repository size, commit-date range, file-length percentiles and share of generated or vendored code (node_modules/, vendor/, protobuf outputs, minified JavaScript).
  • Duplicate share against public code: share of files or functions that are exact or near-duplicates of code in your public reference corpus.
  • Residual secrets and license findings: count of verified secrets per thousand files and count of files carrying copyleft or unknown license headers.

Testing duplication and contamination against your public corpus

Duplication against public code is the metric that most often turns a promising sample into a poor purchase. Near-duplicate content is common in training corpora, and removing it reduces memorized output and train-test overlap [6]. If 30% of a "proprietary" sample turns out to be vendored open-source libraries or forks of public repos, you are paying for data you already have.

Run the check against the corpus you actually train on. Teams that pretrain on public code commonly use The Stack or a derivative as the reference, since it is large, near-deduplicated and documents an opt-out process [7]. Use MinHash with locality-sensitive hashing over normalized token shingles for file-level near-duplicates, then exact hashing at function level, and report both shares. Our guide to deduplicating code training data covers forks, vendored copies and threshold choice in depth.

For evaluation data, even a small amount of contamination can invalidate a held-out set, so treat it as a gating check rather than a tolerable percentage. OpenAI stopped reporting SWE-bench Verified, citing contamination and flawed tasks, as of October 2026 [4]. SWE-Bench Pro responded by drawing tasks from copyleft public repositories and held-out commercial codebases [5]. If a sample is meant for a held-out evaluation set, check every repository, commit hash and issue text against GitHub search, Software Heritage and known mirrors, as described in code benchmark contamination and held-out SWE evaluation sets from private repositories.

Residual secrets and license findings as gating checks

Treat any verified live credential in a sample as a fail for the preparation process, not a minor defect. Language models can memorize and emit verbatim training sequences [8], so an AWS access key, a .env file or a private key block in training data is an exposure path. Run at least two scanners with different detection approaches, for example TruffleHog (pattern detectors with live credential verification) or Gitleaks (regex and entropy rules) alongside a custom pattern set for the supplier's internal token formats, and scan full history rather than the final tree. The secrets removal guide explains how to verify removal across git history.

License scanning answers a different question: whether third-party code inside the supplier's repositories carries terms the supplier cannot pass on. Run ScanCode Toolkit or a comparable scanner and flag GPL, AGPL, LGPL and SSPL headers, SPDX identifiers and copied snippets with no header. Findings are not automatically disqualifying, but they must be explained before signing; see copyleft contamination in licensed code and code ownership due diligence.

Representativeness: proving the sample matches the full delivery

A sample is only evidence if it was drawn from the same repositories, periods and preparation pipeline as the full delivery. A hand-picked set of the supplier's three best services will outperform the 400-repository delivery on every metric above. Ask how the sample was selected and request a manifest of the full population (repository identifiers, languages, commit counts, date ranges) so you can compare distributions.

Prefer a random or stratified draw specified by you, such as 5% of repositories stratified by language and size band. Then compare the sample's distribution profile with the manifest. If the full delivery is 60% Java but the sample is 90% TypeScript, your build and test rates do not transfer. Also confirm the same redaction and secrets pipeline ran on both, since a sample cleaned by hand is not a test of the production process.

A pre-agreed acceptance protocol you can send with the sample request

Agree the protocol before the sample arrives, so neither side moves the goalposts after seeing results. The table below shows the structure. Thresholds depend on your use case: training data for code completion tolerates lower test-run rates than agent-task data, where every instance must be validated.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckHow measuredExample pass thresholdFail action
Build successPinned container, lockfiles, public mirror onlyAt least 85% of repos buildSupplier explains failures by category
Test-run rateNative test command; pass, fail, skip, error countsAt least 70% execute; skips under 20%Exclude non-running repos from pricing basis
Validated tasksFail-before, pass-after on reference patchAt least 95% of task instancesReject unvalidated instances
LinkageCommits or PRs resolving to issue or reviewAt least 60% of PRs linkedRenegotiate scope or price basis
Public duplicate shareMinHash LSH vs. your reference corpus, plus function-level exact hashUnder 10% of filesSupplier removes overlap before delivery
Verified secretsTwo scanners, full historyZero verified live secretsStop; supplier reruns preparation
License findingsScanCode, SPDX and header scanAll copyleft hits explained in writingEscalate to counsel
RepresentativenessSample vs. manifest distributionLanguage shares within 10 pointsRequest a new buyer-specified draw

Time-box the evaluation, typically to a fixed number of working days, and record the harness version, container image digests and scanner versions so results can be reproduced. Spot-check any human-assigned labels, such as review-comment categories or task difficulty ratings, with a blind second reviewer; even widely used benchmark test sets carry label error rates averaging at least 3.3% [9]. The generic mechanics of scoping a pilot are covered in how to run a data pilot with a supplier.

Sample terms, access and deletion

Expect sample terms to be narrower than the eventual license: evaluation-only, limited, revocable and without redistribution. NVIDIA's published sample data license for evaluation is one example of this shape [2]. Read the sample terms for whether you may train a throwaway model on the sample, because a small fine-tuning run is often the most informative test and some evaluation licenses forbid it.

Plan deletion before the sample arrives. Keep it in a separate bucket or project with access logging, avoid copying it into shared feature stores, and log the deletion when the trial ends, including derived artifacts such as embeddings, dedup indexes and fine-tuned checkpoints. If a supplier will not release code at all, consider supplier-hosted or enclave evaluation.

Turning sample results into a purchase decision

Sample metrics tell you whether the data is usable; a value estimate tells you whether it is worth the price. Convert results into a usable-volume figure, such as building repositories times linked PRs times the non-duplicate share, and price against that rather than raw repository count. Pair it with the approach in estimating a dataset's value before purchase and the supplier-level checks in evaluating data supplier quality.

If you are still writing the request, start with how to specify a code dataset request so build and test requirements appear in the brief rather than surfacing at evaluation. The code and software engineering data hub maps the other code data types, and the buyer intake at SourceX is where you can describe the repositories, history and tests you need.

Evaluate code data from US companies with SourceX

SourceX sources operational datasets, including engineering records, from US companies on request and manages the commercial process from assessment through licensing. Each dataset is rights-reviewed, comes with per-dataset diligence materials, and is released only with the supplying company's approval and under a license defining records, uses, term and delivery. Describe the code data you need to evaluate.

Frequently asked questions

How large should a code dataset sample be?

Large enough to estimate each metric within a useful margin across every stratum you care about. Ten repositories cannot tell you a build rate by language if the full delivery spans six languages. Specify sample size per language and size band rather than as a single number.

Should I fine-tune on the sample to measure lift?

Only if the sample terms allow training and the sample is big enough to move a held-out metric. For most code samples, harness metrics plus duplication checks are more decisive than a small fine-tuning run, which can be dominated by noise.

Does this protocol transfer to video dataset samples?

The structure transfers: pre-agreed thresholds, a buyer-specified draw, a time box and deletion. The metrics change to decode success, frame rate and resolution distribution, annotation agreement and duplicate clips. Many video sets ship as WebDataset tar shards in which files sharing a basename form one sample [10], so check that every sample in a shard has all expected components.

Sources

  1. Hugging Face Forums, "Is selling datasets way harder than building them? Or is it just me?". https://discuss.huggingface.co/t/is-selling-datasets-way-harder-than-building-them-or-is-it-just-me/177440
  2. NVIDIA, "NVIDIA Sample Data License for Evaluation" (2026). https://developer.download.nvidia.com/licenses/nvidia-sample-data-license-for-evaluation-2026.01.19.pdf
  3. Jimenez et al. (arXiv:2310.06770), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
  4. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  5. Scale AI et al. (arXiv:2509.16941), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
  6. Lee et al. (arXiv:2107.06499), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
  7. Kocetkov et al., BigCode (arXiv:2211.15533), "The Stack: 3 TB of permissively licensed source code" (2022). https://www.alphaxiv.org/abs/2211.15533v1
  8. Carlini et al. (arXiv:2012.07805), "Extracting Training Data from Large Language Models" (2020). https://arxiv.org/pdf/2012.07805
  9. Northcutt, Athalye, Mueller (arXiv:2103.14749), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  10. WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data