Code and software engineering data
Evaluating a Code Dataset Sample Before You License It
Quick answer
Evaluate a code dataset sample by running it, not reading it. Before the sample arrives, agree a time-boxed protocol and written pass thresholds for build success rate, test-run rate, validated-task rate, issue-to-commit linkage rate, duplicate share against public code, residual secret findings and license-scan results. Then check that the sample was drawn from the same repositories and periods as the full delivery. A sample that passes objective checks earns a license negotiation; adjectives like "high quality" do not.
By SourceX Editorial · Updated
Why code samples need executable checks rather than inspection
A code sample proves value only when it builds, runs and links to the history you are paying for. Practitioners in community discussions report the same pattern across data categories: objective, checkable claims (valid format, duplicate rate, a measured pass rate) hold up in diligence, while suitability adjectives do not [1]. For code, the checkable claims are unusually cheap to verify, because a compiler, a test runner and a git log either succeed or fail.
Inspection misses the failure modes that matter most for coding agents. A repository can look clean in a viewer and still fail mvn package because a private Artifactory mirror was stripped, or pass pytest only because the suite is mostly skipped. Commit history can be squashed so issue links vanish. None of that shows up in a README; all of it shows up in a 48-hour harness run.
The metrics that decide a code data sample
Seven metrics cover most purchase decisions for training and evaluation code data. Define each one precisely in writing, because "build rate" means different things for a monorepo, a Gradle multi-project build and a set of Python packages.
- Build success rate: share of repositories (or build targets) that compile or install in a pinned container from the supplied lockfiles (
package-lock.json,poetry.lock,go.sum,Cargo.lock) with no network access beyond a public package mirror. - Test-run rate: share of repositories where the test command executes and reports results; separately record pass, fail, skip and error counts, since a 100% "run" with 80% skips is weak signal.
- Validated-task rate: for issue-and-fix or agent-task data, share of task instances where designated tests fail before the reference patch and pass after it. This is the fail-to-pass construction SWE-bench uses to turn real GitHub issues and pull requests into checkable tasks [3].
- Linkage rate: share of commits or pull requests that resolve to an issue, ticket or review thread (Jira key in the commit message, GitHub
Fixes #123, Gerrit Change-Id). - Distribution profile: language mix by bytes and by file, repository size, commit-date range, file-length percentiles and share of generated or vendored code (
node_modules/,vendor/, protobuf outputs, minified JavaScript). - Duplicate share against public code: share of files or functions that are exact or near-duplicates of code in your public reference corpus.
- Residual secrets and license findings: count of verified secrets per thousand files and count of files carrying copyleft or unknown license headers.
Testing duplication and contamination against your public corpus
Duplication against public code is the metric that most often turns a promising sample into a poor purchase. Near-duplicate content is common in training corpora, and removing it reduces memorized output and train-test overlap [6]. If 30% of a "proprietary" sample turns out to be vendored open-source libraries or forks of public repos, you are paying for data you already have.
Run the check against the corpus you actually train on. Teams that pretrain on public code commonly use The Stack or a derivative as the reference, since it is large, near-deduplicated and documents an opt-out process [7]. Use MinHash with locality-sensitive hashing over normalized token shingles for file-level near-duplicates, then exact hashing at function level, and report both shares. Our guide to deduplicating code training data covers forks, vendored copies and threshold choice in depth.
For evaluation data, even a small amount of contamination can invalidate a held-out set, so treat it as a gating check rather than a tolerable percentage. OpenAI stopped reporting SWE-bench Verified, citing contamination and flawed tasks, as of October 2026 [4]. SWE-Bench Pro responded by drawing tasks from copyleft public repositories and held-out commercial codebases [5]. If a sample is meant for a held-out evaluation set, check every repository, commit hash and issue text against GitHub search, Software Heritage and known mirrors, as described in code benchmark contamination and held-out SWE evaluation sets from private repositories.
Residual secrets and license findings as gating checks
Treat any verified live credential in a sample as a fail for the preparation process, not a minor defect. Language models can memorize and emit verbatim training sequences [8], so an AWS access key, a .env file or a private key block in training data is an exposure path. Run at least two scanners with different detection approaches, for example TruffleHog (pattern detectors with live credential verification) or Gitleaks (regex and entropy rules) alongside a custom pattern set for the supplier's internal token formats, and scan full history rather than the final tree. The secrets removal guide explains how to verify removal across git history.
License scanning answers a different question: whether third-party code inside the supplier's repositories carries terms the supplier cannot pass on. Run ScanCode Toolkit or a comparable scanner and flag GPL, AGPL, LGPL and SSPL headers, SPDX identifiers and copied snippets with no header. Findings are not automatically disqualifying, but they must be explained before signing; see copyleft contamination in licensed code and code ownership due diligence.
Representativeness: proving the sample matches the full delivery
A sample is only evidence if it was drawn from the same repositories, periods and preparation pipeline as the full delivery. A hand-picked set of the supplier's three best services will outperform the 400-repository delivery on every metric above. Ask how the sample was selected and request a manifest of the full population (repository identifiers, languages, commit counts, date ranges) so you can compare distributions.
Prefer a random or stratified draw specified by you, such as 5% of repositories stratified by language and size band. Then compare the sample's distribution profile with the manifest. If the full delivery is 60% Java but the sample is 90% TypeScript, your build and test rates do not transfer. Also confirm the same redaction and secrets pipeline ran on both, since a sample cleaned by hand is not a test of the production process.
A pre-agreed acceptance protocol you can send with the sample request
Agree the protocol before the sample arrives, so neither side moves the goalposts after seeing results. The table below shows the structure. Thresholds depend on your use case: training data for code completion tolerates lower test-run rates than agent-task data, where every instance must be validated.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Check | How measured | Example pass threshold | Fail action |
|---|---|---|---|
| Build success | Pinned container, lockfiles, public mirror only | At least 85% of repos build | Supplier explains failures by category |
| Test-run rate | Native test command; pass, fail, skip, error counts | At least 70% execute; skips under 20% | Exclude non-running repos from pricing basis |
| Validated tasks | Fail-before, pass-after on reference patch | At least 95% of task instances | Reject unvalidated instances |
| Linkage | Commits or PRs resolving to issue or review | At least 60% of PRs linked | Renegotiate scope or price basis |
| Public duplicate share | MinHash LSH vs. your reference corpus, plus function-level exact hash | Under 10% of files | Supplier removes overlap before delivery |
| Verified secrets | Two scanners, full history | Zero verified live secrets | Stop; supplier reruns preparation |
| License findings | ScanCode, SPDX and header scan | All copyleft hits explained in writing | Escalate to counsel |
| Representativeness | Sample vs. manifest distribution | Language shares within 10 points | Request a new buyer-specified draw |
Time-box the evaluation, typically to a fixed number of working days, and record the harness version, container image digests and scanner versions so results can be reproduced. Spot-check any human-assigned labels, such as review-comment categories or task difficulty ratings, with a blind second reviewer; even widely used benchmark test sets carry label error rates averaging at least 3.3% [9]. The generic mechanics of scoping a pilot are covered in how to run a data pilot with a supplier.
Sample terms, access and deletion
Expect sample terms to be narrower than the eventual license: evaluation-only, limited, revocable and without redistribution. NVIDIA's published sample data license for evaluation is one example of this shape [2]. Read the sample terms for whether you may train a throwaway model on the sample, because a small fine-tuning run is often the most informative test and some evaluation licenses forbid it.
Plan deletion before the sample arrives. Keep it in a separate bucket or project with access logging, avoid copying it into shared feature stores, and log the deletion when the trial ends, including derived artifacts such as embeddings, dedup indexes and fine-tuned checkpoints. If a supplier will not release code at all, consider supplier-hosted or enclave evaluation.
Turning sample results into a purchase decision
Sample metrics tell you whether the data is usable; a value estimate tells you whether it is worth the price. Convert results into a usable-volume figure, such as building repositories times linked PRs times the non-duplicate share, and price against that rather than raw repository count. Pair it with the approach in estimating a dataset's value before purchase and the supplier-level checks in evaluating data supplier quality.
If you are still writing the request, start with how to specify a code dataset request so build and test requirements appear in the brief rather than surfacing at evaluation. The code and software engineering data hub maps the other code data types, and the buyer intake at SourceX is where you can describe the repositories, history and tests you need.
Evaluate code data from US companies with SourceX
SourceX sources operational datasets, including engineering records, from US companies on request and manages the commercial process from assessment through licensing. Each dataset is rights-reviewed, comes with per-dataset diligence materials, and is released only with the supplying company's approval and under a license defining records, uses, term and delivery. Describe the code data you need to evaluate.
Frequently asked questions
How large should a code dataset sample be?
Large enough to estimate each metric within a useful margin across every stratum you care about. Ten repositories cannot tell you a build rate by language if the full delivery spans six languages. Specify sample size per language and size band rather than as a single number.
Should I fine-tune on the sample to measure lift?
Only if the sample terms allow training and the sample is big enough to move a held-out metric. For most code samples, harness metrics plus duplication checks are more decisive than a small fine-tuning run, which can be dominated by noise.
Does this protocol transfer to video dataset samples?
The structure transfers: pre-agreed thresholds, a buyer-specified draw, a time box and deletion. The metrics change to decode success, frame rate and resolution distribution, annotation agreement and duplicate clips. Many video sets ship as WebDataset tar shards in which files sharing a basename form one sample [10], so check that every sample in a shard has all expected components.
Sources
- Hugging Face Forums, "Is selling datasets way harder than building them? Or is it just me?". https://discuss.huggingface.co/t/is-selling-datasets-way-harder-than-building-them-or-is-it-just-me/177440
- NVIDIA, "NVIDIA Sample Data License for Evaluation" (2026). https://developer.download.nvidia.com/licenses/nvidia-sample-data-license-for-evaluation-2026.01.19.pdf
- Jimenez et al. (arXiv:2310.06770), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- Scale AI et al. (arXiv:2509.16941), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- Lee et al. (arXiv:2107.06499), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
- Kocetkov et al., BigCode (arXiv:2211.15533), "The Stack: 3 TB of permissively licensed source code" (2022). https://www.alphaxiv.org/abs/2211.15533v1
- Carlini et al. (arXiv:2012.07805), "Extracting Training Data from Large Language Models" (2020). https://arxiv.org/pdf/2012.07805
- Northcutt, Athalye, Mueller (arXiv:2103.14749), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.