Code and software engineering data
Held-Out Coding Agent Evaluation Sets from Licensed Private Repositories
Quick answer
A private SWE benchmark dataset is a set of issue-to-fix tasks, each with a pinned repository snapshot, a runnable environment and fail-to-pass tests, built from code that has never been public. It measures problem solving rather than recall of public GitHub. To source one, license tasks from several commercial codebases, split by repository or by time, deliver the eval split under its own access list, cap tasks per repository, and refresh from post-cutoff commits on a schedule.
By SourceX Editorial · Updated
Why public SWE benchmarks stop measuring what you care about
Public coding benchmarks lose signal because their repositories, issues and merged fixes sit in the same public corpora that pretraining crawls ingest. The original SWE-bench drew 2,294 tasks from 12 popular Python repositories on GitHub [6], and the 500-task SWE-bench Verified subset was human-screened for overly specific tests, underspecified issues and unreliable environment setup [5]. In February 2026 OpenAI said Verified was increasingly contaminated, that score gains increasingly reflected training-time exposure, and that it had stopped reporting the benchmark [4].
The gap between familiar and less-exposed sets is large. Scale reported that, at the launch of SWE-Bench Pro, the strongest models resolved about 23% of its public tasks under a standardized scaffold, against the 70%+ the same models typically reached on SWE-bench Verified [2]. Leaderboards move, so read the current public-subset board with its date before quoting any figure [3]. As of October 2026, treat any score on a fully public, multi-year-old Python benchmark as an upper bound on capability, not an estimate.
The broader point comes from general LLM evaluation: test items leak into newer models' training data and make a fixed benchmark obsolete quickly [7]. Canary strings such as BIG-bench's GUID help corpus builders filter known test files [8], but they only work when every crawler honors them and nobody pastes the content elsewhere. A held-out set from never-public code removes the leak path instead of hoping filters catch it.
The three-tier pattern you can copy for internal evals
The most practical design keeps one public tier for comparability, one held-out tier you control, and one commercial tier from code no model could have seen. SWE-Bench Pro uses exactly this structure: 1,865 problems from 41 actively maintained repositories, with a public set (731 instances, 11 repositories) and a held-out set (858 instances, 12 repositories) built only from strong-copyleft GPL repositories, plus a commercial set (276 instances, 18 repositories) from private startup codebases [1][2]. Held-out and commercial problems are not publicly accessible, though commercial-set results are published [1].
The GPL choice is a contamination hedge, not a guarantee: copyleft code is less likely to be included in commercial training corpora, but it is still public [1]. Only the commercial tier is structurally unseen. For an internal program, the mapping looks like this:
- Public tier: an existing benchmark subset you run for external comparability and sanity checks.
- Held-out tier: tasks you build or license and never release, used for model selection and regression tracking.
- Commercial private tier: licensed tasks from proprietary codebases, used as the headline capability measure and refreshed periodically.
Our comparison of private evaluation sets and public benchmarks covers when the private tiers pay for themselves across domains; this page focuses on the code-specific sourcing decisions.
Split policy: by repository or by time, never at random
Split the evaluation set by repository or by commit date, because random task-level splits leak repository-specific knowledge into training. Two tasks from the same repository share build files, module names, coding conventions and often the same maintainers' fix patterns. If one lands in training and its sibling in evaluation, the model is graded partly on memorized project structure.
Use a repository-level split when the eval must measure generalization to unfamiliar codebases, which is the usual question for a coding agent sold to many customers. Use a time-level split, where all eval tasks come from commits after a cutoff, when you want to track one codebase over time or measure performance on fresh work in familiar code. Combining both is stronger: eval repositories never appear in training, and eval tasks come from commits after the latest model training cutoff you need to test.
Deliver the eval split as a separate artifact with its own access list, storage location and audit log. When eval and training data share a bucket or a manifest, they eventually share a dataloader. Run the near-duplicate and fork checks described in our guide to detecting code benchmark contamination against your training corpus before the first scored run.
How many tasks, and from how many repositories
As a rule of thumb, a useful private coding eval needs a few hundred tasks spread across at least a dozen repositories, with a hard cap on tasks per repository. The size question is statistical: with a resolve rate near 30%, the standard error on 200 tasks is roughly 3 percentage points, so smaller sets cannot reliably detect the 2 to 4 point regressions most teams care about. Report confidence intervals, and resample by repository rather than by task, because tasks within a repository are correlated.
Repository diversity matters as much as count. SWE-bench's 12 Python repositories were all popular open-source libraries [6]; SWE-Bench Pro widened coverage to business applications, B2B services and developer tools across several languages [1]. For your own set, match the distribution of your users: languages, frameworks, monorepo versus polyrepo layout, code-base size, test-suite runtime and build systems such as Maven, Gradle, Bazel or npm workspaces.
Practical caps that keep a set honest:
- No single repository above roughly 10% of tasks, so one codebase cannot dominate the score.
- Each target language represented by at least three repositories from different owners.
- Task difficulty spread by lines and files changed, with a meaningful share of multi-file fixes.
- Issue text written by the original engineers, not regenerated by an LLM from the diff.
What each task must contain to be gradable
Every task needs a frozen base commit, the original issue or ticket text, a reference patch, fail-to-pass and pass-to-pass tests, and a container that builds offline. SWE-bench grades resolution by running tests after the model edits the codebase [6], and Verified's main fixes were for tests that were too specific and environments that failed to build [5]. Private commercial code makes both problems worse: internal package registries, licensed SDKs, secrets in configuration and flaky integration tests are common.
Ask suppliers to show that each environment builds without network access to their internal systems, that tests pass on the reference patch and fail on the base commit at least three times in a row, and that secrets were scanned across full git history before export. Our pages on SWE task environments with tests and building issue-to-fix pairs from private repositories go deeper on validation.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value | Why it matters |
|---|---|---|
| task_id | cx-ledger-0417 | Stable ID for regression tracking across runs |
| repo_pseudonym | repo_ledger_svc | Hides supplier identity from graders and model logs |
| split | heldout_commercial_2026Q3 | Ties the task to one access list and refresh cycle |
| base_commit_date | 2026-08-14 | Proves the task postdates the model cutoff you test |
| language / build | Kotlin / Gradle 8 | Feeds diversity quotas |
| issue_text | Original Jira description, PII redacted | Keeps the real ambiguity engineers faced |
| files_changed / lines | 3 / 64 | Difficulty stratification |
| fail_to_pass | LedgerReconcileTest.testPartialRefund | Primary grading signal |
| pass_to_pass | 212 tests | Catches fixes that break other behavior |
| env_image_digest | sha256:... | Reproducible offline build |
| validation_runs | 3/3 fail on base, 3/3 pass on patch | Filters flaky tests |
| canary | Per-set GUID in every file | Lets you detect leakage into crawled data [8] |
Access isolation and who should hold the eval
Keep the evaluation split where training pipelines cannot reach it, and be deliberate about which third parties can see it. Practical controls include a separate storage account or bucket with no training-role read access, a per-person access list, execution-only harnesses that return pass or fail without exposing patches, and logs of every scored run. Never send eval tasks to hosted model APIs without contract terms that bar retention and training on prompts.
Research on private evaluation has also flagged a conflict-of-interest risk: when the party that curates a private eval also supplies training data, or has commercial ties to the models being evaluated, the results can be biased in ways outsiders cannot audit [9]. If the same vendor builds your training data and your held-out eval, require separation of the task sources and the staff who write them, or source the eval from a different origin.
Licensing terms specific to evaluation-only use
An evaluation-only license on source code should define the permitted use narrowly and say what is forbidden, because evaluation data can be cheaper precisely when training is excluded. Generic terms to ask suppliers for include: use limited to scoring and error analysis; a no-training covenant covering fine-tuning, distillation and reward models; no publication of task text or patches, with an explicit rule on publishing aggregate scores; retention and deletion terms; and a list of who may access the set. See whether data can be licensed for evaluation only for the general question, and review ownership and copyleft risk in third-party or vendored code via our code ownership due diligence guide.
Refresh: keep the set ahead of training cutoffs
A held-out set decays as soon as models train on data that overlaps it, so plan a refresh cadence at the start. LiveBench's answer for general LLM tasks is to add new questions regularly from recent sources [7]; the coding equivalent is to pull new tasks from commits merged after the latest training cutoff you evaluate. A quarterly or half-year cycle is a reasonable starting point for teams with continuous releases; align it with your model release cadence.
Retire tasks deliberately. Keep a frozen anchor subset for longitudinal comparison, rotate the rest, and flag any task whose patch or issue text later appears in public code search, a fork or a mirror. Budget for the supplier relationship this implies: refresh only works if the supplying company keeps producing and approving new tasks.
Buyer checklist for a private SWE eval set
Illustrative example: invented to show structure; it does not describe an available dataset.
- Repositories never public, including forks, mirrors and package-published source
- Repository-level or time-level split documented; eval split delivered separately
- At least a dozen repositories, per-repository cap, languages matched to users
- Each task builds offline; tests validated against base and reference patch
- Secrets and personal data removed from code, history and issue text
- Evaluation-only terms, no-training covenant and publication rules agreed
- Separate access list, execution-only harness and run logs
- Refresh cadence and retirement rule written into the plan
How SourceX helps you source private coding eval data
SourceX sources operational datasets from US companies, including engineering records, and manages the commercial process from licensing through ongoing purchases. Data is sourced on request, not held in stock, so a request does not guarantee a match; you describe the code, languages and tasks you need rather than naming companies, and every release is approved by the supplying company. You can describe your coding eval requirements to SourceX and see the broader map of code and software engineering datasets and the coding agent use case.
Start a private coding eval request
Each dataset SourceX delivers is rights-reviewed for ownership and consents and comes under a license that defines records, uses, term and delivery; personal details are removed or replaced before delivery, and nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows after an executed agreement. Tell SourceX what held-out coding tasks you need.
Sources
- Scale AI (arXiv:2509.16941), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- Scale AI, "SWE-Bench Pro" (2025). https://scale.com/blog/swe-bench-pro
- Scale AI, "SWE-Bench Pro (Public Dataset) leaderboard". https://scale.com/leaderboard/swe_bench_pro_public
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
- Jimenez et al. (arXiv:2310.06770; ICLR 2024), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
- White et al. (arXiv:2406.19314; ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- Srivastava et al. (arXiv:2206.04615), "Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models" (2022). https://arxiv.org/pdf/2206.04615
- arXiv:2503.04756, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.