Code and software engineering data
SWE Task Environments with Tests: What to Require from a Coding Agent Dataset
Quick answer
A SWE task environment dataset is only useful if every task builds and runs its tests inside your infrastructure, without reaching back to the supplier. Require, per task, a pinned base image or build recipe, committed lockfiles, mirrored private dependencies, a single test command, recorded results at the base commit and after the reference patch, a time and resource budget, and a flake report from repeated runs. Reject tasks that need network access or undeclared services during grading.
By SourceX Editorial · Updated
Issue text and patches define what a task asks; this page covers the runtime contract that makes the task executable and gradeable. For how to build and validate the task content itself, see issue-to-fix pairs from private repositories, and for the wider category map see the code and software engineering data hub.
Why patches alone are not an RL environment
A patch plus an issue is a supervised example; an environment adds a runnable codebase, a test harness and a reward signal an agent can query repeatedly. SWE-bench established the execution-based pattern: each task pairs an issue and a repository snapshot with a reference pull request, and a candidate fix is judged by running the repository's own unit tests, including tests that fail before the reference change and pass after it [1]. Reproducing that runtime for every task is the expensive part, and it is the part suppliers most often under-deliver.
Public benchmarks show how environment defects corrupt results. OpenAI's annotation of SWE-bench found overly specific unit tests, underspecified issues and unreliable development-environment setup, which could cause correct solutions to be scored as failures [2]. As of October 2026, OpenAI no longer reports SWE-bench Verified, having stopped in February 2026 citing contamination [3], so its scores should not be treated as a clean reference. A purchased dataset inherits the same failure modes unless the contract names them.
The minimum per-task runtime contract
Every task should ship a machine-readable manifest that lets your harness build, run and grade it without human interpretation. A useful pattern, and the one the open-source SWE-bench harness follows, is to separate a base image (language and tooling), an environment image (repository-specific dependencies) and a thin per-task layer, all driven by one evaluation entry point. Ask suppliers for an equivalent layering so shared layers cache and per-task deltas stay small.
The fields below are a reasonable minimum to accept. Anything missing becomes a support ticket after delivery, when the supplier's engineers have moved on.
Illustrative example: invented to show structure; it does not describe an available dataset.
task_id: billing-svc-0412
repo_snapshot: billing-svc.tar.zst # full tree at base commit, .git included or stripped per license
base_commit: 9f3c2e1
language: python3.11
image:
base: python:3.11.9-slim-bookworm@sha256:<digest> # pinned by digest, never a floating tag
build_recipe: env/Dockerfile # delivered even if a built image is also shipped
platform: linux/amd64
dependencies:
lockfiles: [poetry.lock]
mirror: deps/wheelhouse/ # every wheel, incl. internal packages
os_packages: env/apt-packages.txt # name=version pins
build:
network: none
command: "pip install --no-index --find-links deps/wheelhouse -e ."
test:
command: "pytest -p no:cacheprovider -q tests/"
fail_to_pass: [tests/test_invoice.py::test_proration_rounding]
pass_to_pass_count: 1184
results_at_base: results/base.junit.xml
results_after_reference_patch: results/gold.junit.xml
runs_for_flake_check: 10
flaky_tests_excluded: [tests/test_export.py::test_s3_timeout]
budget:
wall_clock_s: 900
cpu: 4
memory_gb: 8
services:
- name: postgres
version: "15.6"
provided_as: sidecar image, seeded from fixtures/db.sql
secrets_required: none
Three fields carry most of the value. fail_to_pass names the tests that fail at the base commit and pass with the reference patch, which is what makes the task verifiable. pass_to_pass_count guards against reward hacking by deleting tests, and the two JUnit XML result files let you confirm both states independently before you train on anything.
Offline builds and private dependency mirrors
An environment that builds only on the supplier's network is not a deliverable. Private repositories routinely pull from internal Artifactory, Nexus, GitHub Packages or a private PyPI or npm registry, plus internal base images in an ECR or GCR account you will never have credentials for. Once the supplier's tokens expire, pip install, npm ci, mvn package or go mod download fails, and every task in the dataset fails with it.
Ask for an offline build proof: a log of each task built from scratch in a container started with --network none, using only artifacts in the delivery. Language-specific mechanisms make this concrete:
- Python: a wheelhouse plus
pip install --no-index --find-links, or a vendoreduvor Poetry cache with hashes in the lockfile. - Node:
npm ciagainst a packed offline cache, or a committed.yarn/cachewith Yarn's zero-install layout. - Java and Kotlin: a pre-populated
~/.m2or Gradle cache andmvn -oorgradle --offline. - Go:
go mod vendorwith-mod=vendor, andGOFLAGSplusGOPROXY=offset in the image. - Rust:
cargo vendorwith a.cargo/config.tomlsource replacement andcargo build --offline.
Mirrored internal packages are code too. They belong in the rights review alongside the main repository, and the secrets removal checklist for code datasets applies to them, because internal client libraries frequently embed endpoints and default credentials.
Determinism: flake rates and what to exclude
A task with a nondeterministic test produces a noisy reward, and in RL a noisy reward can do more harm than leaving the task out. Require the supplier to run every task's full test command several times at both the base commit and after the reference patch, and to report per-test outcome counts. Any test in the fail_to_pass set that ever flips should disqualify the task, and flaky pass_to_pass tests should be listed and excluded from grading.
Common sources of flakiness to ask about by name:
- Wall-clock dependence: tests that use
datetime.now(), time zones or daylight-saving boundaries without a frozen clock (freezegun,libfaketime). - Ordering: hash-seed or dict ordering,
pytest-randomly, parallel runners such aspytest-xdistsharing fixtures or ports. - Timing: sleeps, retries and timeouts tuned for the supplier's hardware, which fail under your CPU quota.
- Hidden network calls: SDK clients that resolve DNS, telemetry, or tests that hit staging APIs.
- Floating-point and locale differences between the supplier's images and yours, including
linux/amd64versusarm64.
The architecture point is practical: images built for linux/amd64 either fail or run under emulation on ARM hosts, and emulation changes timing enough to surface latent flakes. Specify your target platform and instance type in the request so the flake check runs where you will actually run.
Image delivery or build-recipe delivery
Shipping built OCI images is faster to adopt; shipping Dockerfiles plus mirrored artifacts is easier to audit and rebuild. Most buyers should require both: the recipe as the source of truth and the image as a convenience, with image digests recorded in the manifest so you can verify they were built from the recipe.
Size planning matters at scale. Without a registry of prebuilt images, every ephemeral training or evaluation worker rebuilds environments from source, which is slow and multiplies the chance of a drifting build. Ask suppliers to share base and environment layers across tasks from the same repository rather than shipping one fat image per task, and to state total registry size before delivery.
Images also carry third-party software. An image bundles a Linux distribution, OS packages, language runtimes and dependencies, each under its own license, separately from the license for the supplier's code. Ask for a software bill of materials (SPDX or CycloneDX, generated with a tool such as Syft) per image, and check it against your redistribution plans; the copyleft contamination guide covers the GPL and AGPL questions that follow.
Services, network isolation and the grading boundary
Grading should run with networking off, and every service a test needs must be declared and delivered. Real business repositories test against PostgreSQL, Redis, Kafka, Elasticsearch, S3-compatible storage or a third-party API. For each, the manifest should state whether it is provided as a sidecar container with seed data, replaced with an in-process fake (moto for AWS, an embedded broker), or mocked at the client boundary.
Keep the agent's environment and the grader separate. HCAST runs humans and agents in identical environments, with agents executing code inside the task container [5]; your version of that principle is that hidden fail_to_pass tests should be copied in only at grading time, after the agent finishes, so the agent cannot read or edit them. Success labels then reduce to the convention used in trajectory studies: a run succeeds when its patch passes the task's tests [6]. If you need richer outcome definitions, see task success labels for agent trajectories.
Private-code tasks, contamination and distribution limits
Tasks built from private commercial repositories are valuable precisely because the code is unlikely to be in pre-training corpora. The SWE-Bench Pro authors argue that permissively licensed public repositories are prime candidates for web-crawled training data, and they keep a commercial subset built from private codebases whose code is not publicly released while still reporting results on it [4]. That pattern shows a supplier can allow evaluation without wide redistribution of its source.
For buyers, the consequence is that the license and the runtime contract interact. A held-out evaluation set may need to stay inside a restricted cluster, while an RL training set needs to be replicated across many workers. Write the permitted environments, copies and retention into the license before engineering builds pipelines around them; the page on held-out SWE evaluation sets from private repos covers the evaluation side, and code benchmark contamination covers leak detection.
Acceptance checklist before you sign off on delivery
Run acceptance on your own hardware, with your own harness, against a random sample before accepting the full set. A practical gate:
Illustrative example: invented to show structure; it does not describe an available dataset.
| Check | How to run it | Reject the task if |
|---|---|---|
| Offline build | Build from recipe with --network none | Any fetch is attempted or the build fails |
| Image provenance | Rebuild and compare to the shipped digest, or inspect layers | Recipe and image disagree with no explanation |
| Base state | Run tests at base commit | Any fail_to_pass test passes |
| Reference state | Apply reference patch, run tests | Any fail_to_pass fails or pass_to_pass count drops |
| Determinism | Repeat both states several times | Any graded test changes outcome |
| Budget | Measure wall clock, CPU and memory | Exceeds declared budget on your target instance type |
| Isolation | Inspect manifest and runtime traffic | Undeclared service, secret or outbound call |
| Rights and hygiene | Secrets scan, SBOM review | Live credentials or unlicensed bundled software |
Sample evaluation before licensing is covered in more depth in evaluating a code dataset sample, and request wording in how to specify a code dataset request. Broader environment design beyond code is in RL environments from business workflows and the RL environment glossary entry.
Where SourceX fits
SourceX sources operational datasets from US companies on request, including engineering records, and manages the commercial process through licensing and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match: you describe the data you need, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, delivered under a license that defines the records, uses, term and delivery, and handed over through private, access-controlled workflows after an executed agreement. SourceX does not train models. See the coding agent use case and software engineering datasets from private repos, or start a buyer request.
Request executable coding tasks with a defined runtime
If you need repository tasks that build offline and grade deterministically, write the runtime contract above into your request. SourceX looks for US businesses that hold the engineering data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted. Describe the executable coding tasks you need.
Sources
- arXiv (Jimenez et al.), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
- OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- arXiv (Scale AI), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- arXiv (METR), "HCAST: Human-Calibrated Autonomy Software Tasks" (2025). https://arxiv.org/pdf/2503.17354
- arXiv, "Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories" (2025). https://arxiv.org/pdf/2506.18824
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.