Code and software engineering data
Issue-to-Fix Pairs from Private Repositories: How to Build and Validate Them
Quick answer
An issue-to-fix pair is a task record that joins a bug report or feature request to the repository at its base commit, the merged reference patch, tests that fail before the patch and pass after it, and tests that must keep passing. Building them means linking issues to pull requests, filtering out changes that teach nothing, re-running tests in a pinned environment, and rejecting tasks whose tests do not check the reported behavior. Expect most merged pull requests to fall out of the funnel.
By SourceX Editorial · Updated
This guide covers the method and the acceptance rules. For what linked engineering history looks like as a dataset category, see software engineering datasets from private repos; for the wider landscape, start at the code and software engineering data hub.
What a valid issue-to-fix task contains
A valid task has five parts that can be re-executed by a third party: the issue text, a repository snapshot, a reference patch, a fail-to-pass test set and a pass-to-pass test set. SWE-bench popularized this shape: it drew 2,294 tasks from real GitHub issues and their linked pull requests across 12 Python repositories, and it judges a model's edit by running tests rather than comparing diffs [1]. The public SWE-bench dataset release names these fields problem_statement, base_commit, patch, test_patch, FAIL_TO_PASS and PASS_TO_PASS; check the current dataset card before you hard-code them.
The split between patch and test_patch matters. The model sees the issue and the repository at base_commit; the harness applies the hidden test changes, then runs the named tests. If the tests ship inside the code patch, or the issue already quotes the fix, the task leaks its own answer.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"instance_id": "acme-billing__ledger-svc-4812",
"repo": "acme-billing/ledger-svc",
"base_commit": "9f3c2e1",
"problem_statement": "Proration credit is negative when a plan downgrade lands on the last day of the cycle...",
"hints_text": "",
"patch": "diff --git a/ledger/proration.py b/ledger/proration.py ...",
"test_patch": "diff --git a/tests/test_proration.py b/tests/test_proration.py ...",
"FAIL_TO_PASS": ["tests/test_proration.py::test_downgrade_last_day_credit_nonnegative"],
"PASS_TO_PASS": ["tests/test_proration.py::test_upgrade_midcycle", "tests/test_invoice.py::test_totals"],
"environment": {"image": "python:3.11-slim", "lockfile": "poetry.lock", "setup_cmd": "poetry install --no-root"},
"linkage": {"issue_id": "LED-4812", "pr_id": 2290, "link_type": "closing_keyword"},
"created_at": "2025-03-14T10:22:05Z",
"split": "train"
}
The linkage block is worth requiring even though public benchmarks omit it. It records how the issue was tied to the pull request (a closing keyword, a tracker key in the branch name, or a manual link), which lets you audit link quality later.
Linking issues to pull requests in private history
Linkage is the first and largest loss point, because private teams rarely use one convention. In public GitHub projects, closing keywords such as "fixes #123" do most of the work. Inside companies, the issue usually lives in Jira, Linear or a support desk, and the link is a ticket key in the branch name, commit message or PR title, or a development-panel integration.
Ask a supplier for the raw linking evidence, not just a joined table. Useful fields are the tracker key, PR number, merge commit SHA, branch name, PR description, review comments and the tracker's own "linked PR" record. When two signals agree (a key in the branch name and a tracker link), treat the pair as high confidence; a key mentioned only in a review comment is weak.
Linked support tickets add a second layer of problem description. If your use case benefits from the customer's wording alongside the engineer's restatement, see SaaS support tickets linked to bug reports. For production errors that start from a stack trace rather than an issue, see stack trace to fix data.
Filters that remove invalid or low-signal tasks
Most merged pull requests should not become tasks, and the filters should be written down before anyone counts yield. The rules below are a practical starting set; tune thresholds per repository.
| Filter | What it catches | Typical detection |
|---|---|---|
| Answer in the issue | Issue text quotes the diff, the exact fix line, or links the PR | Diff-line overlap between problem_statement and patch; URL match to the PR |
| Dependency-only change | Version bumps in requirements.txt, package.json, go.mod, pom.xml with no source edits | Changed paths limited to manifests and lockfiles |
| Formatting-only change | Black, Prettier or gofmt runs; import reordering | Diff is empty after formatter plus whitespace normalization |
| No behavior test | Fix merged without a test that exercises it | test_patch empty, or no test changes its outcome |
| Multi-issue PR | One PR closes several unrelated issues | More than one linked issue key; changed files span unrelated modules |
| Revert or reverted fix | Fix later reverted or superseded | Later commit with "Revert" referencing the merge SHA |
| Vague issue | Title only, "doesn't work", or a screenshot with no text | Minimum length plus a human or model specificity check |
| Generated or vendored code | Edits to generated clients, minified bundles, vendored libraries | Path patterns such as vendor/, dist/, *_pb2.py |
OpenAI's audit of SWE-bench found several of these failure types in a published benchmark: tests that were overly specific to one implementation, issue descriptions too underspecified to solve, and environments that did not set up reliably [2]. Its human-verified subset kept 500 tasks [2]. Treat that as evidence that even curated sets carry invalid tasks, so a supplier's raw linked history will carry more.
Validating the tests, not just running them
A task is valid only when the reference patch flips the fail-to-pass tests and those tests check the behavior the issue describes. Running the suite once on the patched code proves almost nothing. Use a three-run protocol in a pinned container:
- Base run. At
base_commitwithtest_patchapplied, everyFAIL_TO_PASStest must fail, and everyPASS_TO_PASStest must pass. - Gold run. With
patchandtest_patchapplied, all listed tests must pass. - Repeat runs. Run both states at least three times. Any test that changes outcome without a code change is flaky and comes off both lists.
Then check relevance. A fail-to-pass test that fails at base because of an import error, a renamed fixture or a missing file is not testing the bug. Read the assertion against the issue: a proration bug should be caught by an assertion on the credit amount, not by a test that only checks the function exists. A small human review sample per repository, or a model-assisted check followed by human adjudication, catches these cases.
Two further checks reduce false acceptance. Mutate the reference patch (revert one hunk at a time) and confirm at least one fail-to-pass test fails again. Also confirm the PASS_TO_PASS set is broad enough to catch a fix that breaks neighboring behavior; a list of two tests in a module with two hundred is too thin for RL reward design. Environment requirements such as lockfiles, service containers and network isolation are covered in SWE task environments with tests.
Reporting the yield funnel and repository balance
A credible issue-to-fix dataset reports its attrition at every stage, so buyers can judge both quality and what was thrown away. Ask for a funnel like the one below for each repository, not only in aggregate.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Stage | Count | Share of previous stage |
|---|---|---|
| Merged PRs in window | 12,400 | n/a |
| Linked to an issue or ticket | 5,100 | 41% |
| Passing content filters | 2,300 | 45% |
| Test patch present and behavior-related | 1,050 | 46% |
| Environment builds at base commit | 760 | 72% |
| Fail-to-pass and pass-to-pass validated, non-flaky | 410 | 54% |
| Accepted after relevance review | 330 | 80% |
Balance matters as much as total count. SWE-Bench Pro reports task counts per repository and aims for balanced task distribution across codebases to prevent single-repository skew [4]. Cap the share any one repository contributes to a training mix, and report per-repository results on any eval split. Otherwise, a model can look strong by learning one codebase's conventions.
Leakage and contamination controls for private sources
Private repositories are valuable mainly because their code and history are unlikely to be in pretraining corpora. SWE-Bench Pro argues that permissively licensed public repositories are prime candidates for web-crawled pretraining data, and it holds out a commercial subset built from privately sourced codebases [4]. OpenAI stopped reporting SWE-bench Verified in February 2026, citing contamination, and as of October 2026 recommends other developers do the same [3].
Private does not mean clean. Check for public mirrors, forks, vendored copies of open-source dependencies, and code that was open-sourced after the task window. Keep train and eval splits separated by repository and by time, not by random task, and tag eval files with a canary string, as BIG-bench did with a GUID in every task file, so they can be filtered out of later corpora [5]. Detection methods are covered in code benchmark contamination, and split design in held-out coding agent evaluation sets.
Before any of this, scan the full git history for secrets and personal data, not just the head commit. Issue text and stack traces often contain customer emails, account numbers and tokens. See secrets in code datasets for history-wide scanning.
Using pairs for SFT, RL with test rewards and agent evals
The same validated task record serves three uses, but each one needs different fields. For supervised fine-tuning, the reference patch plus the PR description and review discussion give a target and rationale; commit histories with change rationale covers that context. For RL, the fail-to-pass and pass-to-pass lists are the reward function, so flaky or irrelevant tests directly corrupt the training signal.
Open releases already publish coding-agent trajectories labeled pass or fail by each task's tests, which shows how much downstream work depends on test quality [6]. For agent evaluation, keep a held-out, repository-disjoint split with no hints text and no review comments exposed to the agent.
What to require from a supplier's linked history
Specify acceptance rules in the request so a supplier can say early whether its history can meet them. A useful request names:
- Languages, build systems and test frameworks, such as Python with pytest, Java with Maven and JUnit 5, or TypeScript with Jest.
- Linkage evidence: tracker exports with keys, PR metadata, merge SHAs and review threads.
- Time window and repository count, plus any repositories to exclude.
- Environment reproducibility: lockfiles, base images, service dependencies and whether tests need network access.
- The filter table and three-run validation protocol above, with per-repository funnel reporting.
- Privacy handling for issue text, logs and commit authors.
The code dataset request specification guide gives a fuller template, and evaluating a code dataset sample explains how to test a sample against these rules. When you are ready to brief a sourcing partner, describe the data you need to SourceX.
SourceX sources operational datasets, including engineering records, from US companies on request; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. For how these records support coding agents, see coding agent training data from private repositories and what software bug-fixing records show.
Sourcing issue-to-fix history for your code models
SourceX looks for US businesses that hold the linked issue, pull request and test history you describe, then manages assessment of data and licensing permissions, a license defining records, uses, term and delivery, and the transaction. Nothing is contracted until a supplier agrees. Describe the issue-to-fix history you need.
Sources
- Jimenez et al. (Princeton, UChicago), arXiv / ICLR 2024, "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
- OpenAI, "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- Scale AI, arXiv:2509.16941, "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- Srivastava et al. (BIG-bench), arXiv:2206.04615, "Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models" (2022). https://arxiv.org/pdf/2206.04615
- Together AI, "CoderForge-Preview." https://together.ai/blog/coderforge-preview
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.