Software companies
How real bugs in private repositories become coding agent benchmarks
By SourceX Editorial · Updated
Short answer
A private repository coding benchmark turns a bug your team already fixed into a scored task: an AI coding agent gets the code as it stood before the fix plus the issue text, and passes only if its change makes the failing tests pass without breaking others. Private code matters because public repositories may already sit in training data.
Key takeaways
- A usable task needs four linked pieces: the issue, a test that fails before the fix, the merged fix and a passing run after it.
- Private repositories are valued for evaluation because their code and fixes are less likely to have been seen in training.
- Reproducible builds and deterministic tests decide how many past bugs become tasks, more than repository size does.
- Benchmark use means the code is run, not just read, so environment and dependency details matter.
- Evaluation licenses should state whether tasks can be published and how long the buyer may keep the code.
What is an issue-to-patch coding benchmark?#
An issue-to-patch coding benchmark is a set of tasks, each built from a real bug report and the code change that fixed it, used to measure whether an AI coding agent can solve the same problem on its own. The agent sees the repository as it stood before the fix and a description of the problem. It never sees the fix or the tests used to grade it.
Public benchmarks in this style, such as SWE-bench, are built from open-source projects on GitHub. Grading is mechanical: the agent's change is applied, the test suite runs, and the task counts as solved only if the tests tied to the bug now pass and the others still do.
That mechanical grading is what makes these benchmarks useful. It also explains why only some bugs qualify. Without a test that captures the bug, there is nothing to grade against.
From a real bug to a scored task#
A real bug becomes a scored task by walking backward from a merged fix to the moment before it, then proving that the fix, and only the fix, turns a failing test green. The steps below are the usual construction flow, whoever ends up doing the work.
- Find a merged pull request that references an issue and changes both application code and tests.
- Check out the parent commit, the state of the repository just before the fix.
- Apply only the test changes and run them; the new or changed tests should fail.
- Apply the fix and run the tests again; the same tests should now pass.
- Run the wider suite to confirm that tests which passed before still pass after the fix.
- Write the problem statement from the issue text, removing any hint of the solution.
- Capture the environment, such as language version, dependency lockfile and build commands, so the task runs the same way every time.
Why AI developers want private code for evaluation#
AI developers want private code for evaluation because a benchmark only measures skill if the model has not already seen the answers. Public repositories, their issues and their merged fixes may have been collected into training data, so a strong score on a public task can partly reflect memory.
History from a private repository is much less likely to have been seen, provided it was never published. That caveat matters: modules later open-sourced, public SDKs and code shared widely with partners lose most of the advantage, so flag them during scoping. Private history also looks more like the work coding agents are bought to do: internal frameworks, legacy modules, business rules that live in code, and conventions no public project shares. Fresh tasks drawn from real commercial software are hard for an evaluation team to produce any other way.
The same property creates the main obligation on the buyer's side. Tasks built from your code stay useful only while they stay out of training data and out of public view, which is why publication and retention terms matter.
Which repository traits make tasks usable?#
Repository traits decide how many past bugs can become tasks, far more than the size of the codebase. A modest service with fast, deterministic tests and tidy pull requests can yield more usable tasks than a large application that needs a staging database to run anything.
| Trait | Why it matters | Quick check |
|---|---|---|
| Tests run in automation | Grading depends on running tests without a person | Continuous integration runs the suite on every pull request |
| Reproducible builds | Every agent attempt must run in the same environment | Lockfiles committed and build steps scripted |
| Deterministic tests | Flaky tests make pass or fail meaningless | Builds rarely need a rerun to go green |
| Issues written in sentences | The problem statement comes from the issue | Bug tickets describe symptoms, not just a title |
| Pull requests that reference issues | Links connect each fix to its problem | Issue keys in pull request titles or branch names |
| Focused fix commits | Mixed refactors blur the expected change | Bug fixes merged separately from cleanups |
| Dependencies that can be fetched | The environment must be rebuilt outside your network | No packages that exist only in a private registry |
| Long history | More past fixes means more candidate tasks | Years of merged pull requests still in the repository |
What usually disqualifies a bug#
A bug is usually disqualified when its fix cannot be verified automatically, or when the task would reveal something it should not. Expect many past fixes to fall out at this stage; that is normal and says little about code quality.
- Fixes verified only by hand, such as visual changes in the user interface.
- Tests that call external services, production databases or paid APIs.
- Fixes bundled with unrelated refactors or dependency upgrades.
- Issues whose text already contains the patch or names the exact line to change.
- Security fixes, where the task would document how a vulnerability worked.
- Fixtures or test data copied from real customer records.
- Code that was ever public, such as an open-sourced library or a published SDK.
Illustrative: a fleet maintenance software company tests its history#
Illustrative: a fictional company sells fleet maintenance software to regional trucking companies. Its engineers track work in Jira, review code in GitHub and run tests in continuous integration on every pull request. The CTO wants to know whether the history could support an evaluation-only license.
An engineer filters for merged pull requests linked to Jira bugs that changed both source and test files, then replays a sample at each parent commit. Fixes in the preventive-maintenance scheduling engine and the parts-catalog import parsers reproduce cleanly: tests fail before the fix and pass after it. Most fixes in the reporting module drop out, because their integration tests need a seeded staging database, and the telematics connector is flagged because an older version was published as a public SDK.
The CTO decides to offer only the reproducible modules, with personal details removed from fixtures and commit metadata, under terms that allow evaluation and rule out training and publication. The rest of the history stays out of scope until its tests can run in isolation.
License terms that matter for benchmark use#
License terms for benchmark use must answer questions a plain training license often leaves open, because evaluation tasks keep their value only while they stay private.
| Term | Question to settle |
|---|---|
| Permitted use | Is the code licensed for evaluation only, for training, or for both? |
| Publication | May the buyer publish tasks, scores or examples, and in what form? |
| Execution environment | Where will the code run, and who can reach that sandbox? |
| Retention and deletion | How long may the buyer keep code and tasks, and how is deletion confirmed? |
| Contamination controls | Must the buyer keep the tasks out of its own training data? |
| Redistribution | Can tasks be shared with other parties, such as outside evaluators? |
How SourceX handles engineering history for evaluation#
SourceX scopes engineering histories through the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Preparation covers full-history secrets scanning and the removal of personal details from fixtures, issue text and commit metadata.
The SourceX Evidence Packet records the permitted use, such as evaluation only, alongside provenance, licensing rights, the privacy record and release authorization. The company approves each step and keeps ownership of its code, which is licensed rather than sold.
Frequently asked questions
Does the buyer need our whole repository?
Not always. Each task needs the repository snapshot at its base commit and enough environment detail to run the tests, but scope can be limited to specific modules or services. Narrower scope means fewer tasks, so the trade-off is worth settling during scoping rather than after delivery.
Can tasks come from feature work as well as bugs?
Yes, when a feature request comes with tests that define the expected behavior. Feature tasks are harder to grade fairly, because several correct designs may exist and tests written for one design can reject another. Bug fixes with clear failing tests remain the most reliable source.
Will the AI model see our tests?
No. The tests tied to each bug are held back from the agent and used only to score its change. That is why tests are part of what is licensed, and why they need the same review for secrets and personal data as the application code.
Will benchmark tasks reveal weaknesses in our product?
Tasks describe bugs that were already found and fixed, often long ago. The larger concern is what issue text and fixtures say about customers, which preparation removes. Publication terms control whether anyone outside the evaluation team ever sees a task.
Do our engineers have to build the tasks?
Not necessarily. Who constructs and validates tasks is agreed during scoping. Your engineers may be asked about build steps, flaky tests or unusual tooling, because that knowledge rarely sits in documentation and it decides how many past fixes reproduce.
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.