Skip to content

Software companies

How real bugs in private repositories become coding agent benchmarks

By SourceX Editorial · Updated

Short answer

A private repository coding benchmark turns a bug your team already fixed into a scored task: an AI coding agent gets the code as it stood before the fix plus the issue text, and passes only if its change makes the failing tests pass without breaking others. Private code matters because public repositories may already sit in training data.

Key takeaways

  • A usable task needs four linked pieces: the issue, a test that fails before the fix, the merged fix and a passing run after it.
  • Private repositories are valued for evaluation because their code and fixes are less likely to have been seen in training.
  • Reproducible builds and deterministic tests decide how many past bugs become tasks, more than repository size does.
  • Benchmark use means the code is run, not just read, so environment and dependency details matter.
  • Evaluation licenses should state whether tasks can be published and how long the buyer may keep the code.

What is an issue-to-patch coding benchmark?#

An issue-to-patch coding benchmark is a set of tasks, each built from a real bug report and the code change that fixed it, used to measure whether an AI coding agent can solve the same problem on its own. The agent sees the repository as it stood before the fix and a description of the problem. It never sees the fix or the tests used to grade it.

Public benchmarks in this style, such as SWE-bench, are built from open-source projects on GitHub. Grading is mechanical: the agent's change is applied, the test suite runs, and the task counts as solved only if the tests tied to the bug now pass and the others still do.

That mechanical grading is what makes these benchmarks useful. It also explains why only some bugs qualify. Without a test that captures the bug, there is nothing to grade against.

From a real bug to a scored task#

A real bug becomes a scored task by walking backward from a merged fix to the moment before it, then proving that the fix, and only the fix, turns a failing test green. The steps below are the usual construction flow, whoever ends up doing the work.

  • Find a merged pull request that references an issue and changes both application code and tests.
  • Check out the parent commit, the state of the repository just before the fix.
  • Apply only the test changes and run them; the new or changed tests should fail.
  • Apply the fix and run the tests again; the same tests should now pass.
  • Run the wider suite to confirm that tests which passed before still pass after the fix.
  • Write the problem statement from the issue text, removing any hint of the solution.
  • Capture the environment, such as language version, dependency lockfile and build commands, so the task runs the same way every time.

Why AI developers want private code for evaluation#

AI developers want private code for evaluation because a benchmark only measures skill if the model has not already seen the answers. Public repositories, their issues and their merged fixes may have been collected into training data, so a strong score on a public task can partly reflect memory.

History from a private repository is much less likely to have been seen, provided it was never published. That caveat matters: modules later open-sourced, public SDKs and code shared widely with partners lose most of the advantage, so flag them during scoping. Private history also looks more like the work coding agents are bought to do: internal frameworks, legacy modules, business rules that live in code, and conventions no public project shares. Fresh tasks drawn from real commercial software are hard for an evaluation team to produce any other way.

The same property creates the main obligation on the buyer's side. Tasks built from your code stay useful only while they stay out of training data and out of public view, which is why publication and retention terms matter.

Which repository traits make tasks usable?#

Repository traits decide how many past bugs can become tasks, far more than the size of the codebase. A modest service with fast, deterministic tests and tidy pull requests can yield more usable tasks than a large application that needs a staging database to run anything.

Which repository traits make tasks usable?
TraitWhy it mattersQuick check
Tests run in automationGrading depends on running tests without a personContinuous integration runs the suite on every pull request
Reproducible buildsEvery agent attempt must run in the same environmentLockfiles committed and build steps scripted
Deterministic testsFlaky tests make pass or fail meaninglessBuilds rarely need a rerun to go green
Issues written in sentencesThe problem statement comes from the issueBug tickets describe symptoms, not just a title
Pull requests that reference issuesLinks connect each fix to its problemIssue keys in pull request titles or branch names
Focused fix commitsMixed refactors blur the expected changeBug fixes merged separately from cleanups
Dependencies that can be fetchedThe environment must be rebuilt outside your networkNo packages that exist only in a private registry
Long historyMore past fixes means more candidate tasksYears of merged pull requests still in the repository

What usually disqualifies a bug#

A bug is usually disqualified when its fix cannot be verified automatically, or when the task would reveal something it should not. Expect many past fixes to fall out at this stage; that is normal and says little about code quality.

  • Fixes verified only by hand, such as visual changes in the user interface.
  • Tests that call external services, production databases or paid APIs.
  • Fixes bundled with unrelated refactors or dependency upgrades.
  • Issues whose text already contains the patch or names the exact line to change.
  • Security fixes, where the task would document how a vulnerability worked.
  • Fixtures or test data copied from real customer records.
  • Code that was ever public, such as an open-sourced library or a published SDK.

Illustrative: a fleet maintenance software company tests its history#

Illustrative: a fictional company sells fleet maintenance software to regional trucking companies. Its engineers track work in Jira, review code in GitHub and run tests in continuous integration on every pull request. The CTO wants to know whether the history could support an evaluation-only license.

An engineer filters for merged pull requests linked to Jira bugs that changed both source and test files, then replays a sample at each parent commit. Fixes in the preventive-maintenance scheduling engine and the parts-catalog import parsers reproduce cleanly: tests fail before the fix and pass after it. Most fixes in the reporting module drop out, because their integration tests need a seeded staging database, and the telematics connector is flagged because an older version was published as a public SDK.

The CTO decides to offer only the reproducible modules, with personal details removed from fixtures and commit metadata, under terms that allow evaluation and rule out training and publication. The rest of the history stays out of scope until its tests can run in isolation.

License terms that matter for benchmark use#

License terms for benchmark use must answer questions a plain training license often leaves open, because evaluation tasks keep their value only while they stay private.

License terms that matter for benchmark use
TermQuestion to settle
Permitted useIs the code licensed for evaluation only, for training, or for both?
PublicationMay the buyer publish tasks, scores or examples, and in what form?
Execution environmentWhere will the code run, and who can reach that sandbox?
Retention and deletionHow long may the buyer keep code and tasks, and how is deletion confirmed?
Contamination controlsMust the buyer keep the tasks out of its own training data?
RedistributionCan tasks be shared with other parties, such as outside evaluators?

How SourceX handles engineering history for evaluation#

SourceX scopes engineering histories through the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Preparation covers full-history secrets scanning and the removal of personal details from fixtures, issue text and commit metadata.

The SourceX Evidence Packet records the permitted use, such as evaluation only, alongside provenance, licensing rights, the privacy record and release authorization. The company approves each step and keeps ownership of its code, which is licensed rather than sold.

Frequently asked questions

Does the buyer need our whole repository?

Not always. Each task needs the repository snapshot at its base commit and enough environment detail to run the tests, but scope can be limited to specific modules or services. Narrower scope means fewer tasks, so the trade-off is worth settling during scoping rather than after delivery.

Can tasks come from feature work as well as bugs?

Yes, when a feature request comes with tests that define the expected behavior. Feature tasks are harder to grade fairly, because several correct designs may exist and tests written for one design can reject another. Bug fixes with clear failing tests remain the most reliable source.

Will the AI model see our tests?

No. The tests tied to each bug are held back from the agent and used only to score its change. That is why tests are part of what is licensed, and why they need the same review for secrets and personal data as the application code.

Will benchmark tasks reveal weaknesses in our product?

Tasks describe bugs that were already found and fixed, often long ago. The larger concern is what issue text and fixtures say about customers, which preparation removes. Publication terms control whether anyone outside the evaluation team ever sees a task.

Do our engineers have to build the tasks?

Not necessarily. Who constructs and validates tasks is agreed during scoping. Your engineers may be asked about build steps, flaky tests or unusual tooling, because that knowledge rarely sits in documentation and it decides how many past fixes reproduce.

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify