Training and evaluation data for coding agents
To train and evaluate a coding agent, you need real engineering tasks with a checkable outcome: an issue, a snapshot of the repository before the change, the merged pull request with its review thread, and the tests that prove it works, ideally from private repositories that were never public. SourceX sources linked issue-to-PR histories, full-history codebases, incident records and code review feedback from established software companies, with rights reviewed and secrets removed to an agreed standard before delivery.
Dataset types to start with
- Software engineering histories (issues, PRs, reviews)
Each issue tied to the commits, pull requests, review threads and CI runs that resolved it converts directly into an issue-to-PR task, with a reference fix and the tests that changed alongside it.
- Proprietary codebases with full history
Full git history with build files and test suites lets you rebuild a repository at any commit, so tasks run in a real environment, including internal frameworks and older languages that public code covers thinly.
- IT service management and incident histories
Incident, problem and change records tie production failures to the alerts, runbooks, postmortems and code changes that resolved them, which an agent needs to debug and operate software, not only write it.
- Human feedback and QA-scored work
Review verdicts, requested changes and the gap between a pull request's first revision and its merged version record how experienced reviewers judge code beyond pass or fail, a source of preference pairs and grading rubrics.
Why this data is hard to get
Public fixes may already be in pretraining data
Benchmarks mined from public repositories reuse issues and merged patches that sit on the open web, often next to forks and write-ups of the fix. A model trained on a later crawl may have seen the answer, so its score mixes recall with problem solving.
Public code is not what companies run
Open source skews toward libraries and developer tools built in public. Internal services, legacy systems, in-house frameworks and the glue between vendor products are rarely published, yet that is the code an enterprise coding agent is asked to change.
Private code carries layered rights
Repositories can hold client code written under work-for-hire terms, contractor code with unclear IP assignment, vendored libraries, copyleft files and export-controlled cryptography. Each has to be cleared or excluded before licensing.
Secrets persist in history
Deleting a key or customer fixture from the current tree leaves it in every earlier commit, so cleanup must scan and rewrite the full history, including commit metadata naming engineers.
The links live in separate systems
Issues sit in the tracker, code on the code host, builds in CI and outages in the incident tool. Exported one by one, they lose which pull request fixed which issue; only the company running all of them can rejoin them.
From merged pull request to verifiable task
A merged pull request that closes an issue is a candidate task. The issue and the repository at the base commit form the prompt, the merged diff is one reference solution, and the tests added or changed in the pull request are the grader. This is the issue-to-patch shape of SWE-bench-style benchmarks, built from code that was never public. Review comments add intermediate supervision: what a reviewer objected to and how the author responded. Failed CI runs before the final one show the feedback an agent will meet in its own loop.
Expect to filter hard: drop dependency bumps, formatting passes, generated code, refactors without a behavioral test, and issues that already contain the patch. Check that the fail-to-pass tests exercise the reported behavior rather than an unrelated assertion added in the same commit. What remains is smaller, but every task has an answer a program can check.
Changes without tests still carry weaker labels: approval without requested changes, a later revert, or an incident that points back to them. IT service management records supply that last link and a second task family: start from an alert, find the cause in the code, propose the fix.
Splitting for training and evaluation
Split by repository or by time, never by random task. Tasks from one repository share code, conventions and sometimes near-identical fixes, so a random split lets the model learn a test item's answer from its training neighbor. Hold out the most recent window of each repository, or whole repositories that never enter training. Ask for the evaluation split as a separate delivery with its own access list, so it cannot drift into a fine-tuning mix. The private evaluation sets page goes further on keeping held-out data out of training.
Specifying a coding data request
State what decides whether a task can run: languages and versions, build system, test coverage, buildable environments or linked histories only, and years of history. Add incident linkage if you need it, the language of issues and review threads, how author identities should be pseudonymized, and whether the data is for training, evaluation or both. SourceX matches the spec against partner repositories; for each candidate you see a manifest and sample tasks, and can check that they build and that their tests behave, before anything is licensed.
What good data looks like
- Each task links an issue to its base commit, merged diff, review thread and CI runs, using IDs that stay stable across systems.
- Tests from the fixing pull request fail on the base commit and pass with the fix applied.
- The repository builds at the base commit from lockfiles, pinned dependencies or a container spec delivered with it.
- The repository, issues and fixes were never published, and any public mirrors, forks or open-sourced parts are disclosed.
- The full history is scanned for credentials and customer data, author identities are pseudonymized, and the scan method is documented.
- Rejected revisions, superseded review rounds and failed CI runs are kept, not only the merged end state.
Questions buyers ask
Does a private repository rule out contamination completely?
No, but it closes the main route. Code that was never on the public web cannot have been crawled, yet parts of a private repository can exist publicly as mirrors, forks, open-sourced components or copied snippets. Ask about each repository's exposure, search public code for distinctive lines from your evaluation fixes, and favor tasks merged after the training cutoffs of the models you test.
Can I get runnable tasks with tests, not just diffs?
Where the partner's repositories support it. A runnable task needs tests that cover the changed behavior and an environment that still builds at the base commit, which older history often fails once dependencies are yanked or base images deleted. Ask for fail-to-pass tests and reproducible builds in your request; qualification checks which repositories and periods meet that bar.
How is client-owned or open source code inside a repository handled?
It is identified before licensing. Code written for a client under work-for-hire terms usually belongs to the client and is excluded unless the client authorizes the license. Vendored third-party code and files under copyleft or other open source licenses are flagged for removal or for handling under their own terms, and the partner confirms its licensing rights before delivery.
Can the data be used for reinforcement learning with execution feedback?
Yes, if the license permits it, so name reinforcement learning in your intended use. Test results and CI outcomes make natural rewards, but an agent can learn to pass weak tests by special-casing inputs or editing the tests. Keep held-out tests the agent never sees, block writes to test files, and review a sample of passing patches by hand.
Can I request a specific language, framework or domain?
Yes, including older languages, embedded firmware and in-house frameworks that are scarce in public code. Supply is not guaranteed: it depends on which software companies hold matching repositories and agree to license them. A narrow stack, or a minimum test coverage requirement, shrinks the pool of candidates.
Tell us what you are building
Describe the model or agent, the tasks it must handle, and the volume, format and permitted use you need. SourceX will match it to partner data.
Updated 3 October 2026.