Software engineering histories for training and evaluating coding agents
A software engineering history dataset is a linked record of how a team changed its code: each issue joined to the commits, pull request, review comments and CI runs that resolved it, the deploy that shipped it and any incident that followed. SourceX sources these histories from software companies' private repositories and trackers, such as GitHub, GitLab, Bitbucket and Jira, removes secrets and personal data, and confirms the partner's rights to the code before anything is licensed.
Dataset manifest
Sourced to your spec- What it is
- Issues linked to commits, pull requests, code review, CI runs, deploys and incidents
- Typical systems
- GitHub Enterprise, GitLab, Bitbucket, Azure DevOps, Jira, Linear, Jenkins
- Typical history
- Varies by partner; fully linked history is usually shorter than commit history
- Modality
- Code diffs and repository snapshots, review text and structured CI and tracker events
- Delivery formats
- Agreed per order; JSONL task records with git bundles or snapshot archives
- Preparation
- Secrets and credentials removed; authors pseudonymized; customer data stripped from fixtures
- Licensing
- Partner confirms rights to in-scope code; client-owned code needs client authorization
- Availability
- Depends on partners with linked, owned and reproducible repositories; not guaranteed
What a delivery contains
Fields vary by source system and are fixed per order. A typical delivery includes:
| Field | Type | What it holds |
|---|---|---|
| task_id | string | Pseudonymous identifier for one unit of work, from issue to merged change. |
| repo | object | Pseudonymous repository ID, languages, build system, test framework and size band. |
| issue | object | The ticket as first filed and as later edited, with type, labels, priority and discussion. |
| base_commit | string | The commit the work started from, so the repository can be restored to the state the developer saw. |
| commits | array | Ordered commits with messages, diffs, timestamps and pseudonymous authors. |
| pull_request | object | Source branch, title, description, linked issues, reviewers, approvals, merge method and timestamps. |
| review_threads | array | Line-anchored review comments and replies, with resolution status and the commit that addressed each. |
| ci_runs | array | Pipeline runs per commit with job results, failing test names and log excerpts. |
| tests | object | Tests that fail before the change and pass after it, plus tests that pass on both sides. |
| deploys | array | Release and deploy events that shipped the change, including rollbacks. |
| incidents | array | Incidents and postmortems linked to the change, where the partner's tooling records the link. |
| env | object | Container definition or setup script that builds the code and runs its tests offline. |
| redactions | array | Each secret or personal-data item removed, with its file, kind and the replacement used. |
Example record
{
"task_id": "swe_6a1f20",
"repo": { "id": "repo_c41d", "languages": ["python", "typescript"], "build": "poetry",
"test_framework": "pytest", "size_band": "250k-500k_loc" },
"issue": { "id": "iss_3392", "type": "bug", "labels": ["billing", "regression"],
"title": "Annual invoices show the wrong proration after a mid-cycle seat change",
"body": "Customer [CUSTOMER_ID] added seats on day 12 of an annual plan; invoice prorated by month." },
"base_commit": "9f2c1e0b7a43d6e18c5f20a9b3d4e7f61a2c8b90",
"commits": [
{ "sha": "e41a07c", "author": "dev_17", "message": "Use daily proration for annual plans",
"diff_ref": "patches/swe_6a1f20/e41a07c.diff", "files_changed": 3 },
{ "sha": "b88d2f1", "author": "dev_17", "message": "Round per line item; handle leap years",
"diff_ref": "patches/swe_6a1f20/b88d2f1.diff", "files_changed": 2 }
],
"pull_request": { "id": "pr_1187", "branch": "fix/annual-proration", "opened_at": "2024-09-03T08:12:40Z",
"merged_at": "2024-09-04T16:55:02Z", "approvals": 2, "merge": "squash" },
"review_threads": [
{ "reviewer": "dev_04", "file": "billing/proration.py", "line": 88, "resolved": true,
"comment": "This divides by 365. What happens for a renewal that spans February 29?",
"addressed_by": "b88d2f1" }
],
"ci_runs": [
{ "commit": "e41a07c", "status": "failed",
"failing": ["tests/billing/test_invoice_totals.py::test_rounding_matches_ledger"] },
{ "commit": "b88d2f1", "status": "passed", "duration_s": 742 }
],
"tests": { "fail_to_pass": ["tests/billing/test_proration.py::test_annual_mid_cycle_seat_add",
"tests/billing/test_proration.py::test_proration_leap_year"],
"pass_to_pass_count": 412 },
"deploys": [ { "env": "production", "at": "2024-09-05T10:02:00Z", "rolled_back": false } ],
"incidents": [],
"env": { "dockerfile_ref": "env/repo_c41d/Dockerfile", "network_required": false },
"redactions": [ { "file": "config/settings.py", "kind": "api_key", "replacement": "[REDACTED]" } ]
}Synthetic record for illustration. Field names, structure and format are agreed per order.
What AI teams use it for
Build contamination-resistant coding benchmarks
Issues from private repositories, paired with the merged fix and the tests it made pass, become SWE-bench-style tasks that were never published, which closes the main route for benchmark contamination.
Train coding agents with verifiable rewards
Fail-to-pass tests in a reproducible environment give an automatic pass or fail signal for reinforcement learning, without a human grading every attempt.
Model code review
Line-anchored comments, the follow-up commits that addressed them and the final approval show what experienced reviewers object to and how changes get fixed.
Learn from failed builds
Failing CI runs followed by the commit that turned them green are natural examples of debugging from logs and test output.
Predict change risk
Changes linked to rollbacks, hotfixes and incidents label which pull requests caused trouble in production, the signal a pre-merge risk model needs.
Use-case guides: Coding agents, Private evaluation sets
What makes this data valuable
Linked end to end
Issue, commits, PR, review, CI and deploy share IDs, so a task can be followed from report to production.
Reproducible environments
Pinned dependencies and a container that builds offline, so tests can be rerun years later.
Discriminating tests
Tests that fail before the fix and pass after it, beyond tests that merely accompany the change.
Substantive review
Review threads with real objections and follow-up commits, not only one-click approvals.
Production outcomes
Rollbacks, hotfixes and incident links show which merged changes held up after release.
Codebase variety
Several repositories, languages, frameworks and team conventions rather than one large monorepo.
Why private engineering histories matter
Public coding benchmarks are mostly built from open-source repositories, and the same repositories are scraped into pretraining corpora. Over time, scores on public benchmarks mix problem-solving with recall of fixes a model may already have seen. Histories from private repositories largely avoid that, because the code, the issues and the fixes were never published.
They also differ from open source in ways that matter for agents. Commercial codebases have deadlines, internal libraries with thin documentation, requirements written by product managers rather than maintainers, and review conventions enforced by a small, stable team. An agent that will work inside a company's repository is better tested on work that looks like it.
Turning a merged pull request into a usable task
Not every merged change makes a good task. Usable ones are self-contained, have a clear problem statement and include tests that would catch a wrong fix. Issues are often written after the fix, or describe only a symptom while the real discussion happened in chat, so the problem statement has to come from the issue and its early comments, never from the fix. Some fixes span several pull requests or arrive bundled with unrelated refactors, and need to be split or dropped.
Tests cause the most trouble. Flaky tests produce false failures, and tests that depend on network access, wall-clock time or internal services cannot run in a sealed environment. Run each candidate several times before and after the fix, and keep only tasks whose results are stable.
Keeping evaluation tasks clean
A private benchmark stays useful only while its tasks stay out of training data, including your own. Split by repository rather than by task: fixes in the same codebase share helpers, naming and even near-identical bugs, so training on one task can hand a model the answer to its neighbor. Forks, shared internal libraries and code copied between services carry the same leak across repositories, so deduplicate before splitting. Keep the gold patches, review threads and commit messages of evaluation tasks out of every training mix.
What to check before licensing
- Confirm the partner owns the code. Agencies and consultancies often build software that belongs to their clients, and code written by contractors generally belongs to the company only if their agreements assign it.
- Ask how secrets were found and removed. Scanning has to cover the full commit history, CI logs and attachments, because a key deleted in a later commit is still in history.
- Check that customer data in test fixtures, database dumps, logs and issue attachments was removed, not only developer names and emails.
- Run a software composition scan on the sample. Vendored libraries and snippets copied from public sources carry their own licenses, including copyleft, and may also sit in public training corpora.
- Rebuild a sample of task environments yourself, and confirm the fail-to-pass tests fail before the fix and pass after it, repeatedly and offline.
- Measure linkage on the sample, meaning the share of pull requests with a linked issue, review comments and CI results, since the links are what turn commits into tasks.
- Agree how the code may be used, for example for training or for evaluation only, and whether tasks or results may be published.
How licensing works through SourceX
- 1
Define
Send the domain, modality, volume, format, timeline and permitted use you need.
- 2
Source
SourceX identifies businesses that hold matching data and are open to licensing it.
- 3
Qualify
Fit, rights and quality are checked, and you review samples before committing.
- 4
License
Scope, permitted use, exclusivity, price and obligations are agreed in writing.
- 5
Deliver
Approved data is prepared, de-identified where required and transferred securely.
Questions buyers ask
Can this data be used to build SWE-bench-style evaluations?
Yes, that is one of its main uses. Each task pairs an issue with the repository at its base commit, the merged fix and the tests that failed before it and passed after it. Because the repositories were never public, the tasks are far less exposed to contamination than benchmarks built from open-source projects, although vendored or copied public code can still overlap with what a model has seen.
How are secrets and credentials removed?
By scanning the full history, not just the latest snapshot. Secret scanners and custom patterns search every commit, CI log and issue attachment for API keys, tokens, private keys, passwords and connection strings, and matches are replaced with placeholders. Rewriting history changes commit hashes, so a delivery should carry a mapping. Treat any key found as live: the partner should rotate it whether or not it ships.
Are CI logs and test results included?
Where the partner's CI system kept them. Run status, job results and failing test names are usually available, while full logs may be truncated or deleted after a retention period. Logs can contain environment variables, internal hostnames and customer data, so they are scanned like code. Ask for log coverage by year if you plan to build debugging tasks from them.
Can the code be run, or is it text only?
It can be run when an environment comes with it. Executable tasks need pinned dependency versions and a container or setup script that builds and runs the tests offline. Older commits often depend on package versions or internal services that no longer exist, so environment rebuilding is scoped per order, and tasks that cannot be reproduced are flagged or dropped.
Does recent history include AI-generated code?
Possibly. Many teams have used AI coding assistants in recent years, and commits rarely record which lines a tool wrote. Some histories mark agent-authored work through bot accounts, co-author trailers or pull request labels, which can be kept as metadata. If you need human-written work, for example to avoid training on model outputs, set a date cutoff or ask for flagged commits to be excluded or labeled.
How far back does linked engineering history go?
It varies by partner and repository. Commits often go back further than anything linked to them, because issue tracking, code review and CI tend to be adopted later than version control, and moves between hosts or trackers can break links. Ask for the fully linked span separately, meaning the years in which issues, pull requests, reviews and CI results all connect.
Related datasets
- Proprietary codebases with full history
Complete private repositories with full version history, build files, tests and docs
- IT service management and incident histories
Incidents, problems, changes and requests with work notes, CI links and outcomes
- Human feedback and QA-scored work
Work items with scores, verdicts and corrections from the people who reviewed them
- Enterprise workflow and task execution histories
Linked task trajectories from request to outcome, across every tool the work touched
Evaluating this data for procurement?
Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.
Request dataset diligenceTell us what your models need
Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.
Updated 3 October 2026. Own data like this? See how companies license it to AI developers.