Skip to content

Software engineering histories for training and evaluating coding agents

A software engineering history dataset is a linked record of how a team changed its code: each issue joined to the commits, pull request, review comments and CI runs that resolved it, the deploy that shipped it and any incident that followed. SourceX sources these histories from software companies' private repositories and trackers, such as GitHub, GitLab, Bitbucket and Jira, removes secrets and personal data, and confirms the partner's rights to the code before anything is licensed.

Dataset manifest

Sourced to your spec
What it is
Issues linked to commits, pull requests, code review, CI runs, deploys and incidents
Typical systems
GitHub Enterprise, GitLab, Bitbucket, Azure DevOps, Jira, Linear, Jenkins
Typical history
Varies by partner; fully linked history is usually shorter than commit history
Modality
Code diffs and repository snapshots, review text and structured CI and tracker events
Delivery formats
Agreed per order; JSONL task records with git bundles or snapshot archives
Preparation
Secrets and credentials removed; authors pseudonymized; customer data stripped from fixtures
Licensing
Partner confirms rights to in-scope code; client-owned code needs client authorization
Availability
Depends on partners with linked, owned and reproducible repositories; not guaranteed

What a delivery contains

Fields vary by source system and are fixed per order. A typical delivery includes:

FieldTypeWhat it holds
task_idstringPseudonymous identifier for one unit of work, from issue to merged change.
repoobjectPseudonymous repository ID, languages, build system, test framework and size band.
issueobjectThe ticket as first filed and as later edited, with type, labels, priority and discussion.
base_commitstringThe commit the work started from, so the repository can be restored to the state the developer saw.
commitsarrayOrdered commits with messages, diffs, timestamps and pseudonymous authors.
pull_requestobjectSource branch, title, description, linked issues, reviewers, approvals, merge method and timestamps.
review_threadsarrayLine-anchored review comments and replies, with resolution status and the commit that addressed each.
ci_runsarrayPipeline runs per commit with job results, failing test names and log excerpts.
testsobjectTests that fail before the change and pass after it, plus tests that pass on both sides.
deploysarrayRelease and deploy events that shipped the change, including rollbacks.
incidentsarrayIncidents and postmortems linked to the change, where the partner's tooling records the link.
envobjectContainer definition or setup script that builds the code and runs its tests offline.
redactionsarrayEach secret or personal-data item removed, with its file, kind and the replacement used.

Example record

{
  "task_id": "swe_6a1f20",
  "repo": { "id": "repo_c41d", "languages": ["python", "typescript"], "build": "poetry",
            "test_framework": "pytest", "size_band": "250k-500k_loc" },
  "issue": { "id": "iss_3392", "type": "bug", "labels": ["billing", "regression"],
             "title": "Annual invoices show the wrong proration after a mid-cycle seat change",
             "body": "Customer [CUSTOMER_ID] added seats on day 12 of an annual plan; invoice prorated by month." },
  "base_commit": "9f2c1e0b7a43d6e18c5f20a9b3d4e7f61a2c8b90",
  "commits": [
    { "sha": "e41a07c", "author": "dev_17", "message": "Use daily proration for annual plans",
      "diff_ref": "patches/swe_6a1f20/e41a07c.diff", "files_changed": 3 },
    { "sha": "b88d2f1", "author": "dev_17", "message": "Round per line item; handle leap years",
      "diff_ref": "patches/swe_6a1f20/b88d2f1.diff", "files_changed": 2 }
  ],
  "pull_request": { "id": "pr_1187", "branch": "fix/annual-proration", "opened_at": "2024-09-03T08:12:40Z",
                    "merged_at": "2024-09-04T16:55:02Z", "approvals": 2, "merge": "squash" },
  "review_threads": [
    { "reviewer": "dev_04", "file": "billing/proration.py", "line": 88, "resolved": true,
      "comment": "This divides by 365. What happens for a renewal that spans February 29?",
      "addressed_by": "b88d2f1" }
  ],
  "ci_runs": [
    { "commit": "e41a07c", "status": "failed",
      "failing": ["tests/billing/test_invoice_totals.py::test_rounding_matches_ledger"] },
    { "commit": "b88d2f1", "status": "passed", "duration_s": 742 }
  ],
  "tests": { "fail_to_pass": ["tests/billing/test_proration.py::test_annual_mid_cycle_seat_add",
                              "tests/billing/test_proration.py::test_proration_leap_year"],
             "pass_to_pass_count": 412 },
  "deploys": [ { "env": "production", "at": "2024-09-05T10:02:00Z", "rolled_back": false } ],
  "incidents": [],
  "env": { "dockerfile_ref": "env/repo_c41d/Dockerfile", "network_required": false },
  "redactions": [ { "file": "config/settings.py", "kind": "api_key", "replacement": "[REDACTED]" } ]
}

Synthetic record for illustration. Field names, structure and format are agreed per order.

What AI teams use it for

Build contamination-resistant coding benchmarks

Issues from private repositories, paired with the merged fix and the tests it made pass, become SWE-bench-style tasks that were never published, which closes the main route for benchmark contamination.

Train coding agents with verifiable rewards

Fail-to-pass tests in a reproducible environment give an automatic pass or fail signal for reinforcement learning, without a human grading every attempt.

Model code review

Line-anchored comments, the follow-up commits that addressed them and the final approval show what experienced reviewers object to and how changes get fixed.

Learn from failed builds

Failing CI runs followed by the commit that turned them green are natural examples of debugging from logs and test output.

Predict change risk

Changes linked to rollbacks, hotfixes and incidents label which pull requests caused trouble in production, the signal a pre-merge risk model needs.

Use-case guides: Coding agents, Private evaluation sets

What makes this data valuable

Linked end to end

Issue, commits, PR, review, CI and deploy share IDs, so a task can be followed from report to production.

Reproducible environments

Pinned dependencies and a container that builds offline, so tests can be rerun years later.

Discriminating tests

Tests that fail before the fix and pass after it, beyond tests that merely accompany the change.

Substantive review

Review threads with real objections and follow-up commits, not only one-click approvals.

Production outcomes

Rollbacks, hotfixes and incident links show which merged changes held up after release.

Codebase variety

Several repositories, languages, frameworks and team conventions rather than one large monorepo.

Why private engineering histories matter

Public coding benchmarks are mostly built from open-source repositories, and the same repositories are scraped into pretraining corpora. Over time, scores on public benchmarks mix problem-solving with recall of fixes a model may already have seen. Histories from private repositories largely avoid that, because the code, the issues and the fixes were never published.

They also differ from open source in ways that matter for agents. Commercial codebases have deadlines, internal libraries with thin documentation, requirements written by product managers rather than maintainers, and review conventions enforced by a small, stable team. An agent that will work inside a company's repository is better tested on work that looks like it.

Turning a merged pull request into a usable task

Not every merged change makes a good task. Usable ones are self-contained, have a clear problem statement and include tests that would catch a wrong fix. Issues are often written after the fix, or describe only a symptom while the real discussion happened in chat, so the problem statement has to come from the issue and its early comments, never from the fix. Some fixes span several pull requests or arrive bundled with unrelated refactors, and need to be split or dropped.

Tests cause the most trouble. Flaky tests produce false failures, and tests that depend on network access, wall-clock time or internal services cannot run in a sealed environment. Run each candidate several times before and after the fix, and keep only tasks whose results are stable.

Keeping evaluation tasks clean

A private benchmark stays useful only while its tasks stay out of training data, including your own. Split by repository rather than by task: fixes in the same codebase share helpers, naming and even near-identical bugs, so training on one task can hand a model the answer to its neighbor. Forks, shared internal libraries and code copied between services carry the same leak across repositories, so deduplicate before splitting. Keep the gold patches, review threads and commit messages of evaluation tasks out of every training mix.

What to check before licensing

  • Confirm the partner owns the code. Agencies and consultancies often build software that belongs to their clients, and code written by contractors generally belongs to the company only if their agreements assign it.
  • Ask how secrets were found and removed. Scanning has to cover the full commit history, CI logs and attachments, because a key deleted in a later commit is still in history.
  • Check that customer data in test fixtures, database dumps, logs and issue attachments was removed, not only developer names and emails.
  • Run a software composition scan on the sample. Vendored libraries and snippets copied from public sources carry their own licenses, including copyleft, and may also sit in public training corpora.
  • Rebuild a sample of task environments yourself, and confirm the fail-to-pass tests fail before the fix and pass after it, repeatedly and offline.
  • Measure linkage on the sample, meaning the share of pull requests with a linked issue, review comments and CI results, since the links are what turn commits into tasks.
  • Agree how the code may be used, for example for training or for evaluation only, and whether tasks or results may be published.

How licensing works through SourceX

  1. 1

    Define

    Send the domain, modality, volume, format, timeline and permitted use you need.

  2. 2

    Source

    SourceX identifies businesses that hold matching data and are open to licensing it.

  3. 3

    Qualify

    Fit, rights and quality are checked, and you review samples before committing.

  4. 4

    License

    Scope, permitted use, exclusivity, price and obligations are agreed in writing.

  5. 5

    Deliver

    Approved data is prepared, de-identified where required and transferred securely.

Questions buyers ask

Can this data be used to build SWE-bench-style evaluations?

Yes, that is one of its main uses. Each task pairs an issue with the repository at its base commit, the merged fix and the tests that failed before it and passed after it. Because the repositories were never public, the tasks are far less exposed to contamination than benchmarks built from open-source projects, although vendored or copied public code can still overlap with what a model has seen.

How are secrets and credentials removed?

By scanning the full history, not just the latest snapshot. Secret scanners and custom patterns search every commit, CI log and issue attachment for API keys, tokens, private keys, passwords and connection strings, and matches are replaced with placeholders. Rewriting history changes commit hashes, so a delivery should carry a mapping. Treat any key found as live: the partner should rotate it whether or not it ships.

Are CI logs and test results included?

Where the partner's CI system kept them. Run status, job results and failing test names are usually available, while full logs may be truncated or deleted after a retention period. Logs can contain environment variables, internal hostnames and customer data, so they are scanned like code. Ask for log coverage by year if you plan to build debugging tasks from them.

Can the code be run, or is it text only?

It can be run when an environment comes with it. Executable tasks need pinned dependency versions and a container or setup script that builds and runs the tests offline. Older commits often depend on package versions or internal services that no longer exist, so environment rebuilding is scoped per order, and tasks that cannot be reproduced are flagged or dropped.

Does recent history include AI-generated code?

Possibly. Many teams have used AI coding assistants in recent years, and commits rarely record which lines a tool wrote. Some histories mark agent-authored work through bot accounts, co-author trailers or pull request labels, which can be kept as metadata. If you need human-written work, for example to avoid training on model outputs, set a date cutoff or ask for flagged commits to be excluded or labeled.

How far back does linked engineering history go?

It varies by partner and repository. Commits often go back further than anything linked to them, because issue tracking, code review and CI tend to be adopted later than version control, and moves between hosts or trackers can break links. Ask for the fully linked span separately, meaning the years in which issues, pull requests, reviews and CI results all connect.

Evaluating this data for procurement?

Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.

Request dataset diligence

Tell us what your models need

Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.

Updated 3 October 2026. Own data like this? See how companies license it to AI developers.

See if you qualify