Code and software engineering data
Dependency Upgrade and API Migration Data for Coding Agents
Quick answer
A useful dependency upgrade dataset is not a pile of version bumps. It is a set of task units where a manifest or lockfile change broke the build or tests, an engineer changed application code to fix it, and CI results exist before and after. Buy those breaking-update-plus-fix units, separate them from bot bumps that merged untouched, insist that the old dependency versions can still be resolved, and prioritize internal library migrations that public repositories cannot supply.
By SourceX Editorial · Updated
Why public upgrade data runs out quickly
Public sources prove the task is learnable but leave little in-house migration data. BUMP, a widely used benchmark, collects reproducible breaking dependency updates in Java projects, and on it the strongest model in the Byam study fully repaired only 27% of the broken builds [1]. That tells you how much headroom remains, and how narrow the public base is.
Breadth is the other gap. VersiCode spans 300 libraries and 2,207 versions and notes that most code datasets ignore library versioning altogether [2]. None of these corpora contain your customers' reality: internal SDKs, in-house frameworks, private forks of open-source packages, and the multi-month platform migrations that enterprise teams actually run. That is the data worth licensing from operating companies' private repositories, typically alongside broader software engineering histories.
The task unit: manifest diff, code fix, CI before and after
Each record should bundle three things, because a model needs all three to learn or be scored. First, the dependency change itself: the diff to package.json and package-lock.json, pom.xml, build.gradle, go.mod and go.sum, Cargo.toml and Cargo.lock, pyproject.toml with poetry.lock or uv.lock, or a Gemfile.lock. Second, the human code changes that made the build pass again. Third, CI evidence: the failing run on the bump commit and the passing run on the fix.
Diff context matters for training, not only for evaluation. A recent code migration study found that supplying diff context significantly improved LLM migration performance [3]. Ask for the upstream changelog or release-notes reference when the supplier's commit or pull request links one, since it is the closest thing to the "documentation an agent would read" for that API deprecation fix.
Illustrative example: invented to show structure; it does not describe an available dataset.
{"task_id": "upg-000412", "ecosystem": "maven", "dependency": "com.example.internal:payments-sdk", "internal_library": true, "from_version": "3.8.2", "to_version": "4.0.0", "semver_jump": "major", "bump_commit": "a1f3c9e", "fix_commits": ["b72d004", "c0e9a11"], "bump_author_type": "bot", "fix_author_type": "human", "files_changed_in_fix": 14, "ci_before": {"status": "failed", "stage": "compile", "log_ref": "logs/upg-000412/bump.txt"}, "ci_after": {"status": "passed", "tests_run": 1882, "tests_failed": 0}, "failure_class": "removed_method", "changelog_ref": "CHANGELOG.md#4.0.0", "registry_mirrored": true, "container_image": "upg-000412:pre"}
Deliver records like this as JSON Lines, one UTF-8 JSON object per line [5], with logs, diffs and repository snapshots referenced by path rather than inlined.
Separating bot bumps from breaking upgrades
The most valuable records are upgrades that broke something and needed human code edits, and most raw upgrade commits are neither. Tools such as Dependabot and Renovate open version-bump pull requests automatically on a schedule, and on a long-neglected repository they can produce a steady stream of them. Most of these patch and minor bumps merge with no code change and teach a model nothing beyond editing a version string.
Filter on observable signals rather than commit-message keywords:
- Author type. Bot author on the bump, at least one human-authored commit before merge.
- Code delta. Changes to source files outside manifests and lockfiles, excluding generated code and vendored directories.
- CI transition. A red run on the bump commit followed by green on the fix, from the same pipeline definition.
- Version distance. Under Semantic Versioning, a major increment signals incompatible API changes [4], so weight major bumps, but keep minor bumps that broke builds, because those are undeclared breaking changes and hard cases.
- Failure class. Compile errors, test failures and runtime configuration errors behave differently, so check that your sample is not all one class.
Keep a small share of clean bumps as negatives if you are training a model to decide when an upgrade is safe to merge as is.
Reproducibility: the registry problem
A migration task you cannot rebuild is only text. Upgrades reference versions that may since have been yanked or unpublished from PyPI or npm, superseded internally, or hosted on a private Artifactory, Nexus or GitHub Packages registry the buyer cannot reach. Bot configurations such as dependabot.yml often reference those private registries, a reminder that the dependency graph lives partly outside the repository.
For RL with build and test rewards, or for a held-out migration eval, ask suppliers how they will make both the old and new versions resolvable: a mirrored snapshot of the required internal packages, prebuilt container images per task, as Java breaking-update benchmarks do [1], or a pinned offline cache. Specify the toolchain versions too (JDK, Node, Python, Go), because a fix that compiles only on a newer runtime will confuse the reward. Our guide to SWE task environments with tests covers harness requirements in more depth.
Large codemod migrations and internal frameworks
Internal library migrations are the scarce part of this category, so ask suppliers to tag them explicitly. A field such as internal_library: true plus the owning team's migration guide or RFC turns a commit into a task an agent can be instructed on. Framework upgrades (Spring Boot 2 to 3, Angular major versions, Django LTS jumps, React class-to-hooks rewrites) often sit between public and internal: the framework is public but the call sites and wrappers are private.
Codemod-driven migrations touching thousands of files need to be split before they are useful. A single 4,000-file pull request is neither a reasonable SFT target nor a fair eval item. Ask whether the supplier can segment by module or package, keep the codemod script as a separate artifact, and isolate the hand-written follow-up commits, which are where the real reasoning lives. For language-level ports rather than library upgrades, see legacy code translation pairs and SQL dialect migration data.
Buyer checklist for a dependency upgrade data request
Use this checklist when you write the request and when you evaluate a sample:
| Requirement | What to ask for | Why it matters |
|---|---|---|
| Task unit | Manifest/lockfile diff, human fix commits, CI before and after | Missing any one breaks SFT targets or rewards |
| Filtering | Bot-only clean bumps excluded or labeled | Avoids a dataset of version-string edits |
| Ecosystems | Named package managers and build tools | Lockfile formats and failure modes differ |
| Internal tagging | Internal library and in-house framework flag | The data public code lacks |
| Reproducibility | Mirrored registries, containers or offline cache | Tasks must build for RL and eval |
| CI logs | Full failing logs, not just status | The error is the agent's main input |
| Splitting | Large codemods segmented by module | Reviewable, fair task sizes |
| Hygiene | Secrets scan over full history; license review of vendored code | Credentials and copyleft travel with lockfiles and vendored code |
| Held-out set | Repositories reserved from training splits | Prevents leakage into your eval |
Pair it with evaluating a code dataset sample, and with the CI build failure logs guide if you need failing logs beyond upgrade events.
How SourceX handles upgrade and migration data requests
SourceX sources operational datasets, including engineering records, from US companies and manages the commercial process through licensing and ongoing purchases. Data is sourced on request rather than held in stock, so a request does not guarantee a match. You describe the data you need, and SourceX looks for US businesses that hold it; every release is approved by the supplying company.
The process runs Find, Assess (the data and its licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails and account numbers are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. Delivery runs through private, access-controlled workflows after an executed agreement. See proprietary code datasets with full Git history, the code data buyer's map, or describe your upgrade data request.
Request dependency upgrade and migration data
If your agent needs breaking-update fixes, internal library migrations and before/after CI from private repositories, describe the ecosystems, task unit and reproducibility needs you require. SourceX looks for US companies that hold matching engineering records, serves AI teams wherever they are based, and manages licensing from assessment through ongoing purchases. Start a buyer request at sourcex.si/buyers.
Sources
- arXiv, "Byam: Fixing Breaking Dependency Updates with Large Language Models" (2025). https://arxiv.org/pdf/2505.07522v3
- arXiv, "VersiCode: Towards Version-controllable Code Generation" (2024). https://arxiv.org/html/2406.07411v1
- arXiv, "What a diff makes: automating code migration with large language models" (2025). https://arxiv.org/html/2511.00160v1
- semver.org, "Semantic Versioning 2.0.0". https://semver.org/
- jsonlines.org, "JSON Lines". https://jsonlines.org/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.