Code and software engineering data
Code and Software Engineering Datasets for LLM Training: A Buyer's Map
Quick answer
Code datasets for LLM training come in eight main forms: whole repositories with version history, issue-to-fix tasks with tests, code review threads, CI build logs, commit rationale, legacy-language code, infrastructure and SQL code, and developer session recordings. Open corpora cover mostly public, permissively licensed code, which is also the code most likely to be in models' pre-training data already. Buyers license private code for what public code lacks: internal services, linked engineering histories and evaluation tasks no model has seen.
By SourceX Editorial · Updated
Eight kinds of code data and what each one trains
Sort code data by the engineering record it comes from, because the record decides the training unit, the use and the main risk.
| Data type | Source systems | Trains or evaluates | Main risk to check | Read next |
|---|---|---|---|---|
| Whole repositories with full history | GitHub Enterprise, GitLab, Bitbucket, Azure Repos, Perforce, Subversion | Pre-training, mid-training, repository-level completion | Third-party and copyleft files; secrets anywhere in history | Repository-level code context |
| Issue-to-fix pairs and task environments | Jira, Linear or GitHub Issues linked to pull requests | Supervised fine-tuning, RL with test-based rewards, agent evaluation | Broken issue-to-fix links; flaky or overly specific tests | Issue-to-fix pairs; SWE task environments with tests |
| Code review threads and outcomes | Pull and merge request comments, approvals, requested changes | Code review models, preference and reward data | Reviewer identities; outcome labels that differ by team | Review comment resolution pairs; preference data from review outcomes |
| CI build logs and failures | Jenkins, GitHub Actions and GitLab CI pipeline runs | Build-repair and debugging agents | Credentials and internal hostnames printed in logs | CI build failure logs |
| Commit messages and design records | Commit messages, PR descriptions, architecture decision records | Change summarization, commit message generation | Messages too thin to explain the change | Commit histories with change rationale |
| Legacy and domain-specific languages | Mainframe COBOL and JCL, SAP ABAP, Verilog and VHDL | Continued pre-training, migration and translation | Client- or vendor-owned code; few runnable tests | COBOL and mainframe code; ABAP custom code; RTL code |
| Infrastructure, pipeline and SQL code | Terraform, Kubernetes manifests, dbt models, Airflow DAGs, stored procedures | Infrastructure-as-code generation, data-engineering and SQL agents | Embedded secrets; customer values in SQL literals | Infrastructure-as-code datasets; production SQL corpora |
| Developer session recordings | IDE, terminal and browser sessions with the resulting change | Coding and computer-use agents | Third-party content and personal data on screen | Developer session recordings |
The issue-to-fix row follows the shape SWE-bench made standard: 2,294 problems drawn from real GitHub issues and the pull requests that resolved them across 12 Python repositories, with each fix checked by tests [1]. A private equivalent means the issue text, base commit, reference diff and an environment where the tests run, as listed on the software engineering histories page.
Two rows are thin in public data. SQL quality depends on real schemas: Spider 2.0 builds its tasks on enterprise databases that often have more than 1,000 columns, on systems such as BigQuery and Snowflake, and some solutions exceed 100 lines of SQL [2]; question-to-SQL pairs are covered in text-to-SQL training data. Public log collections such as Loghub gather runtime logs from distributed systems, supercomputers, operating systems and other software [3], not build-and-test runs tied to the commit that fixed them.
Why public repositories no longer settle the question
Public code is cheap for pre-training but unreliable for evaluation and thin on the work coding agents do. The SWE-Bench Pro authors argue that widely used open-source repositories, especially permissively licensed ones, are prime candidates for the web-crawled corpora used in pre-training, so benchmarks built from public GitHub repositories are hard to keep clean [4]. They built public and held-out sets only from repositories under strong copyleft (GPL) licenses and added a commercial set from codebases acquired from startups, publishing results while that code stays private [4].
In a post dated 23 February 2026, OpenAI said it had stopped reporting SWE-bench Verified because score gains increasingly reflected exposure to the benchmark during training; it recommends SWE-bench Pro for now and calls Pro's contamination considerably lower but not perfect [5]. The LiveBench authors cite evidence that model performance on Codeforces problems drops sharply for problems released after a model's training cutoff [6].
Open baselines exist. The Stack collects 3.1 TB of permissively licensed source code in 30 programming languages, filtered by repository license, with an "Am I in The Stack" search and a removal process [7]. The Stack v2, built from the Software Heritage archive, is distributed as a gated Hugging Face dataset; check its dataset card for access terms and download steps [8]. The Common Pile v0.1 includes code among the 30 openly licensed and public-domain sources in its 8 TB collection [9].
Developer community content is gated too: Stack Exchange posts are licensed CC BY-SA, yet in 2024 the company put its data dump behind a login and an agreement not to use the content to train AI models [10]. Public license labels are weak evidence: the Data Provenance Initiative found licenses omitted for more than 70% of popular datasets on hosting sites and error rates above 50% [11].
| Question | Open code corpora | Licensed private code |
|---|---|---|
| Rights basis | Each file's open-source license, plus dataset terms and opt-outs | The code owner's license defining records, uses and term |
| Exposure to existing models | High: widely used public repositories are prime candidates for pre-training crawls [4] | Low when no public mirror or fork exists; verify it |
| What it covers | Public libraries, frameworks and popular languages | Internal services, legacy estates, linked issue, review and CI history |
| Fit for held-out evaluation | Weak | Strong when access stays controlled |
Synthetic code, the third route, carries its provider's output terms; as of October 2026, Anthropic's help center, for example, says its terms do not allow outputs to train models that compete with its own [12]. Compare the routes in open code datasets vs licensed private code and synthetic vs licensed real code.
What each training stage needs from a code record
Specify the training stage first: pre-training buys deduplicated tokens, fine-tuning buys linked task pairs, reinforcement learning buys runnable environments, and evaluation buys tasks no model has seen.
| Stage | Unit to specify | Acceptance test before you pay |
|---|---|---|
| Pre-training and mid-training | Tokens by language after exact and near-duplicate removal; a file-level license inventory | Overlap with public code corpora; share of vendored, generated and minified files |
| Supervised fine-tuning | Issue to diff, review comment to resolving change, change to commit message | Every link resolves; each diff applies cleanly to its base commit |
| RL with test-based rewards | Tasks with a base commit, fail-to-pass and pass-to-pass tests, and a container image | Tests fail before and pass after the reference patch on repeated offline runs |
| Coding-agent evaluation | Held-out tasks from repositories with no public copy | Searches for public mirrors and forks; contract terms that keep tasks out of training mixes |
| Code review and security models | Review threads with outcomes; fixes for long-patched vulnerabilities | Outcome labels defined the same way across teams; no unpatched vulnerability details |
Test quality decides the value of RL and evaluation tasks. When it introduced SWE-bench Verified, a 500-task human-validated subset built with the SWE-bench authors, OpenAI cited overly specific unit tests, underspecified problem descriptions and unreliable environment setup in the original tasks, and shipped a Docker-based harness [13]. Ask for the same screening on private tasks; see held-out evaluation sets from private repositories and the LLM evaluation datasets hub.
Forks, vendored dependencies and copied snippets make code heavily duplicated. In natural-language corpora, Lee et al. found that deduplicated training data cut memorized output about tenfold [14]; see deduplicating code training data and data sourcing for post-training teams.
Seven risks that separate licensable code from risky code
Private code is worth licensing only if the supplier can show it owns the code and has removed what it may not share, and the buyer can keep it from resurfacing verbatim.
| Risk | What goes wrong | Evidence to ask for | Guide |
|---|---|---|---|
| Ownership | Contractor code without an IP assignment, acquired code without transferred rights, agency code owned by clients | Origin per repository (in-house, contractor, acquired, client) and the document behind it | Code ownership due diligence; chain of title |
| Copyleft and third-party code | GPL or AGPL files, vendored libraries and pasted snippets under other terms | Per-file license findings with SPDX identifiers and the action taken on each | Copyleft contamination; SPDX-based dataset bills of materials |
| Secrets | Keys deleted from the current tree remain in older commits, branches, tags and CI logs | Scan scope covering every commit and ref, findings by category, replacement method | Secrets in code datasets |
| Personal data | Author names and emails in commit metadata; customer records in test fixtures and comments | De-identification method and a checked sample | De-identifying source code; employee-authored records |
| Export controls | Some code, such as cryptographic code, may be subject to export control regulations that restrict its transfer or release | A per-repository export screen, especially for non-US buyers | Export controls on source code |
| Contamination | A "private" repository has a public mirror, fork or leaked copy | Searches against public code hosts and an overlap report | Code benchmark contamination |
| Regurgitation | The trained model reproduces licensed code verbatim, exposing the supplier | Deduplication report, training-run controls, output filtering | Preventing verbatim reproduction |
Verbatim extraction is documented: Carlini et al. recovered hundreds of verbatim training sequences from GPT-2, including code and UUIDs [15].
N-gram overlap is the most common decontamination technique, with thresholds that vary by lab, such as 13-gram matches or 50-character overlaps [16]. The LMSYS team showed that rephrased test items slip past n-gram checks, and its LLM-based decontaminator found benchmark overlap in corpora including The Stack [17].
Keep provenance per repository, because model developers' disclosure duties depend on knowing where training data came from. As of October 2026, under the EU AI Act, providers of general-purpose AI models must publish a summary of their training content; the European Commission published a template for this on 24 July 2025, which calls for listing main data collections and other sources [18]. California's AB 2013 has required developers of generative AI systems made available to Californians to post training-data documentation since 1 January 2026, including whether datasets contain copyrighted or licensed material or personal information [19].
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Who holds private code, and who must approve its release
Private code sits with three kinds of holders, and each has a different approval chain. Software product companies usually own their code but embed third-party components. In-house teams at non-software businesses hold internal tools and legacy estates; customizations of a vendor platform need a check of the vendor's terms. Agencies and contractors often build code their clients own, so the client's authorization comes first.
SourceX sources operational datasets from US companies, including engineering records, and manages the commercial process, from the licensing agreement to ongoing purchases. Buyers describe the data rather than the businesses: SourceX looks for US companies that hold it, and every release is approved by the supplying company. Datasets are sourced on request, not held in stock, and a request does not guarantee a matching dataset.
Every dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and is delivered under a license that defines the records, permitted uses, term and delivery. Personal details such as names and emails are removed or replaced before delivery, with the method recorded and a processed sample checked, though no de-identification method is perfect. SourceX does not source scraped public web content, does not train AI models and does not publish prices. You can describe the private code or engineering records you need.
Record-level detail is on the pages for proprietary codebases with full git history, coding agent training data, licensing source code, licensing code reviews and pull requests and whether AI labs buy code.
Mistakes code data buyers make
These mistakes come from measuring the wrong unit, accepting the wrong shape of data or trusting an unchecked label.
- Pricing by lines of code. Counts that include vendored, generated and minified files inflate volume; see what drives the price of licensed code data.
- Accepting a snapshot when you need history. A tree of the latest files has no commits, no issue links and no base commits for tasks; see delivering repositories with full history.
- Trusting the repository LICENSE file for every file. Copied snippets and vendored directories carry their own terms, and dataset-level labels are often missing or wrong [11].
- Comparing scores across benchmark versions. Task sets and harnesses differ between SWE-bench, SWE-bench Verified, and SWE-bench Pro, so scores do not transfer. OpenAI stopped reporting on SWE-bench Verified due to flawed tests and contamination, recommending SWE-bench Pro instead [13][5].
Start here: code data guides by buying step
The guides below follow the order of a purchase, from writing the request to settling rights. For other data types, start from the SourceX guide to AI data.
| Buying step | Guide |
|---|---|
| Write the request | How to specify a code dataset request |
| Check data before licensing | Evaluating a code dataset before you license it |
| Translate legacy systems | Legacy code translation pairs |
| Train agents on work beyond code | AI agent training data hub |
| Settle rights and terms | AI training data licensing hub |
| Plan post-training mixes | Fine-tuning datasets hub |
Need private code or engineering histories for your model?
Describe the languages, repository types, history depth, linked records (issues, reviews, CI runs), test requirements, volume and permitted uses you need. SourceX looks for US companies that hold that code, checks their licensing permissions, and manages the license and delivery. Send SourceX your code dataset specification.
Guides in this section
- CI Build Failure Datasets: Logs, Fixes and Flaky TestsSource CI build failure datasets for coding agents: failing runs linked to logs, root cause, fixing commits and flaky-test reruns, plus a spec checklist.
- COBOL Datasets for LLM Training: Complete Mainframe EstatesWhat a usable COBOL training dataset must contain: copybooks, JCL, CICS maps, DB2 DDL, EBCDIC handling, ownership checks and realistic evaluation.
- COBOL-to-Java Translation Datasets from Real MigrationsHow to source aligned COBOL-to-Java and legacy migration code pairs with equivalence evidence for training and evaluating code translation models.
- Code Benchmark Contamination: Detecting Repo and Fork LeaksHow coding benchmarks such as HumanEval and SWE-bench leak via forks, mirrors and rephrasings, and which checks catch it in training and eval data.
- Code Dataset Request Spec: Languages, History, Builds, TestsA field-by-field template for code data requests: languages, versions, VCS history depth, issue and PR linkage, buildability, tests and acceptance metrics.
- Copyleft Code in AI Training Data: GPL, AGPL, SnippetsHow GPL, AGPL and copied snippets hide in private codebases, and the inventory, SPDX tagging and disposition to require before licensing code data.
- Developer Session Recordings for Coding Agent TrainingHow to specify human coding agent trajectories: IDE events, terminal I/O, screen and narration, linked to diffs and tests, plus consent and redaction.
- Evaluating a Code Dataset Sample: Build, Test, Dedup ChecksHow to evaluate a code dataset sample before licensing: build and test rates, validated tasks, dedup against public code, secrets and license scans.
- Issue-to-Fix Pairs: Building and Validating TasksHow to build issue-to-fix pairs from private repositories: task anatomy, filters, fail-to-pass test validation, yield funnels and leakage controls.
- Private SWE Benchmark Datasets from Licensed RepositoriesHow to source a held-out coding agent eval set from never-public repositories: repository mix, split policy, task counts, access isolation and refresh.
- Removing Secrets from Code Training Data: Buyer ChecksHow buyers verify a licensed code dataset is free of secrets: full git history scope, detectors, typed placeholders, hash maps and acceptance rescans.
- Source Code Ownership Due Diligence for AI Data DealsWho owns each part of a private codebase? Check employee and contractor assignments, client work, acquired code and inbound OSS before licensing it for AI.
- SWE Task Environments with Tests: The Runtime ContractWhat to require from an executable coding task dataset: pinned images, lockfiles, offline builds, deterministic tests, flake reports and a task manifest.
- ABAP and SAP Custom Code Datasets for LLM TrainingHow to source licensed ABAP custom code for LLM training: customer namespace scope, transport history, ATC remediation pairs and DDIC without table data.
- Code Deduplication for LLM Training: Forks and Near-DupesHow to dedupe code training data: forks, vendored copies, generated files, MinHash on normalized tokens, and measuring net-new tokens in licensed code.
- Code Preference Data from Review Outcomes and RevertsBuild code preference pairs and reward model data from real review outcomes: merged vs first revisions, rejected PRs, reverts and incident fixes.
- Code Review Comment Datasets: Comment-to-Fix Pairs for AIHow to source code review comment data: line-anchored comments, addressing commits, resolution labels and ignored-comment negatives for AI code reviewers.
- Commit Message Datasets with Change Rationale for AIHow to source commit message datasets where the why is recorded: PR descriptions, linked tickets and ADRs, informativeness filters and redaction checks.
- Data Pipeline Code Datasets: dbt, Airflow and Spark for AIWhat to require in a data engineering code dataset: whole dbt projects, Airflow DAGs and Spark jobs with manifests, lineage, tests and failure-fix history.
- Dependency Upgrade Datasets for Migration Coding AgentsHow to source dependency upgrade and API migration data: breaking-update commits with lockfile diffs, human fixes and before/after CI for training agents.
- Export Controls on Source Code for AI Training: Buyer ChecksHow EAR encryption rules, ITAR technical data and deemed exports apply to licensed source code, and what screening non-US AI buyers should expect.
- Infrastructure-as-Code Datasets for AI: Terraform to CISource Terraform, Kubernetes, Helm, Ansible and CI pipeline code paired with plan output, policy findings and fixes to train and evaluate IaC agents.
- Jupyter Notebook Datasets for Data Science AgentsHow to source enterprise Jupyter notebook datasets with code, outputs and narrative for training and evaluating data science agents, and what to check.
- Open Code Datasets vs Licensed Private Code for LLM TrainingCompare openly licensed code corpora with licensed private repositories: license audits, attribution and opt-out duties, coverage gaps and contamination.
- Preventing LLMs from Reproducing Licensed Training CodeHow code models memorize licensed proprietary code, how to measure verbatim extraction, and the dedup, filter and contract controls suppliers expect.
- Production SQL Query Corpora and Stored Procedures for AIHow to source real SQL query logs, stored procedures and schema DDL for LLM training: fields to require, masking literals, dedup and execution stats.
- Repository-Level Code Data: Cross-File Context GuideHow to package repository-level code data: dependency-ordered packing, build metadata, symbol indexes, monorepo sampling and leak-free cross-file evals.
- SQL Dialect Translation Datasets from Real MigrationsHow to source SQL dialect translation data: PL/SQL and T-SQL conversion pairs, DDL context, and result-reconciliation labels for training and evaluation.
- Stack Trace to Fix Datasets for Debugging AgentsHow to source production stack traces linked to the commits that fixed them: release tags, grouping, scrubbing, link rates and eval splits for agents.
- Synthetic Code Data vs Licensed Real Code for LLM TrainingWhen model-generated code is enough for SFT, RL tasks and evals, and when licensed real code is needed: output terms, contamination and coverage.
- Unit Test Generation Data: Focal Code, Tests, CoverageHow to source unit test generation data: focal method to test mappings, per-test coverage, mutation scores, flakiness flags and fixture de-identification.
- Verilog, SystemVerilog and VHDL Datasets for LLM TrainingWhat to require in licensed RTL data for chip design AI: testbenches, assertions, coverage, lint and synthesis reports, EDA Tcl scripts and IP exclusions.
- Vulnerability Fix Datasets from Private Code: Buyer GuideHow to source vulnerable-then-fixed code pairs from private repositories, with CWE labels, finding sources and verification tests for secure code models.
- What Drives the Price of Licensed Proprietary Code DataWhy licensed code data prices vary: net-new tokens after dedup, history linkage, buildability, test strength, validated tasks, use rights and exclusivity.
Sources
- Jimenez et al. (Princeton, UChicago; ICLR 2024; arXiv:2310.06770), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
- arXiv (arXiv:2411.07763; ICLR 2025), "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows" (2024). https://www.arxiv.org/pdf/2411.07763
- arXiv (arXiv:2008.06448), "Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics" (2020). https://arxiv.org/pdf/2008.06448
- arXiv (arXiv:2509.16941), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- arXiv (arXiv:2406.19314; ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
- Kocetkov et al. (BigCode; arXiv:2211.15533), "The Stack: 3 TB of permissively licensed source code" (2022). https://export.arxiv.org/abs/2211.15533?context=cs
- BigCode (Hugging Face dataset repository), "bigcode/the-stack-v2-dedup". https://huggingface.co/datasets/bigcode/the-stack-v2-dedup
- arXiv (arXiv:2506.05209), "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
- DevClass, "Stack Exchange restricts access to dump of user-contributed data as critics complain license permits reuse for any purpose" (2024). https://devclass.com/2024/07/30/stack-exchange-restricts-access-to-dump-of-user-contributed-data-as-critics-complain-license-permits-reuse-for-any-purpose
- Longpre et al. (arXiv:2310.16787; journal version: Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Anthropic (Claude Help Center; vendor page, evidence of one provider's terms), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- OpenAI (with the SWE-bench authors), "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
- Lee et al. (ACL 2022; arXiv:2107.06499), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- Carlini et al. (USENIX Security 2021), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
- arXiv (arXiv:2406.04244), "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
- LMSYS Org, "LLM Decontaminator (blog post, 14 November 2023)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.