Skip to content

Code and software engineering data

Code and Software Engineering Datasets for LLM Training: A Buyer's Map

Quick answer

Code datasets for LLM training come in eight main forms: whole repositories with version history, issue-to-fix tasks with tests, code review threads, CI build logs, commit rationale, legacy-language code, infrastructure and SQL code, and developer session recordings. Open corpora cover mostly public, permissively licensed code, which is also the code most likely to be in models' pre-training data already. Buyers license private code for what public code lacks: internal services, linked engineering histories and evaluation tasks no model has seen.

By SourceX Editorial · Updated

Eight kinds of code data and what each one trains

Sort code data by the engineering record it comes from, because the record decides the training unit, the use and the main risk.

Data typeSource systemsTrains or evaluatesMain risk to checkRead next
Whole repositories with full historyGitHub Enterprise, GitLab, Bitbucket, Azure Repos, Perforce, SubversionPre-training, mid-training, repository-level completionThird-party and copyleft files; secrets anywhere in historyRepository-level code context
Issue-to-fix pairs and task environmentsJira, Linear or GitHub Issues linked to pull requestsSupervised fine-tuning, RL with test-based rewards, agent evaluationBroken issue-to-fix links; flaky or overly specific testsIssue-to-fix pairs; SWE task environments with tests
Code review threads and outcomesPull and merge request comments, approvals, requested changesCode review models, preference and reward dataReviewer identities; outcome labels that differ by teamReview comment resolution pairs; preference data from review outcomes
CI build logs and failuresJenkins, GitHub Actions and GitLab CI pipeline runsBuild-repair and debugging agentsCredentials and internal hostnames printed in logsCI build failure logs
Commit messages and design recordsCommit messages, PR descriptions, architecture decision recordsChange summarization, commit message generationMessages too thin to explain the changeCommit histories with change rationale
Legacy and domain-specific languagesMainframe COBOL and JCL, SAP ABAP, Verilog and VHDLContinued pre-training, migration and translationClient- or vendor-owned code; few runnable testsCOBOL and mainframe code; ABAP custom code; RTL code
Infrastructure, pipeline and SQL codeTerraform, Kubernetes manifests, dbt models, Airflow DAGs, stored proceduresInfrastructure-as-code generation, data-engineering and SQL agentsEmbedded secrets; customer values in SQL literalsInfrastructure-as-code datasets; production SQL corpora
Developer session recordingsIDE, terminal and browser sessions with the resulting changeCoding and computer-use agentsThird-party content and personal data on screenDeveloper session recordings

The issue-to-fix row follows the shape SWE-bench made standard: 2,294 problems drawn from real GitHub issues and the pull requests that resolved them across 12 Python repositories, with each fix checked by tests [1]. A private equivalent means the issue text, base commit, reference diff and an environment where the tests run, as listed on the software engineering histories page.

Two rows are thin in public data. SQL quality depends on real schemas: Spider 2.0 builds its tasks on enterprise databases that often have more than 1,000 columns, on systems such as BigQuery and Snowflake, and some solutions exceed 100 lines of SQL [2]; question-to-SQL pairs are covered in text-to-SQL training data. Public log collections such as Loghub gather runtime logs from distributed systems, supercomputers, operating systems and other software [3], not build-and-test runs tied to the commit that fixed them.

Why public repositories no longer settle the question

Public code is cheap for pre-training but unreliable for evaluation and thin on the work coding agents do. The SWE-Bench Pro authors argue that widely used open-source repositories, especially permissively licensed ones, are prime candidates for the web-crawled corpora used in pre-training, so benchmarks built from public GitHub repositories are hard to keep clean [4]. They built public and held-out sets only from repositories under strong copyleft (GPL) licenses and added a commercial set from codebases acquired from startups, publishing results while that code stays private [4].

In a post dated 23 February 2026, OpenAI said it had stopped reporting SWE-bench Verified because score gains increasingly reflected exposure to the benchmark during training; it recommends SWE-bench Pro for now and calls Pro's contamination considerably lower but not perfect [5]. The LiveBench authors cite evidence that model performance on Codeforces problems drops sharply for problems released after a model's training cutoff [6].

Open baselines exist. The Stack collects 3.1 TB of permissively licensed source code in 30 programming languages, filtered by repository license, with an "Am I in The Stack" search and a removal process [7]. The Stack v2, built from the Software Heritage archive, is distributed as a gated Hugging Face dataset; check its dataset card for access terms and download steps [8]. The Common Pile v0.1 includes code among the 30 openly licensed and public-domain sources in its 8 TB collection [9].

Developer community content is gated too: Stack Exchange posts are licensed CC BY-SA, yet in 2024 the company put its data dump behind a login and an agreement not to use the content to train AI models [10]. Public license labels are weak evidence: the Data Provenance Initiative found licenses omitted for more than 70% of popular datasets on hosting sites and error rates above 50% [11].

QuestionOpen code corporaLicensed private code
Rights basisEach file's open-source license, plus dataset terms and opt-outsThe code owner's license defining records, uses and term
Exposure to existing modelsHigh: widely used public repositories are prime candidates for pre-training crawls [4]Low when no public mirror or fork exists; verify it
What it coversPublic libraries, frameworks and popular languagesInternal services, legacy estates, linked issue, review and CI history
Fit for held-out evaluationWeakStrong when access stays controlled

Synthetic code, the third route, carries its provider's output terms; as of October 2026, Anthropic's help center, for example, says its terms do not allow outputs to train models that compete with its own [12]. Compare the routes in open code datasets vs licensed private code and synthetic vs licensed real code.

What each training stage needs from a code record

Specify the training stage first: pre-training buys deduplicated tokens, fine-tuning buys linked task pairs, reinforcement learning buys runnable environments, and evaluation buys tasks no model has seen.

StageUnit to specifyAcceptance test before you pay
Pre-training and mid-trainingTokens by language after exact and near-duplicate removal; a file-level license inventoryOverlap with public code corpora; share of vendored, generated and minified files
Supervised fine-tuningIssue to diff, review comment to resolving change, change to commit messageEvery link resolves; each diff applies cleanly to its base commit
RL with test-based rewardsTasks with a base commit, fail-to-pass and pass-to-pass tests, and a container imageTests fail before and pass after the reference patch on repeated offline runs
Coding-agent evaluationHeld-out tasks from repositories with no public copySearches for public mirrors and forks; contract terms that keep tasks out of training mixes
Code review and security modelsReview threads with outcomes; fixes for long-patched vulnerabilitiesOutcome labels defined the same way across teams; no unpatched vulnerability details

Test quality decides the value of RL and evaluation tasks. When it introduced SWE-bench Verified, a 500-task human-validated subset built with the SWE-bench authors, OpenAI cited overly specific unit tests, underspecified problem descriptions and unreliable environment setup in the original tasks, and shipped a Docker-based harness [13]. Ask for the same screening on private tasks; see held-out evaluation sets from private repositories and the LLM evaluation datasets hub.

Forks, vendored dependencies and copied snippets make code heavily duplicated. In natural-language corpora, Lee et al. found that deduplicated training data cut memorized output about tenfold [14]; see deduplicating code training data and data sourcing for post-training teams.

Seven risks that separate licensable code from risky code

Private code is worth licensing only if the supplier can show it owns the code and has removed what it may not share, and the buyer can keep it from resurfacing verbatim.

RiskWhat goes wrongEvidence to ask forGuide
OwnershipContractor code without an IP assignment, acquired code without transferred rights, agency code owned by clientsOrigin per repository (in-house, contractor, acquired, client) and the document behind itCode ownership due diligence; chain of title
Copyleft and third-party codeGPL or AGPL files, vendored libraries and pasted snippets under other termsPer-file license findings with SPDX identifiers and the action taken on eachCopyleft contamination; SPDX-based dataset bills of materials
SecretsKeys deleted from the current tree remain in older commits, branches, tags and CI logsScan scope covering every commit and ref, findings by category, replacement methodSecrets in code datasets
Personal dataAuthor names and emails in commit metadata; customer records in test fixtures and commentsDe-identification method and a checked sampleDe-identifying source code; employee-authored records
Export controlsSome code, such as cryptographic code, may be subject to export control regulations that restrict its transfer or releaseA per-repository export screen, especially for non-US buyersExport controls on source code
ContaminationA "private" repository has a public mirror, fork or leaked copySearches against public code hosts and an overlap reportCode benchmark contamination
RegurgitationThe trained model reproduces licensed code verbatim, exposing the supplierDeduplication report, training-run controls, output filteringPreventing verbatim reproduction

Verbatim extraction is documented: Carlini et al. recovered hundreds of verbatim training sequences from GPT-2, including code and UUIDs [15].

N-gram overlap is the most common decontamination technique, with thresholds that vary by lab, such as 13-gram matches or 50-character overlaps [16]. The LMSYS team showed that rephrased test items slip past n-gram checks, and its LLM-based decontaminator found benchmark overlap in corpora including The Stack [17].

Keep provenance per repository, because model developers' disclosure duties depend on knowing where training data came from. As of October 2026, under the EU AI Act, providers of general-purpose AI models must publish a summary of their training content; the European Commission published a template for this on 24 July 2025, which calls for listing main data collections and other sources [18]. California's AB 2013 has required developers of generative AI systems made available to Californians to post training-data documentation since 1 January 2026, including whether datasets contain copyrighted or licensed material or personal information [19].

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Who holds private code, and who must approve its release

Private code sits with three kinds of holders, and each has a different approval chain. Software product companies usually own their code but embed third-party components. In-house teams at non-software businesses hold internal tools and legacy estates; customizations of a vendor platform need a check of the vendor's terms. Agencies and contractors often build code their clients own, so the client's authorization comes first.

SourceX sources operational datasets from US companies, including engineering records, and manages the commercial process, from the licensing agreement to ongoing purchases. Buyers describe the data rather than the businesses: SourceX looks for US companies that hold it, and every release is approved by the supplying company. Datasets are sourced on request, not held in stock, and a request does not guarantee a matching dataset.

Every dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and is delivered under a license that defines the records, permitted uses, term and delivery. Personal details such as names and emails are removed or replaced before delivery, with the method recorded and a processed sample checked, though no de-identification method is perfect. SourceX does not source scraped public web content, does not train AI models and does not publish prices. You can describe the private code or engineering records you need.

Record-level detail is on the pages for proprietary codebases with full git history, coding agent training data, licensing source code, licensing code reviews and pull requests and whether AI labs buy code.

Mistakes code data buyers make

These mistakes come from measuring the wrong unit, accepting the wrong shape of data or trusting an unchecked label.

  • Pricing by lines of code. Counts that include vendored, generated and minified files inflate volume; see what drives the price of licensed code data.
  • Accepting a snapshot when you need history. A tree of the latest files has no commits, no issue links and no base commits for tasks; see delivering repositories with full history.
  • Trusting the repository LICENSE file for every file. Copied snippets and vendored directories carry their own terms, and dataset-level labels are often missing or wrong [11].
  • Comparing scores across benchmark versions. Task sets and harnesses differ between SWE-bench, SWE-bench Verified, and SWE-bench Pro, so scores do not transfer. OpenAI stopped reporting on SWE-bench Verified due to flawed tests and contamination, recommending SWE-bench Pro instead [13][5].

Start here: code data guides by buying step

The guides below follow the order of a purchase, from writing the request to settling rights. For other data types, start from the SourceX guide to AI data.

Buying stepGuide
Write the requestHow to specify a code dataset request
Check data before licensingEvaluating a code dataset before you license it
Translate legacy systemsLegacy code translation pairs
Train agents on work beyond codeAI agent training data hub
Settle rights and termsAI training data licensing hub
Plan post-training mixesFine-tuning datasets hub

Need private code or engineering histories for your model?

Describe the languages, repository types, history depth, linked records (issues, reviews, CI runs), test requirements, volume and permitted uses you need. SourceX looks for US companies that hold that code, checks their licensing permissions, and manages the license and delivery. Send SourceX your code dataset specification.

Guides in this section

Sources

  1. Jimenez et al. (Princeton, UChicago; ICLR 2024; arXiv:2310.06770), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
  2. arXiv (arXiv:2411.07763; ICLR 2025), "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows" (2024). https://www.arxiv.org/pdf/2411.07763
  3. arXiv (arXiv:2008.06448), "Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics" (2020). https://arxiv.org/pdf/2008.06448
  4. arXiv (arXiv:2509.16941), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/pdf/2509.16941
  5. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  6. arXiv (arXiv:2406.19314; ICLR 2025), "LiveBench: A Challenging, Contamination-Limited LLM Benchmark" (2024). https://www.arxiv.org/pdf/2406.19314
  7. Kocetkov et al. (BigCode; arXiv:2211.15533), "The Stack: 3 TB of permissively licensed source code" (2022). https://export.arxiv.org/abs/2211.15533?context=cs
  8. BigCode (Hugging Face dataset repository), "bigcode/the-stack-v2-dedup". https://huggingface.co/datasets/bigcode/the-stack-v2-dedup
  9. arXiv (arXiv:2506.05209), "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
  10. DevClass, "Stack Exchange restricts access to dump of user-contributed data as critics complain license permits reuse for any purpose" (2024). https://devclass.com/2024/07/30/stack-exchange-restricts-access-to-dump-of-user-contributed-data-as-critics-complain-license-permits-reuse-for-any-purpose
  11. Longpre et al. (arXiv:2310.16787; journal version: Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  12. Anthropic (Claude Help Center; vendor page, evidence of one provider's terms), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
  13. OpenAI (with the SWE-bench authors), "Introducing SWE-bench Verified" (2024). https://openai.com/index/introducing-swe-bench-verified/
  14. Lee et al. (ACL 2022; arXiv:2107.06499), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  15. Carlini et al. (USENIX Security 2021), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  16. arXiv (arXiv:2406.04244), "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
  17. LMSYS Org, "LLM Decontaminator (blog post, 14 November 2023)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
  18. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  19. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data