Code and software engineering data
Open Code Datasets vs Licensed Private Code: Rights, Opt-Outs and Coverage
Quick answer
Openly licensed code corpora such as The Stack and the code portions of the Common Pile are a credible base layer for commercial LLM training, but they are not rights-free: you inherit license metadata errors, attribution and notice duties, opt-out removals and platform terms [1][2][3][4]. They also miss what enterprise code models most need: internal services, legacy stacks, linked review and CI history, and code your evaluation set has never seen. Licensed private code fills those gaps under a negotiated license.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What openly licensed code corpora actually give you
Open code corpora give you broad, cheap coverage of popular languages and public libraries, filtered by declared license. The Stack, released by BigCode, assembled about 3.1 TB of permissively licensed source code across 30 programming languages, selecting repositories by their detected license [3]. The Common Pile v0.1 is an 8 TB collection of public domain and openly licensed text from 30 sources that include code alongside papers, books and other domains [1].
For pre-training, that is a strong foundation: syntax, idioms, standard library usage and the long tail of open-source frameworks. It is also well documented, which helps when you write a training-content summary or a model card. The weakness is that "permissively licensed" is a property assigned by a pipeline, not a guarantee about every file.
Why "open" does not mean "clean" for commercial training
Open code datasets carry rights risk because license labels are frequently missing or wrong at the repository and file level. The Data Provenance Initiative audited more than 1,800 datasets and reported license omission rates above 70% and license error rates above 50% on popular dataset hosting sites [2]. Code adds its own failure modes on top of that general finding.
Common code-specific failures to test for:
- Repo-level label, file-level reality. A repository tagged MIT can vendor a GPL or AGPL file under
third_party/or paste a Stack Overflow snippet; repository-level filtering misses both. See copyleft contamination in licensed code datasets. - Missing LICENSE file. Code with no license grants no reuse rights by default; detection pipelines treat this differently.
- Relicensing over history. Projects move from Apache-2.0 to source-available licenses (or the reverse); a snapshot may predate or postdate the change.
- Forks and mirrors. A fork can carry a different declared license from its upstream, and mirrors can strip license headers.
- Secrets and personal data. Public commits contain API keys, emails in
AUTHORSfiles and author metadata in git logs; review secrets in code datasets.
Platform terms can also sit on top of content licenses. When Stack Exchange restricted access to its data dump in 2024, critics pointed out that the Creative Commons license on the content still permitted reuse; the dispute was about access terms, not the license itself [4]. Read both the content license and the terms of the place you get the data from. For a general audit method, use auditing open dataset licenses before commercial training.
Attribution, notice and opt-out duties you take on
Using open code creates ongoing obligations, not a one-time clearance. Permissive licenses such as MIT, BSD and Apache-2.0 still carry conditions, typically preserving copyright and license notices, and Apache-2.0 adds NOTICE file handling. Whether and how those conditions apply to model weights and outputs is unsettled, so most teams keep per-file license and provenance metadata so they can answer the question later.
Opt-outs add a moving target. The Stack published an "Am I in The Stack" tool and a process for developers to request removal of their code [3]. If you train on a version of such a dataset, decide in advance whether you will refresh to versions that honor later removals, and record which snapshot each model used. Check each dataset's current terms; they change between versions.
Regulation pushes in the same direction. As of October 2026, under Article 53(1)(c) of the EU AI Act, general-purpose AI model providers must have a copyright policy that identifies and complies with reservations of rights under Article 4(3) of the DSM Directive, and Article 53(1)(d) requires a public summary of training content [7]; these duties have applied since 2 August 2025. The Commission published the summary template on 24 July 2025 [9], and the GPAI Code of Practice copyright chapter describes how signatories maintain that policy [8]. In California, AB 2013 required developers of generative AI systems offered to Californians to post training-data documentation by January 1, 2026 [10]. See EU text and data mining opt-outs for the opt-out mechanics.
Coverage gaps open corpora cannot close
Public code over-represents libraries, tutorials and popular frameworks and under-represents how companies actually build software. The gaps that matter most for enterprise code models are structural, not just a matter of volume:
- Internal services and glue code: service-to-service clients, internal SDKs, feature-flag logic, auth middleware and migration scripts that never leave private monorepos.
- Legacy stacks: COBOL, PL/I, JCL, RPG, older Java EE and stored-procedure-heavy SQL are thin in public data. See COBOL and mainframe code datasets and production SQL query corpora.
- Linked history: ticket to branch to commits to review comments to CI result to merge, which supports issue-to-fix pairs and code review comment resolution data.
- Repository-scale context: whole-repo snapshots with build files, dependency lockfiles and test suites that run, needed for repository-level code context.
- Data and infrastructure code: dbt models, Airflow DAGs, Terraform and Helm charts tied to real production constraints.
Open corpora can include some of this, but rarely with the linked metadata (ticket IDs, reviewer decisions, CI status) that makes it usable for SFT, preference or agent training.
Contamination and held-out evaluation
Public code is a poor source of evaluation data because it is likely already in your pre-training mix. The SWE-Bench Pro authors argue that widely used permissively licensed repositories are prime candidates for inclusion in web-crawled pre-training corpora, which undermines benchmarks built from them [5]. Their response was to build 1,865 problems from 41 repositories split into public, held-out and commercial subsets, using copyleft repositories for the public and held-out subsets and drawing the commercial subset from private codebases that are not publicly accessible [5][6].
The lesson for buyers is that licensed private code is the main route to evaluation sets your model cannot have memorized, provided the license restricts redistribution and the supplier's code was never mirrored publicly. Check forks, package registry uploads and vendored copies before you trust a private repository as held-out. See code benchmark contamination and held-out coding agent evaluation sets from private repositories.
Decision table: when open data is enough and when to license private code
Open data suffices when your target is general completion in popular languages; licensed private code earns its cost when you need coverage, linked history or held-out evaluation that public data cannot supply.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Need | Open corpora usually enough? | What licensed private code adds | How to test before buying |
|---|---|---|---|
| General completion in Python, JS, Go | Yes, after license and dedup filtering | Little beyond style diversity | Compare held-out loss with and without a candidate sample |
| Enterprise internal services and monorepo patterns | Partly | Real internal APIs, cross-service calls, build systems | Near-dedup the sample against your corpus; measure novel share |
| COBOL, JCL, PL/I, RPG | Rarely | Volume and production-grade programs | Count tokens per language versus your current mix |
| Issue-to-fix and review data for agents | Rarely, without linked metadata | Ticket, diff, review and CI linkage | Check join rate between tickets, commits and CI runs |
| Held-out evaluation | No, contamination risk | Code with no public copies | Search hashes and distinctive identifiers against public mirrors |
| Low regulatory footprint | Depends on audit quality | One negotiated license per source, defined uses | Map each source to license, term and permitted use |
Dedup-against-your-corpus check (worked procedure). Before pricing a private code sample, run it through the same exact and near-duplicate pipeline you use for training: normalize whitespace and comments, hash files (for example SHA-256 for exact matches), compute MinHash signatures over token shingles for near duplicates, and report the share of files and tokens with no match in your existing mix. A sample whose "novel" share is small is mostly paying you for code you already have. Pair that with a held-out loss comparison on a small fine-tune to estimate value per token. Our guide to evaluating a code dataset sample covers the rest of the sample review.
How to structure the mix and the paperwork
Most code model teams end up with a layered mix: an audited open base layer for pre-training, licensed private code for coverage gaps and SFT, and licensed private code held out for evaluation. Keep per-file provenance (source, license, snapshot date, opt-out status) for the open layer, and per-source license terms for the licensed layer, so you can answer AI Act or AB 2013 documentation questions and remove material if a source changes terms.
For licensed private code, the key diligence questions differ from open data. Who owns the code, including contractor and acquired code, is covered in code ownership due diligence; trade-secret, confidentiality and regurgitation terms are covered in proprietary source code licenses for AI training. Specify languages, history depth, build and test requirements up front using the code dataset request specification. For the broader comparison across data types, see licensed vs synthetic vs scraped data and the code cluster hub.
SourceX sources operational datasets, including engineering records, from US companies on request, so you describe the code you need rather than picking from a catalog; a request does not guarantee a match. Details on licensing proprietary code datasets with full git history and on whether AI labs buy code are on our owner pages, and you can start a request on the buyer page.
Sourcing licensed private code to complement open corpora
SourceX looks for US businesses that hold the code data you describe, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery, and it is delivered under a license that defines records, uses, term and delivery. Describe the private code data your model needs.
Sources
- arXiv (Kandpal et al., EleutherAI and collaborators), "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/pdf/2310.16787
- arXiv (Kocetkov et al., BigCode), "The Stack: 3 TB of permissively licensed source code" (2022). https://export.arxiv.org/abs/2211.15533?context=cs
- DevClass, "Stack Exchange restricts access to dump of user-contributed data as critics complain license permits reuse for any purpose" (2024). https://devclass.com/2024/07/30/stack-exchange-restricts-access-to-dump-of-user-contributed-data-as-critics-complain-license-permits-reuse-for-any-purpose
- arXiv (Scale AI), "SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?" (2025). https://arxiv.org/abs/2509.16941v1
- Scale AI, "SWE-Bench Pro" (2025). https://scale.com/blog/swe-bench-pro
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.