Data licensing for AI training
Open data licenses for commercial AI training: a compatibility matrix
Quick answer
Most open dataset licenses allow commercial use for AI training if you meet their conditions. Public-domain dedications (CC0, PDDL) and attribution licenses (CC BY, ODC-By, CDLA-Permissive-2.0, MIT, Apache-2.0) form the lowest-risk core, provided notices travel with copies. Share-alike licenses (CC BY-SA, ODbL, CDLA-Sharing-1.0) permit commercial training but differ on whether share-alike reaches released weights. Non-commercial, no-derivatives, research-only and unlicensed data need new rights or counsel review. Check every license at its source, not on the hub tag.
By SourceX Editorial · Updated
Five license terms that decide commercial training
A commercial-training decision turns on five terms in each license: whether commercial use is allowed, what attribution it requires, whether share-alike reaches what you release, which uses it forbids, and whether any term reaches trained models or outputs. The matrix below scores each license on those terms; for the wider contract picture, see the AI training data licensing guide.
- Commercial use. Does the grant cover a company building a paid product? Non-commercial (NC) and research-only terms do not.
- Attribution. What notice must accompany copies, and does a released dataset, model card or product count as a copy?
- Share-alike. If you share adapted material, must it carry the same license?
- Use restrictions. Behavioral-use lists, field-of-use limits and click-through access terms layered on top of the content license.
- Reach to models and outputs. Some custom terms extend restrictions to trained models. JParaCrawl's research-only terms, for example, exclude translators trained on the corpus from commercial use [1].
Open licenses were drafted for copying and sharing content, not for training, so the hardest question has no settled answer: is a trained model adapted material? The builders of the GPT-NL Public Corpus describe the legal position as unclear and filtered out CC NonCommercial and ShareAlike sources so the corpus would stay usable for commercial training [2]. That is a conservative reading, not a legal requirement, and the matrix flags where the reading changes the answer.
Jurisdiction also changes how much a condition binds. One peer-reviewed chapter argues that US fair use may override Creative Commons conditions for training, while EU text-and-data-mining rules and the AI Act add duties [3]. Article 53(1)(c) of the EU AI Act requires general-purpose model providers to keep a copyright policy that honors rights reservations made under Article 4(3) of the DSM Directive [4]. As of October 2026, the UK's text and data analysis exception, section 29A of the Copyright, Designs and Patents Act 1988, covers only non-commercial research [5]; license or rely on fair use explains how teams weigh these routes.
The matrix: 15 license types scored for commercial training
The table sorts common dataset licenses into four tiers for a team training a commercial model. Tiers describe review effort, not legal conclusions.
| License | Commercial training | Attribution duty | Share-alike or downstream reach | Tier |
|---|---|---|---|---|
| CC0 1.0 | Yes (public-domain dedication) | None required [6] | None | A |
| PDDL 1.0 | Yes (public-domain dedication for databases) | None required | None | A |
| CC BY 4.0 | Yes | Credit the source, link the license, note changes | None | A |
| ODC-By 1.0 | Yes | Attribution notice for the database | None | A |
| CDLA-Permissive-2.0 | Yes | Keep the license text with shared data | None on the data; read its clause on computational results | A |
| MIT | Yes | Keep the copyright and license notice in copies | None | A |
| Apache-2.0 | Yes | Keep the license, any NOTICE file and change notes | None | A |
| CC BY-SA 4.0 | Yes | As CC BY | Shared adapted material carries the same or a compatible license; reach to weights unsettled [2] | B |
| ODbL 1.0 | Yes | Notice | Applies to adapted databases you use publicly; read how it treats works produced from the database | B |
| CDLA-Sharing-1.0 | Yes | Keep the license text | Applies to data you share; read its results clause | B |
| CC BY-ND 4.0 | Yes, but adapted material may not be shared | As CC BY | 4.0 allows adaptations you keep private; earlier ND versions grant no adaptation right; preprocessing may count as adaptation [7] | C |
| RAIL family (e.g., OpenRAIL) | Depends on the variant | Per license | Use-based restrictions; check flow-down to your users | C |
| CC BY-NC, BY-NC-SA, BY-NC-ND | No, without separate permission [2] | As CC BY | NC-SA adds share-alike | D |
| Research-only custom terms | No [8] | Per terms | Some extend to trained models [1] | D |
| None stated or "Unspecified" | Unknown; treat as all rights reserved | Unknown | Unknown | D |
Tier key: A, train and keep notices. B, train, but decide before releasing data or weights. C, counsel review before training. D, do not train a commercial model without new rights.
Notes on the rows that cause the most rework:
- Software licenses on data (MIT, Apache-2.0). Both are permissive, and code corpora such as The Stack were built by keeping only repositories under permissive licenses [10]. Both expect the license text, and for Apache-2.0 any NOTICE file, to travel with copies, which reposted copies can drop.
- Share-alike licenses. Share-alike is not non-commercial: The People's Speech corpus is released under CC BY-SA for academic and commercial use [11]. Whether weights count as adapted material is covered in share-alike data licenses and trained models. ODbL and the CDLA pair were written for databases and computational use, each with its own clause on derived outputs; read CDLA-Permissive-2.0 and CDLA-Sharing-1.0 explained before relying on them.
- No-derivatives. A 2026 audit of African-language corpora argues that a NoDerivs clause silently prohibits tokenization and annotation [7], both routine preprocessing steps. The 4.0 ND licenses bar sharing adapted material rather than making it, so the license version and whether a processed copy will be shared decide how much this matters.
- Non-commercial and research-only. Microsoft's MS MARCO page states that its datasets are intended for non-commercial research purposes only [8]. How NC terms apply to training is covered in CC BY-NC and other non-commercial datasets.
- RAIL-family licenses. These attach a list of prohibited uses. Check whether the list must flow down to your model's users and whether it conflicts with your product terms.
Element-by-element analysis of the Creative Commons family is in Creative Commons licenses and AI training.
Combining licenses in one training mix
Licenses in a mix conflict at different moments. NC and research-only terms limit the use itself, so they matter even if nothing leaves your cluster, while share-alike and attribution duties are triggered when you share: publishing a merged dataset, shipping data inside a product and, on some readings, releasing weights. Three patterns cover most mixes:
- Per-record licenses. Articles in the PubMed Central Open Access Subset carry their own machine-readable Creative Commons or similar licenses, and not every PMC article is available for reuse [12]. Filter on each record's license field, not the collection name.
- Benchmark bundles. The BEIR paper lists component licenses that include CC BY 4.0, Apache 2.0, CC BY-SA 3.0 and 4.0, GPL v3 and non-commercial terms [13]. Treat a bundle as many datasets; which public retrieval datasets allow commercial use works through the common ones.
- Incompatible pairs. CC BY-SA and CC BY-NC material cannot be merged into one adapted dataset under a single license [7]: share-alike requires the same terms downstream, and NC adds a restriction those terms do not contain. Whether shipping them side by side as separately licensed files avoids the conflict is a question for counsel.
Which terms bite at each release point
| Mix | Internal training only | Publishing the merged dataset | Releasing weights or a hosted model |
|---|---|---|---|
| Tier A only | Keep an attribution file | Carry every notice and license text | Ship the attribution file with the model card |
| Tier A plus CC BY-SA | Proceed | Adapted CC BY-SA material goes out under CC BY-SA or a compatible license | Decide with counsel whether weights are adapted material |
| CC BY-SA plus CC BY-NC | NC part blocks commercial use | Cannot be merged under one license [7] | Exclude the NC part first |
| Anything plus CC BY-ND | 4.0 allows private adaptations, older versions do not; preprocessing may be adaptation [7] | No adapted release | Counsel review |
| Anything plus research-only terms | Follow the terms; some reach trained models [1] | Check redistribution terms | Only if the terms allow it |
| Anything plus no license | Exclude until a license is found | Exclude | Exclude |
Releasing weights raises its own notice and flow-down questions; see releasing open-weight models trained on licensed data.
Why a hub's license tag is not the license
A dataset hub's license field is metadata entered by whoever uploaded the data, so it can be missing, wrong or broader than the rights in the underlying content. On Hugging Face the license sits in the YAML block at the top of the dataset card, where it drives search filters [14].
Audits show how often that field misleads:
- The Data Provenance Initiative reports license omission above 70% and error rates above 50% on popular dataset hosting sites [15]. Its preprint found that 66% of the Hugging Face licenses it analyzed fell in a different use category from the author's intended license, often a more permissive one [9].
- A 2021 study notes that a public dataset may be hosted in several places and built from several sources, each with its own license. It found potential license-violation risks in five of six widely used image datasets if used to build commercial software [16].
- An audit of 364,000 Hugging Face datasets and 1.6 million models found that 35.5% of model-to-application transitions dropped restrictive clauses by relicensing under permissive terms [17].
Access terms are a second layer. Stack Exchange user content is licensed under CC BY-SA, yet in 2024 the company placed its data dump behind a login and an agreement not to use it for AI training, which critics said conflicted with the license [18]. Read click-through terms on their own; see click-through dataset licenses on data marketplaces and gated and custom-licensed datasets.
Generator terms are a third layer. The Data Provenance Initiative found that newer synthetic data tend to be restrictively licensed [15]. As evidence of market practice, one model provider's help center says, as of October 2026, that its terms do not allow outputs to be used to train competing models [19]. Material generated without sufficient human input may also fall outside copyright, so a Creative Commons label on it may describe rights nobody holds [3]; open instruction and preference datasets applies these checks to fine-tuning sets.
Worked example: screening a five-source fine-tuning mix
A license screen records, per source, the license you can prove, any terms layered on top, and a decision tied to how the model will ship. The record below covers a mix for a support assistant that will be sold as a hosted product.
Illustrative example: invented to show structure; it does not describe an available dataset.
mix_id: FT-SUPPORT-07
release_plan: hosted_api # internal_only | hosted_api | open_weights
sources:
- id: S1
declared_license: CC0-1.0
evidence: LICENSE file in upstream repository, copy stored 2026-09-30
decision: include
- id: S2
declared_license: CC-BY-SA-4.0
access_terms: download agreement prohibits AI training
decision: hold # access contract conflicts with intended use
- id: S3
declared_license: apache-2.0 # dataset card YAML only, no LICENSE file
generator: commercial model API; output terms restrict competing models
decision: counsel_review
- id: S4
declared_license: mixed
components: {CC-BY-4.0: 2, CC-BY-SA-3.0: 4, GPL-3.0: 1, non_commercial: 3}
decision: split # NC components removed
- id: S5
declared_license: per_record
record_filter: license in [CC0-1.0, CC-BY-4.0]
decision: include_filtered
attribution_file: ATTRIBUTION.md # lists S1 too, although CC0 requires no credit
How the decisions follow from the matrix:
- S1 passes at tier A. Listing it in the attribution file still helps when disclosure rules ask for dataset sources.
- S2 waits. CC BY-SA is tier B, but the access agreement, not the content license, conflicts with training.
- S3 goes to counsel. The only license evidence is a card tag, and the generator's terms may restrict the intended use.
- S4 is split. The NC components leave; the share-alike and copyleft components stay, subject to counsel's view, because a hosted API does not publish the merged dataset.
- S5 is filtered to records whose own license is CC0 or CC BY.
If the release plan changed to open weights, S4's share-alike components would go back to review. Re-run the screen whenever the release plan changes.
What to record so the screen survives an audit
Keep one license record per source and one decision per release, because EU and California disclosure rules ask what data trained a model and where it came from. As of October 2026, Article 53(1)(d) of the EU AI Act requires general-purpose model providers to publish a summary of training content using the AI Office template [4]. California's AB 2013 required developers of generative AI systems made available to Californians to post training-data documentation by 1 January 2026 and before each later release or substantial modification, covering dataset sources or owners and whether datasets include copyrighted or licensed material [20].
Record these fields for each source:
- SPDX-style license identifier, version and license URL
- Evidence: the LICENSE file, terms page or contract, with retrieval date and a stored copy
- Upstream sources and their licenses, for compilations
- Access or click-through terms accepted, by whom and when
- Generator model and its output terms, for synthetic data
- The exact attribution string the license requires
- Flags: commercial use, share-alike, no-derivatives, use restrictions, reach to models
- Decision, reviewer, date and the releases it covers (dataset, weights, API)
Tooling covers parts of this. The Data Provenance Explorer traces and filters provenance for popular fine-tuning collections [9], and the license-drift study released a rule engine covering almost 200 SPDX and model-specific clauses [17]. For the audit procedure, see auditing open dataset licenses before commercial training; to keep decisions current, use an AI training data register.
When open-licensed data runs out
Open licenses cover enough general text to pretrain capable models, but rarely the domain records a commercial model is built to handle. The Common Pile v0.1 gathers 8 TB of openly licensed and public-domain text from 30 sources, and its authors report that 7B models trained on it perform competitively with models trained on unlicensed text at similar compute [21]. The Data Provenance Initiative found the opposite pattern in lower-resource languages, creative tasks and newer synthetic data, where datasets tend to be restrictively licensed [15].
Operational records such as support tickets, engineering histories and finance workflows sit inside companies with no public license at all. The routes are to negotiate commercial rights for research-only datasets, to weigh the real cost of "free" datasets, or to agree a dataset license with the holder. SourceX's comparison of licensed, synthetic and scraped training data sets out the trade-offs.
SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing agreement. Each dataset goes through rights review and is delivered under a license that defines which records are included and what they can be used for. Datasets are sourced on request, so a request does not guarantee a match; you can describe the records you need to SourceX.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Need training rights an open license cannot give?
Describe the data you need and the uses you need licensed. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license and delivery. Submit your licensing requirements.
Sources
- NTT Communication Science Laboratories, "JParaCrawl". https://www.rd.ntt/cs/team_project/icl/lirg/jparacrawl/
- arXiv:2604.00920, "GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training" (2026). https://arxiv.org/pdf/2604.00920
- IntechOpen, "Chapter on Creative Commons licensing and AI training (online-first chapter 1233275; exact title to be confirmed)". https://www.intechopen.com/online-first/1233275
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- UK Intellectual Property Office (unofficial consolidation hosted on GOV.UK), "Copyright, Designs and Patents Act 1988 (consolidated), section 29A: Copies for text and data analysis for non-commercial research". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
- arXiv:2501.11003, "Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo" (2025). https://arxiv.org/pdf/2501.11003
- arXiv:2606.28867, "Preprint auditing Creative Commons licensing in African NLP corpora (exact title to be confirmed)" (2026). https://arxiv.org/abs/2606.28867
- Microsoft, "MS MARCO: Datasets". https://microsoft.github.io/msmarco/Datasets.html
- Longpre et al., arXiv:2310.16787, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Kocetkov et al. (BigCode), arXiv:2211.15533, "The Stack: 3 TB of permissively licensed source code" (2022). https://export.arxiv.org/abs/2211.15533?context=cs
- arXiv:2111.09344, "The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage" (2021). https://arxiv.org/pdf/2111.09344
- National Library of Medicine (Data.gov catalog entry), "PubMed Central Open Access Subset (PMC OA)". https://catalog.data.gov/dataset/pubmed-central-open-access-subset-pmc-oa
- Thakur et al., arXiv:2104.08663, "BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
- Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
- MIT Media Lab (publication record for Longpre et al.), "A large-scale audit of dataset licensing and attribution in AI (Nature Machine Intelligence 6, 2024)" (2024). https://www.media.mit.edu/publications/a-large-scale-audit-of-dataset-licensing-and-attribution-in-ai/
- Rajbahadur et al., arXiv:2111.02374 (v4), "Can I use this publicly available dataset to build commercial AI software? Most likely not" (2021). https://arxiv.org/abs/2111.02374v4
- arXiv:2509.09873, "From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem" (2025). https://arxiv.org/abs/2509.09873v1
- DevClass, "Stack Exchange restricts access to dump of user-contributed data as critics complain license permits reuse for any purpose" (2024). https://devclass.com/2024/07/30/stack-exchange-restricts-access-to-dump-of-user-contributed-data-as-critics-complain-license-permits-reuse-for-any-purpose
- Anthropic (Claude Help Center), "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- California Legislature (California Legislative Information), "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Kandpal et al., arXiv:2506.05209 (NeurIPS 2025 Datasets and Benchmarks Track), "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.