Skip to content

Data licensing for AI training

Creative Commons licenses and AI training: BY, SA, NC and ND element by element

Quick answer

Creative Commons licenses do not prohibit AI training, but their conditions only apply when the training needs copyright permission. Where an exception covers the use, such as US fair use or an EU text and data mining exception, the BY, SA, NC and ND terms may not be triggered at all [1]. When permission is needed, BY requires reasonable attribution, SA and ND turn on whether a model or output is "adapted material," and NC bars commercial use. Buyers should audit every record's license version and keep attribution metadata.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

When CC conditions apply to model training

CC license conditions bind only acts that would otherwise need the licensor's permission, so the first question is which legal regime governs your training run [1]. CC's own guidance says that where training is covered by an exception or limitation, the licensee does not need the license and its conditions have limited reach [1]. That is the core of the steward's position, and it is why CC has said it does not plan an outright training ban [3].

The exception depends on where copies are made and for what purpose. In the US, the Copyright Office's Part 3 report, still a pre-publication version as of October 2026, treats many training steps (collection, curation, copying into training sets) as implicating the reproduction right and leaves fair use to a fact-specific analysis [8]. In the EU, commercial training relies on the Article 4 DSM exception, which a rightsholder can override with a machine-readable reservation; general-purpose model providers must have a policy to honor those reservations under AI Act Article 53(1)(c) [6]. Our guide to TDM exceptions by country covers the jurisdictional detail.

For a buyer, the practical consequence is that you cannot treat "CC-licensed" as either a free pass or a hard stop. If you rely on an exception, the license still matters as evidence of the rightsholder's expectations and as the fallback if the exception fails. If you rely on the license, every condition below has to be met at the scale of millions of records.

BY: attribution at corpus scale

Attribution is the one condition shared by all six CC licenses, and it is the condition most likely to fail in a training pipeline [2]. CC 4.0 lets attribution be given "in any reasonable manner based on the medium, means, and context," including by linking to a resource that holds the required information. That flexibility is what makes dataset-level attribution workable: a public manifest that lists creator, title or source URL, license and version, and modification notice for each work.

The failure mode is lost metadata, not bad intent. The Data Provenance Initiative audit of more than 1,800 text datasets reported license omission rates above 70% and license error rates above 50% on popular dataset hosting sites [5]. Once a crawl or aggregator drops the creator field, a corpus built on CC BY content can no longer satisfy BY, regardless of what the original license said.

Whether a trained model or its outputs must carry attribution is less settled. CC's guidance indicates attribution concerns arise mainly where outputs reproduce or closely adapt a licensed work [1]. Keep the manifest anyway: it is cheap at ingestion, near impossible to rebuild later, and it also feeds disclosure duties such as the EU training-content summary template of 24 July 2025 [7] and California AB 2013 documentation, which was due by January 1, 2026 [10].

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "img-000418273",
  "source_url": "https://example.org/photos/418273",
  "creator": "J. Rivera",
  "title": "Pump housing, side view",
  "license": "CC BY-SA",
  "license_version": "4.0",
  "license_url": "https://creativecommons.org/licenses/by-sa/4.0/",
  "modifications": "resized to 1024px; EXIF stripped; caption rewritten",
  "retrieved_at": "2026-09-14",
  "basis_for_use": "license",
  "split": "pretrain",
  "exclude_from_redistribution": false
}

SA: share-alike and whether a model is adapted material

ShareAlike requires that adapted material you share be licensed under the same or a compatible license, so the decisive question is whether a trained model, or a model output, counts as adapted material [2]. CC licenses define adapted material as content derived from the licensed work in which it is "translated, altered, arranged, transformed, or otherwise modified" in a way that requires permission. Model weights are statistical parameters, and most commentators doubt that weights trained across millions of works are an adaptation of any one of them [9].

That doubt is not a resolution. The IViR study commissioned by Open Future concludes that share-alike clauses have limited bite on model training under current law, while flagging that outputs which reproduce protectable expression can still be adaptations [4]. Treat SA as low risk for weights but material for three things: redistributing the training corpus itself, releasing fine-tuning sets built from SA text, and outputs that regurgitate SA content.

Share-alike also differs across license families. ODbL and CDLA-Sharing have their own "produced work" and derivative rules, covered in our page on share-alike licenses and trained models.

NC: the commercial-use line

NonCommercial restricts use "primarily intended for or directed towards commercial advantage or monetary compensation," and that definition turns on the user's purpose rather than its corporate form [2]. A commercial lab training a product model on CC BY-NC content is the clearest case of conflict where the license, rather than an exception, is the basis for use. An academic group that later spins out a company is the hardest case, because the purpose shifts after copies were made.

Exceptions can matter more than the license here. The UK's s29A text and data analysis exception is limited to non-commercial research, so it does not rescue a commercial training run, and EU Article 4 reliance depends on the absence of a reservation [6]. The open data license compatibility matrix maps NC terms against other open licenses for commercial training.

ND: no derivatives, private adaptation and release

NoDerivatives forbids sharing adapted material, but in version 4.0 it still permits creating adaptations for private use [2]. If training is an act requiring permission, a model trained privately on ND works and never distributed sits closer to that private-use allowance than to a breach. Distribution of the weights, or outputs that closely reproduce an ND work, is where the condition starts to bite.

Earlier versions are stricter in wording. Versions 2.0 and 3.0 frame the ND grant around reproducing and distributing the work "as incorporated in Collections," with less explicit room for private adaptation, so buyers should not apply 4.0 reasoning to 3.0 records.

Version differences: 2.0, 3.0 and 4.0

License version changes the analysis, so record it per item rather than per source site. Version 4.0 is the first international suite that expressly licenses sui generis database rights, which matters for EU-origin datasets where the database right, not copyright, may be the restriction on extraction. Version 4.0 also adds a cure mechanism: the license is reinstated automatically if a violation is fixed within 30 days of discovering it, which is relevant when an attribution manifest is repaired after release.

Version 3.0 ported licenses vary by jurisdiction, and some 3.0 ports treat database rights differently from the unported text. Versions 3.0 and earlier terminate automatically on breach with no comparable cure mechanism. In a corpus that mixes versions, each subset carries its own version's terms where the license is your basis for use, so apply the strictest terms wherever versions cannot be separated.

Decision table: CC elements by training use

The table below maps each element to the training acts most buyers perform, assuming the license rather than an exception is the basis for use.

Illustrative example: invented to show structure; it does not describe an available dataset.

ElementInternal pre-trainingSFT on CC textReleasing weightsRedistributing the corpusTypical buyer control
BYAllowed with attribution keptAllowed with attribution keptAttribution via manifest link advisableAttribution required per itemPer-record creator, URL, license, version
SAAllowedAllowedUnsettled whether weights are adapted material [4]Must use same or compatible licenseSegregate SA subsets; flag in split
NCConflict for commercial purposeConflict for commercial purposeConflict for commercial purposeNon-commercial onlyExclude from commercial runs
NDCloser to private adaptation (4.0)Closer to private adaptation (4.0)Higher risk if outputs reproduce worksVerbatim copies onlyExclude from released fine-tuning sets

Buyer checklist for CC-licensed corpora

A CC-heavy corpus is defensible only if each record's license can be proven at audit time.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • License per record: element set and version (for example "CC BY-NC-SA 3.0 IGO"), not a site-level label.
  • Basis for use: "license" or a named exception, per jurisdiction where copies are made.
  • Attribution fields: creator, title, source URL, license URL and modification notes retained through deduplication and filtering.
  • Opt-out screening: robots.txt, TDM reservation protocols and machine-readable signals checked at crawl time and logged [6].
  • Third-party content: stock photos, quoted text and embedded media inside CC pages often carry other licenses; see third-party content in licensed corpora.
  • Regurgitation testing: sample outputs against SA and ND subsets before any release.
  • Disclosure readiness: fields mapped to the EU training-content summary [7] and AB 2013 documentation [10].

The open dataset license audit guide covers the audit workflow in more depth.

CC Signals and emerging preference frameworks

Creative Commons has proposed CC Signals, a framework for rightsholders to state conditions for AI training such as attribution, financial contribution or open-model release, rather than a new license that bans training [3]. As reported in 2025, the signals are meant as reciprocity expectations layered over existing law, and their legal force depends on how they are implemented [3]. As of October 2026, treat them as a developing framework: log any signal present at collection time, but do not assume it carries contractual weight until the steward's published status confirms it.

When CC corpora are not enough

CC content is broad but thin on the operational records that enterprise models need, and it rarely comes with a single counterparty who can warrant ownership. SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, documents, and finance and legal workflows, rather than holding inventory. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery; buyers can describe the data they need. For the contract side, see the licensing hub, the data warranties guide and the data licensing glossary entry.

Licensing training data beyond Creative Commons

When CC-licensed corpora leave gaps, a direct license from the business that holds the data gives you a named counterparty and defined terms. SourceX finds US companies that hold the data you describe, assesses data and licensing permissions, and nothing is contracted until the supplier agrees. Start with a request at sourcex.si/buyers.

Frequently asked questions

Does Creative Commons allow AI training?

None of the six CC licenses prohibits training, and CC's guidance says the conditions apply only where the use needs copyright permission [1]. Whether you rely on the license or an exception changes which conditions you must meet.

Does CC BY require attribution inside model outputs?

Not as a rule; the concern arises mainly when an output reproduces or adapts a specific licensed work [1]. A dataset-level manifest linked from the model documentation is the common way to satisfy attribution for the training copies themselves.

Can CC0 content be treated differently?

Yes. CC0 waives copyright and related rights to the extent possible, so there are no BY, SA, NC or ND conditions to track, though third-party rights in the content, such as privacy or trademark, are not waived.

Sources

  1. Creative Commons, "Using CC-licensed works for AI training". https://creativecommons.org/using-cc-licensed-works-for-ai-training-2/
  2. Creative Commons, "Understanding CC licenses". https://creativecommons.org/share-your-work/understanding-cc-licenses/
  3. heise online, "No ban planned: Creative Commons is working on licenses for AI training" (2025). https://heise.de/-10461593
  4. Open Future (study by IViR, University of Amsterdam), "The Impact of Share Alike/Copyleft Licensing on Generative AI" (2024). https://openfuture.eu/publication/the-impact-of-share-alike-copyleft-licensing-on-generative-ai/
  5. Longpre et al., Nature Machine Intelligence 6 (2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://www.nature.com/articles/s42256-024-00878-8
  6. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  7. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  8. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  9. TechnoLlama (Andres Guadamuz), "Creative Commons and AI training". https://www.technollama.co.uk/creative-commons-and-ai-training
  10. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data