Skip to content

Data licensing for AI training

Share-alike data licenses (CC BY-SA, ODbL, CDLA-Sharing) and trained models

Quick answer

Share-alike obligations attach to model weights or outputs only if the model or output counts as the thing each license regulates: "Adapted Material" under CC BY-SA, a "Derivative Database" or "Produced Work" under ODbL, or published "Enhanced Data" under CDLA-Sharing. None of the three texts names models, and no court has settled the question as of October 2026. Treat SA data as a release-gating risk: segregate it, log it per training run, and decide before training whether you could live with the license flowing to your weights.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What each share-alike trigger actually regulates

Each family hooks share-alike to a different legal object, so the same training run can be clean under one license and exposed under another. The 2024 IViR study commissioned by Open Future is a detailed analysis of the question; it concludes that whether share-alike reaches a model depends on tracing how a work was used in training, on whether a copyright exception covered that use, and on whether recognizable traces of the work survive in outputs [1][2].

  • CC BY-SA 4.0. ShareAlike applies when you share Adapted Material, defined by reference to modifications that require permission under copyright and similar rights. If training and the resulting weights do not require permission (for example under fair use or a TDM exception), the license conditions arguably never engage; Creative Commons' own guidance notes that license conditions only matter where a license is needed, while taking a cautious line on downstream sharing [3].
  • ODbL 1.0. Section 4.4 requires that a publicly used Derivative Database be offered under the ODbL; section 4.3 asks for an attribution notice on a Produced Work, but a publicly used Produced Work made from a Derivative Database also triggers an offer of that database; and a database reconstructed from Produced Works falls back under the ODbL [4][5]. A model can be argued to be either, and the classification, plus whether you altered the data first, decides whether you owe an offer of the derived database or merely a notice.
  • CDLA-Sharing-1.0. The CDLA family separates the Data from Results of computational use. Sharing obligations bind recipients who publish the data or their enhancements to it, which is why CDLA is described as copyleft for data rather than for models [6][7][8].

For a clause-level walk-through of the CDLA texts, see CDLA-Permissive-2.0 and CDLA-Sharing-1.0 explained; for the BY, NC and ND elements, see Creative Commons licenses and AI training element by element.

Weights versus outputs: two separate questions

Weights and outputs carry different share-alike risk and should be analyzed separately. Weights are the stronger candidate for "adaptation" or "derivative database" arguments when the model can regurgitate substantial parts of the SA corpus; outputs are exposed only when a specific output reproduces or closely adapts a specific SA work or extracts a substantial part of an SA database [1].

Weights. The argument that a model is Adapted Material needs the model to embody protected expression from the work, which memorization research makes plausible for duplicated or long-tail text. Under ODbL, the stronger argument is that a model trained to answer queries over a database (a geocoder trained on OpenStreetMap, for example) is a Derivative Database or functions as one. Embedding indexes and retrieval stores built from SA data are closer still: a vector store of chunked CC BY-SA text is hard to describe as anything other than a copy or adaptation.

Outputs. Most outputs of a general model will not contain enough of any one SA work to be an adaptation of it. The risk concentrates in narrow cases: code or prose regurgitated near-verbatim, map tiles or address lists reconstructed from ODbL data, and RAG answers that quote SA passages at length. These are detectable with n-gram overlap and canary checks, which gives you a practical control.

Share-alike is a license condition, so it only binds you if you needed the license in the first place. Where training is covered by an exception or by fair use, the licensor's conditions may never attach, but that is a legal position you must be able to defend, not a default [1][9].

As of October 2026, the U.S. Copyright Office's Part 3 report remains a pre-publication version; it concludes that many acts in training implicate reproduction rights and that fair use turns on the facts of each use [9]. In the EU, the commercial TDM exception is subject to rights reservations, and Open Future's proposal is precisely to pair a reservation with share-alike conditions so that copyleft flows down to model recipients [2]. In the UK, the CDPA s29A exception still covers only non-commercial research, so a commercial lab training on UK-sourced SA data generally falls back on the license.

Database rights add a layer: ODbL is built around the EU sui generis database right and contract, so an "it's just facts" argument that works for copyright may not dispose of ODbL obligations in Europe [5]. Counsel should treat ODbL as the family least likely to be neutralized by a copyright exception.

Risk ranking for commercial and open releases

Risk depends more on what you release and how than on which SA license sits in the corpus. The table ranks common scenarios from lowest to highest exposure, as a starting point for review rather than a legal conclusion.

Illustrative example: invented to show structure; it does not describe an available dataset.

ScenarioSA family in corpusLikely trigger argumentRelative riskTypical control
Internal eval set, never distributedAnyNo sharing or public useLowAccess controls, no release
Closed API model, SA text under 1% of pre-training tokens, dedup appliedCC BY-SAOutputs as adaptationsLow to moderateRegurgitation filters, attribution log
Open-weight release, SA text in pre-training mixCC BY-SAWeights as Adapted MaterialModerateDocument SA share, keep release license compatible
Fine-tuned model specialized on one SA corpusCC BY-SA / CDLA-SharingWeights closely derived from one sourceModerate to highPrefer permissive or licensed substitute
Geospatial model or index trained on ODbL data, publicly servedODbLDerivative Database, 4.4 offerHighOffer derived database or redesign
Public RAG store of chunked SA documentsCC BY-SA / ODbLCopy or adaptation, plainlyHighLicense the store under SA terms or exclude

The common failure mode is the fine-tune row: a team uses an SA corpus as its main SFT source, releases weights under a custom restrictive license with field-of-use bans, and creates a direct conflict with "no additional restrictions" language. If you plan restrictive terms, see Releasing open-weight models trained on licensed data and field-of-use restrictions.

Practical mitigations before training

The cheapest mitigation is a decision made at ingestion, because removing SA data from trained weights after the fact usually means retraining. Build the controls into your data pipeline rather than into the release checklist.

  1. Tag license at the record level. Carry an SPDX identifier (for example CC-BY-SA-4.0, ODbL-1.0, CDLA-Sharing-1.0) and a source URL on every shard, and fail the job if the field is null.
  2. Segregate SA shards. Keep SA material in named mixtures so you can report its token share per run and rebuild without it.
  3. Decide the release license first. A permissive release license carries less conflict risk than one with use restrictions, but if the weights were held to be Adapted Material, CC BY-SA would still require a BY-SA or compatible license, not Apache-style terms.
  4. Test for regurgitation. Run exact and near-duplicate overlap checks against the SA shards and keep the results with the model card.
  5. Keep attribution ready. Even where share-alike does not attach, BY and ODbL 4.3 attribution duties may; publish a source list with the release.
  6. Prefer negotiated licenses for core sources. A bilateral license that names model training, weights and outputs removes the interpretive question; see derivative and successor model rights.

The open data license compatibility matrix covers how SA terms combine with NC, ND and permissive sources in one mix, and the open dataset license audit covers how to verify what a dataset card claims.

Sample training-data license register entry

A register entry should let a reviewer answer, in one line, whether an SA source could reach a given release. The schema below is a minimal version teams can extend.

Illustrative example: invented to show structure; it does not describe an available dataset.

source_id: wiki-derived-qa-v3
spdx_license: CC-BY-SA-4.0
legal_object_triggered: "Adapted Material (s.3(b)) if weights or outputs adapt protected expression"
use_stage: [pretraining]
token_share_pct: 0.6
dedup_applied: minhash_0.8
regurgitation_test: {method: 13-gram overlap, max_overlap: 0.002, run_id: eval-2026-09-30}
release_targets: [api_closed, open_weights_apache2]
sa_conflict_with_release_license: false
attribution_published_at: "MODEL_CARD.md#data-sources"
counsel_review: {status: approved, note: "Re-review if used for SFT"}

When negotiated data avoids the question

Bilateral licenses avoid share-alike ambiguity because the parties write down what the model may do. A negotiated license can state whether weights, outputs and successor models are covered, which no public copyleft license does today [1].

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, and pricing and allowed uses are agreed in that license before anything is contracted. For the broader framework, start at the AI training data licensing guide, the AI data hub, or the data licensing glossary entry; buyers with a specific requirement can describe the data they need.

Sourcing training data without share-alike ambiguity

If share-alike terms make a corpus hard to use in your release plan, a negotiated license is the usual alternative. SourceX sources operational data from US companies on request, rights-reviews each dataset and agrees allowed uses in a license before delivery; a request does not guarantee a match. Describe the data you need at SourceX for buyers.

Sources

  1. Institute for Information Law (IViR), University of Amsterdam, "Mapping the Impact of Share Alike/Copyleft Licensing on Machine Learning and Generative AI" (2024). https://ivir.nl/publicaties/download/Share-Alike-and-ML.pdf
  2. Open Future, "The Impact of Share Alike/CopyLeft Licensing on Generative AI" (2024). https://openfuture.eu/publication/the-impact-of-share-alike-copyleft-licensing-on-generative-ai/
  3. Creative Commons, "Using CC-licensed works for AI training". https://creativecommons.org/using-cc-licensed-works-for-ai-training-2/
  4. OpenStreetMap Wiki, "ODbL 1.0 text". https://wiki.openstreetmap.org/wiki/OSMFJ/ODbL/1.0/text
  5. Wikipedia, "Open Database License". https://en.wikipedia.org/wiki/Open_Database_License
  6. LF AI & Data Foundation, "Community Data License Agreement (CDLA) repository". https://github.com/lfai/CDLA
  7. The Linux Foundation Projects (cdla.dev), "CDLA FAQ". https://cdla.dev/faq-resources/faq/
  8. data.world, "Common license types for datasets". https://docs.data.world/en/214274-common-license-types-for-datasets.html
  9. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data