Data licensing for AI training
CDLA-Permissive-2.0 and CDLA-Sharing-1.0 explained for AI training data
Quick answer
CDLA-Permissive-2.0 lets you use, modify and share the licensed data for any purpose, including commercial model training, and it places no obligations on "Results" such as trained models [1][2]. The main condition is passing the license text along when you share the data itself. CDLA-Sharing-1.0 is the copyleft sibling: publishing modified data triggers a duty to share it under the same terms [4]. Neither label proves the licensor held the rights to grant, so check upstream content separately.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What the Community Data License Agreement family is
The Community Data License Agreements are Linux Foundation data licenses written for data rather than code or creative works, with AI and machine learning use as an explicit design target [1]. The family has three texts buyers meet in practice: CDLA-Permissive-1.0 and CDLA-Sharing-1.0 (the original 2017 pair) and CDLA-Permissive-2.0, released in June 2021 as a shorter, simpler replacement for the permissive 1.0 text [1]. The license texts and context documents live in the LF AI & Data repository [4].
Each version has a stable identifier for tooling. SPDX lists CDLA-Permissive-2.0 under that short identifier [2], and dataset hubs use lower-case tags such as cdla-sharing-1.0 in their license lists. Those identifiers are what your ingestion pipeline should key on, because a free-text "CDLA" string does not tell you which variant, or which obligations, apply.
What CDLA-Permissive-2.0 grants and requires
CDLA-Permissive-2.0 grants a broad right to use, modify and share the data, and attaches only a light notice condition to sharing the data onward [1][2]. In practice the operative terms reduce to four points:
- Grant. A recipient may use, modify and share the data, with no field-of-use or commercial-use limit [1][2].
- Notice when sharing data. If you redistribute the data, make the license text available with it [2].
- Attribution. Attribution to data providers is encouraged but is not a condition of the license [3].
- Results. Outcomes of computational analysis, including trained models, carry no obligations under the license [1].
The text also includes the usual warranty disclaimer and limitation of liability. That matters for procurement: the data comes as-is, with no representation that the contributor owned what it contributed, so any assurance you need has to come from somewhere else (see data warranties buyers should require).
How "Results" are treated, and why it matters for trained models
The defining feature for AI buyers is that CDLA-Permissive-2.0 separates the data from the Results of analyzing it, and expressly frees the Results [1]. The Linux Foundation's launch release states that the license defines Results to include machine learning models and imposes no obligations on their use or sharing [1]. So weights, embeddings, evaluation scores and model outputs derived from CDLA-Permissive-2.0 data can be shipped under your own terms.
This is the question most open licenses never answer. A CC BY 4.0 license, for example, does not say whether a model is an "adapted material," which is why attribution and share-alike analysis for models stays contested. With CDLA-Permissive-2.0 you do not have to win that argument for the license's own terms.
Two limits remain. First, the Results carve-out only covers what the CDLA licensor could license; it does not cure third-party copyright, privacy or contract claims in the underlying content. Second, if you redistribute the training data itself (for example, publishing a cleaned corpus next to your model), the notice condition applies to that redistribution [2].
How CDLA-Sharing-1.0 differs
CDLA-Sharing-1.0 is a copyleft data license: recipients can use and modify the data, but if they publish modified data they must share those changes under the same agreement [4][5]. The repository contrasts this directly with the permissive version, which carries no obligation to share changes [4].
For model teams, the practical questions are what counts as "publishing" and whether a model is caught. The sharing obligation in the 1.0 text is triggered by publishing data, not by internal computational use, and the 1.0 texts also distinguish computational results from the data. Read the operative definitions in the hosted text [4] rather than relying on summaries, because the 1.0 drafting is longer and more defined-term heavy than 2.0. Our share-alike licenses and trained models guide covers how CDLA-Sharing compares with CC BY-SA and ODbL on that question.
Where CDLA-Sharing-1.0 bites in practice is data pipelines: if you clean, relabel or augment a CDLA-Sharing dataset and then release that enhanced set (to a benchmark, a partner or a public hub), plan to release it under CDLA-Sharing-1.0 too. Keep CDLA-Sharing derivatives in a separately tagged lineage so a later corpus release does not inherit the obligation by accident.
CDLA vs CC BY for datasets
CDLA-Permissive-2.0 is usually simpler for model training than CC BY 4.0 because it settles the trained-model question and does not make attribution a condition [1][3]. CC BY 4.0 is still the more common dataset license, and it is workable, but it leaves more for counsel to interpret.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Question a buyer asks | CDLA-Permissive-2.0 | CDLA-Sharing-1.0 | CC BY 4.0 |
|---|---|---|---|
| Commercial training allowed? | Yes [1] | Yes, subject to sharing terms on published data [4] | Yes, with attribution |
| Trained model addressed? | Yes: Results carry no obligations [1] | Results distinguished from data; verify text [4] | Not addressed; contested |
| Attribution a condition? | No, encouraged [3] | Notice terms in 1.0 text; verify [4] | Yes |
| Share-alike on modified data? | No [4] | Yes, when published [4][5] | No |
| Redistribute data | Include license text [2] | Same license for published changes [4] | Attribution and license notice |
| Mixed inputs | CC0 can be included; CC BY terms still apply to that material [3] | Check compatibility per source | Per-source attribution |
The last row is the one teams miss. The CDLA FAQ notes that CC0 data can sit inside a CDLA-Permissive-2.0 dataset, but CC BY 4.0 attribution requirements still apply when that material is shared [3]. A single CDLA label on a compiled dataset does not erase the obligations of its components. For the broader picture, see the open data license compatibility matrix.
Checking upstream rights behind a CDLA label
A CDLA license is only as good as the contributor's right to grant it, so treat the label as a claim to verify, not a clearance. The Data Provenance Initiative audited more than 1,800 text datasets and reported license omission rates above 70% and license error rates above 50% on popular hosting sites [6]. Mislabeling is the base rate, not the exception.
Common failure modes for CDLA-labeled datasets:
- Scraped content relabeled. A contributor wraps web text or images it did not own in CDLA-Permissive-2.0; the license cannot pass rights it never had.
- Annotation vs source split. Labels are CDLA, but the underlying documents keep their original copyright or terms of service.
- Version drift. A dataset card says "CDLA" while the repository
LICENSEfile says CDLA-Permissive-1.0, which has different notice drafting. - Personal data. No data license addresses GDPR, CCPA or HIPAA obligations attached to records about people.
For providers placing general-purpose models on the EU market, Article 53(1)(c) of the AI Act requires a copyright-compliance policy [7], and the Commission's 24 July 2025 template sets the baseline format for the Article 53(1)(d) public summary of training content [8]. As of October 2026, a permissive dataset license helps document provenance for those obligations, but it does not substitute for a source-level record.
A CDLA intake checklist for ML and legal teams
The fastest way to clear a CDLA dataset is a short, repeatable intake record per dataset version.
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset: example-support-intents-v3
license_spdx: CDLA-Permissive-2.0 # from LICENSE file, not the card
license_text_hash: sha256:... # pin the exact text received
variant_confirmed_by: legal@buyer # 1.0 vs 2.0 vs Sharing checked
upstream_sources:
- component: utterances
origin: contributor-authored # not scraped
rights_basis: contributor attestation on file
- component: intent labels
origin: crowd annotation
rights_basis: annotator work-for-hire terms
embedded_third_party_licenses: [CC0-1.0] # CC BY parts need attribution kept
personal_data_review: pii scan + sample check
intended_use: [pre-training, sft, eval]
redistribution_planned: false # if true, ship license text
sharing_variant_lineage: none # tag any CDLA-Sharing derivatives
Record the license_spdx value from the actual license file and pin its hash, since hubs and cards drift [6]. If you work with data catalogs, map this record into your dataset metadata so the provenance trail follows the data into training manifests (see the Data Provenance Standards explainer).
When an open CDLA dataset is not enough
Open CDLA datasets suit benchmarks, research corpora and narrow fine-tuning sets, but they rarely contain the proprietary operational records that make enterprise models useful. When you need support histories, engineering tickets or finance workflows, you are negotiating a bespoke license, not adopting a public one, and terms such as allowed uses, term and delivery have to be written. Our guide to the AI training rights grant and the data licensing hub cover that path, and the dataset licensing glossary entry defines the core terms.
SourceX sources operational datasets from US companies on request and manages the licensing process; it holds no stock, so a request does not guarantee a match. You can describe the data you need to SourceX when open sets run out.
Sourcing licensed data beyond CDLA-Permissive-2.0
SourceX sources operational datasets from US companies, rights-reviews each one, and delivers it under a license that defines records, uses, term and delivery. Nothing is contracted until the supplying company agrees. Start a data request with SourceX.
Frequently asked questions
Can I use CDLA-Permissive-2.0 data to train a commercial model?
Yes, the license allows use and modification without a commercial-use limit, and it places no obligations on Results such as models [1][2]. Confirm the contributor actually had rights to the underlying content.
Do I have to credit the data provider?
Under CDLA-Permissive-2.0, attribution is encouraged but not required [3]. If the dataset embeds CC BY material, that material's attribution terms still apply when you share it [3].
Does CDLA-Sharing-1.0 force me to open-source my model?
The share-alike duty attaches to published data changes [4][5]. Whether any model release is caught depends on the 1.0 definitions, so have counsel read the hosted text [4] before releasing derivatives.
Sources
- The Linux Foundation, "Enabling Easier Collaboration on Open Data for AI and ML with CDLA-Permissive-2.0" (2021). https://www.linuxfoundation.org/press/press-release/enabling-easier-collaboration-on-open-data-for-ai-and-ml-with-cdla-permissive-2-0
- SPDX, "Community Data License Agreement Permissive 2.0". https://spdx.org/licenses/CDLA-Permissive-2.0.html
- Community Data License Agreement (cdla.dev), "CDLA FAQ". https://cdla.dev/faq-resources/faq/
- LF AI & Data (GitHub), "CDLA repository". https://github.com/lfai/CDLA
- data.world, "Common license types for datasets". https://docs.data.world/en/214274-common-license-types-for-datasets.html
- Longpre et al. (Nature Machine Intelligence 6, 2024), "A large-scale audit of dataset licensing and attribution in AI". https://www.nature.com/articles/s42256-024-00878-8
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.