Skip to content

Data licensing for AI training

CC BY-NC and other non-commercial datasets: can a company train on them?

Quick answer

Usually not, once the model serves the business. CC BY-NC restricts the purpose of a use, not the type of user, and Creative Commons reads it as reaching every step that needs permission: copying the data, training and sharing the model. That makes CC BY-NC commercial model training a poor bet for anything you sell, deploy or release, and leaves internal research at a company in a gray zone. Custom research-only terms are often stricter. The workable routes are a commercial license or replacement data.

By SourceX Editorial · Updated

What the NonCommercial condition restricts

NonCommercial restricts why you use the material, not who you are: the CC 4.0 licenses define it as use "not primarily intended for or directed towards commercial advantage or monetary compensation." A for-profit company is not excluded by name, but a use aimed at its products, customers or revenue is.

Three consequences shape every training decision:

  • The test runs per use. A university lab and a startup apply the same definition. The startup's uses are simply more often directed at commercial advantage.
  • The conditions bind only where you need permission. The 4.0 legal code says the license does not apply where a copyright exception or limitation covers your use. Creative Commons' guidance, "Using CC-Licensed Works for AI Training", makes the same point and notes that its advice may lead to over-compliance.
  • Where permission is needed, every licensed step must stay noncommercial. The same guidance treats making training copies, training and distributing the model as stages that must each be noncommercial.

Check the version and variant too. BY-NC-SA adds share-alike, BY-NC-ND bars sharing adapted material, and older versions are separate legal texts: the BEIR benchmark paper lists SciFact under CC BY-NC 2.0, for example [1]. How the BY, SA and ND elements behave in training is covered in Creative Commons licenses and AI training, element by element.

How NC reads on the training work a company actually does

The closer a use sits to revenue, a product decision or a released artifact, the harder it is to call noncommercial. The table assumes the use needs copyright permission.

ScenarioSteps that touch NC materialReading under the NC purpose testPractical position
University co-author trains on NC data for a published paperCopies, training, publicationFits the test if the work is not directed at a company's advantageGenerally within the grant; check who funds and directs the work
Company researchers test a method on NC data, then delete the modelCopies, trainingGray: corporate research usually serves future productsCounsel sign-off, a written purpose and a deletion record
Scoring your own model on an NC benchmark before releaseCopies, inference on test itemsThe license has no evaluation exception, and the score informs a commercial decisionTreat as commercial; see benchmark license checks
Full or LoRA fine-tuning of a model behind a paid product or an internal operations toolCopies, training, deploymentDirected at commercial advantageOutside the grant
NC text in a pre-training mixture for a model served by APICopies, training, servingDirected at commercial advantageOutside the grant; filter before the run
Reward model, quality classifier or data filter trained on NC data for a commercial pipelineCopies, training an auxiliary modelIndirect, but it serves the commercial modelTreat as commercial
Synthetic data from an NC-trained model used to train a commercial modelOutputs of the NC-trained modelUnsettled under CC; some custom terms name derived data expressly [2]Treat as tainted unless the terms say otherwise
Releasing weights trained on NC data under a permissive licenseDistributionThe release may be noncommercial, but every commercial user downstream inherits the questionCarry the NC terms forward, or do not release
Training for a client under a research contractCopies, training, deliverablePaid work directed at a company's advantageOutside the grant; UK guidance on a parallel exception says contract research for an outside company is unlikely to be non-commercial [3]

The internal-research row carries the most risk for startups: a prototype that only informs which product to build can still be read as directed at commercial advantage, and deleting the model leaves the copies, logs and results behind.

Do NC terms follow the trained weights?

Under CC licenses nobody can answer that with certainty, because whether trained weights are adapted material of their training data is an unsettled copyright question. Some custom research-only licenses write the answer down, and it is yes.

  • Unsettled under copyright. The authors of the GPT-NL Public Corpus describe legislation and case law as currently unclear, and their license filter excludes NC- and SA-licensed material [4].
  • Settled by contract. NTT's terms for JParaCrawl, a large public English-Japanese parallel corpus, limit it to research and exclude data derived from it and translators trained on it from commercial use; commercial licensing goes through NTT [2].
  • Settled by the model card. When a published model's card carries an NC dataset license forward to the weights, anyone building on that model takes the restriction with it.

The engineering consequence is that an NC source cannot be subtracted later. Removing it means retraining from a checkpoint saved before the source entered the mix, so keep NC experiments in separate runs or in LoRA adapters you can delete, not in merged weights. The same logic governs licensed data when a contract ends; see what happens to trained models when a data license ends.

Research-only terms that go further than CC BY-NC

Custom licenses define "commercial" by profit, by product development or by derived output instead of CC's purpose test, and several reach trained models expressly. Read each one on its own terms rather than mapping it to CC BY-NC.

TermsRestriction as writtenReaches models or derived data?Commercial route
CC BY-NC 4.0Use not primarily directed at commercial advantage or monetary compensationUnsettledSeparate permission from each rightsholder
MS MARCO (Microsoft)Non-commercial research purposes only, provided without extending any license or other intellectual property rights [5]No rights are extended to rely onNone found in our sources; ask Microsoft
JParaCrawl (NTT)Research use onlyYes: derived data and trained translators are excluded from commercial use [2]Contact NTT
LDC corpora, non-member organizationsNo use to develop or test products for commercialization, or in any commercial product, per LDC notices [6]Covers product testing, not only trainingFor-profit membership or a commercial license agreement [7]
Planet E&R basic terms of serviceNoncommercial means "use of Content for purposes that do not include uses from which Licensee will derive profits" [8]Not addressed by the definition itself; a profit test can also catch internal toolsCommercial license from the provider

The LDC for-profit membership agreement shows what a commercial grant can look like: a license for linguistic and language-based research and technology development at listed sites, with the right to incorporate portions of the data into the member's own work products, including for commercial purposes, to the extent copyright law and user agreements allow [7]. Planet's terms show one vendor's wording, not a standard. Click-through and hub terms are covered in gated and custom-licensed datasets on model hubs.

Where NC data hides in a commercial training mix

NC terms usually enter a commercial pipeline through aggregates, mirrors and mislabeled cards rather than a deliberate choice, so the check has to run per upstream source, not per download.

  • Missing and wrong licenses are common. The Data Provenance Initiative audited more than 1,800 text datasets and reported license omission above 70% and license error rates above 50% on popular dataset hosting sites [9].
  • Datasets inherit licenses from many sources. A study of six widely used public image datasets found potential license-violation risks in five of them if used commercially, and notes that a dataset may be built from multiple sources with different licenses [10].
  • One dataset, several regimes. The SeeFar geospatial dataset applies per-source licenses: WorldStrat imagery from Airbus is CC BY-NC 4.0, while Sentinel data is CC BY-SA 3.0 IGO [11].
  • Benchmark suites mix terms. The BEIR paper lists MS MARCO under MIT while describing it as for non-commercial research [1], and Microsoft's own page restricts it to non-commercial research [5].
  • Literature corpora license per article. Articles in NLM's PubMed Central Open Access Subset carry their own license terms, so check each article's license statement.
  • Popular SFT conversation sets. A third-party catalog lists ShareGPT 52K as CC BY-NC 4.0 and LMSYS-Chat-1M as non-commercial; confirm on the original cards before relying on either [12]. Alternatives are compared in open instruction and preference datasets that allow commercial fine-tuning.

A repeatable method for tracing upstream sources is in auditing open dataset licenses before commercial training.

Worked example: tracing one NC source through a fine-tuning stack

One NC source can reach five artifacts in a post-training pipeline, and a per-source audit record is how a team finds them before release.

Illustrative example: invented to show structure; it does not describe an available dataset.

A startup fine-tunes a permissively licensed base model into a support agent. Its SFT mix holds 40,000 conversations, 6,000 of them from a conversation set whose card says CC BY-NC 4.0. The same set supplied 2,400 preference pairs for a reward model, the SFT model generated a synthetic batch, and a policy was trained against the reward model. The audit record for that one source looks like this:

source_id: conv-set-07
card_license_field: cc-by-nc-4.0
license_text_checked: CC BY-NC 4.0 legal code linked from upstream repository
rightsholders: many individual contributors; no single licensor
restriction_reaches: [training_copies, training, model_distribution]
custom_terms_on_models: none found
used_in:
  - {artifact: sft-run-12, role: sft_training, rows: 6000}
  - {artifact: rm-run-03, role: reward_model_training, pairs: 2400}
  - {artifact: synth-batch-05, role: generated_by, parent: sft-run-12}
downstream: [policy-run-04, support-agent-v2]
commercial_route: none identified
decision: exclude_and_retrain
retrain_from: base checkpoint before SFT
approved_by: [ml_lead, counsel]
decided_on: 2026-10-09

Remediation here is bounded: drop the 6,000 rows, retrain SFT from the base checkpoint, discard the synthetic batch, retrain the reward model and rerun the policy. The same source inside a pre-training mixture would cost far more to remove, so the screen belongs before the first run.

Because CC conditions bind only where copyright permission is needed, an exception or fair use can matter more than the NC element. As of October 2026, none of the UK, US or EU regimes lets a commercial team rely on an exception with confidence before training on NC material.

  • UK. Section 29A of the Copyright, Designs and Patents Act 1988 permits copies for computational analysis only for non-commercial research, and those copies may not be transferred or used for other purposes without permission [13]. UK IPO guidance warns that researchers whose purpose is not solely non-commercial are very likely to be infringing [3]. The government's 18 March 2026 copyright and AI report describes s29A as applying only when its conditions are met [14]. See UK AI training and copyright in 2026.
  • US. The Copyright Office's Part 3 report, still a May 2025 pre-publication version, concludes that many training acts may be prima facie infringing unless an exception such as fair use applies [15]. On 29 September 2026 the Third Circuit held that ROSS's use of Westlaw material to train a non-generative legal research tool was not fair use [16]. In Kadrey v. Meta, a June 2025 ruling in the Northern District of California found fair use on that record, while other claims continue [17]. Fair use is a defense to copyright infringement, not by itself to a claim that you broke terms you accepted to download a gated dataset; license or rely on fair use covers how teams weigh the two.
  • EU. The general text-and-data-mining exception in Article 4 of Directive 2019/790 does not apply where rightsholders have expressly reserved their rights, and providers of general-purpose AI models must keep a copyright policy that complies with those reservations and publish a summary of training content [18]. Our sources do not settle whether an NC license counts as a reservation.

Getting commercial rights or replacing the data

When an NC source matters to your model, there are three routes: buy commercial rights from whoever can grant them, switch to openly licensed data filtered for NC terms, or license comparable records from the businesses that hold them.

  1. Ask the rightsholder. A CC license is non-exclusive, so the licensor can grant separate commercial terms, and institutional corpora often have a priced route. For one Chinese-English patent corpus, LDC's catalog listed, as of October 2026, US$25 for not-for-profit research use and US$5,000 for for-profit organizations under a Commercial License Agreement, with a lower listed fee for LDC for-profit members [19]. The route fails when rights are spread across thousands of contributors, as with user-shared chat logs. Negotiation is covered in research-only datasets: getting commercial rights.
  2. Use corpora filtered for commercial use. GPT-NL's license filter drops NC and SA material [4]. The Common Pile v0.1 assembles eight terabytes of public-domain and openly licensed text from 30 sources, and its authors report that their 7B Comma v0.1 models perform competitively with models trained on unlicensed text at similar compute budgets [20]. Still verify each source's license.
  3. License operational data directly. Support and sales histories, engineering records and finance workflows rarely exist under any open license. SourceX sources operational datasets from US companies on request: it looks for businesses that hold the data you describe, checks each supplier's licensing permissions in a rights review, and delivers under a license that defines which records are included, what they can be used for and for how long. Datasets are not held in stock, so a request does not guarantee a match; describe the data you need to SourceX. The trade-offs are compared in licensed vs synthetic vs scraped AI training data, and lower-cost options in data licensing for AI startups on a budget.

NC screening checklist before a training run

Run this screen on every source before it enters a pre-training mixture, an SFT set, a reward-model set or an eval suite.

  • Read the license text, not the card's license field, and record the version and any added terms.
  • Trace every upstream source in aggregates, mirrors and benchmark suites.
  • Classify each source as commercial-use, CC NC, custom research-only or unknown, and treat unknown as NC until resolved.
  • For custom terms, note whether they reach derived data, trained models or product testing.
  • Map every artifact that touched NC material: checkpoints, adapters, reward models, filters, synthetic sets and eval reports.
  • Keep any counsel-approved research use of NC data in isolated runs, with a written purpose, an approver and a deletion date.
  • Before release, confirm the model card and any EU training-content summary match the sources actually used [18].
  • Choose a replacement route for every NC source the product depends on.

Compare licenses side by side in the open data license compatibility matrix, or start from the AI training data licensing guide.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Need training data you can use commercially?

Describe the records you need and the uses your models require, from fine-tuning to deployment. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license, delivery and future purchases. Start a data request with SourceX.

Sources

  1. Thakur et al., "BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models" (2021). https://arxiv.org/pdf/2104.08663
  2. NTT Communication Science Laboratories, "JParaCrawl". https://www.rd.ntt/cs/team_project/icl/lirg/jparacrawl/
  3. UK Intellectual Property Office, "Exceptions to copyright: Research" (2014). https://assets.publishing.service.gov.uk/government/uploads/system/uploads/attachment_data/file/375954/Research.pdf
  4. arXiv:2604.00920, "GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training" (2026). https://arxiv.org/pdf/2604.00920
  5. Microsoft, "MS MARCO Datasets". https://microsoft.github.io/msmarco/Datasets.html
  6. LINGUIST List, "LINGUIST List 28.4892 (Linguistic Data Consortium announcement)". https://linguistlist.org/issues/28/4892/
  7. Linguistic Data Consortium, "LDC For-Profit Membership Agreement". https://Catalog.Ldc.Upenn.Edu/license/ldc-for-profit-membership.pdf
  8. Planet Labs, "Basic Terms of Service (E&R, revised May 2023)" (2023). https://assets.planet.com/docs/E&R_Basic_ToS_Revised_May_2023.pdf
  9. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  10. Rajbahadur et al., "Can I use this publicly available dataset to build commercial AI software? Most likely not" (arXiv 2021). https://arxiv.org/abs/2111.02374v4
  11. Registry of Open Data on AWS, "SeeFar". https://registry.opendata.aws/seefar/
  12. LLM Configurator (third-party dataset catalog), "ShareGPT 52K: LLM Instruction / SFT Dataset". https://llmconfigurator.com/en/datasets/sharegpt
  13. UK Intellectual Property Office (unofficial consolidation hosted on GOV.UK), "Copyright, Designs and Patents Act 1988 - Consolidated (section 29A)". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
  14. UK Government (GOV.UK), "Report on Copyright and Artificial Intelligence" (2026). https://www.gov.uk/government/publications/report-and-impact-assessment-on-copyright-and-artificial-intelligence/report-on-copyright-and-artificial-intelligence
  15. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  16. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (precedential opinion)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  17. Akin Gump Strauss Hauer & Feld LLP, "Second District Court Rules AI Training Can Be Fair Use (Kadrey v. Meta)" (2025). https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/second-district-court-rules-ai-training-can-be-fair-use
  18. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  19. Linguistic Data Consortium, "Chinese-English Parallel Sentences Extracted from Patents (LDC2016T22)". https://catalog.ldc.upenn.edu/LDC2016T22
  20. Kandpal et al., "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data