Skip to content

Data licensing for AI training

Research-only datasets: getting commercial rights for AI training

Quick answer

A research-only license almost never covers shipping a model, so a team that prototyped on one has four routes: license commercial rights from the rights holder, join a consortium or distributor tier that sells commercial licenses, re-collect equivalent data under your own consents, or replace the corpus with licensed operational data. Start by reading how the license defines commercial use and "results," then check whether the original contributors' consents even allow a commercial grant. Budget for the possibility that the answer is no.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What a research-only license actually forbids

Research-only terms usually reach further than training: they cover the trained model, internal evaluation of commercial products and indirect business benefit. Qualcomm's dataset research license, for example, defines Research Use as non-profit, treats models trained on the data as "Results" governed by the license, and defines commercial use broadly enough to include indirect benefit [1]. TU Delft's MatchNMingle EULA counts testing a commercial system as commercial use [2], which closes the "we only benchmarked on it" argument.

Other well-known research corpora follow the same pattern. ImageNet's access terms limit use to non-commercial research and education [6], the IAM Handwriting Database is free only for non-commercial research and requires registration [7], and NTT's JParaCrawl excludes both derived data and translators trained on it from commercial use [5]. Statutory exceptions do not rescue you either: UK CDPA s29A permits text and data analysis copies only for non-commercial research [9], and the UK IPO's guidance says contract research for an outside company is unlikely to qualify [10]. As of October 2026, s29A remains limited to non-commercial research.

Before choosing a route, extract five clauses from the license text: the definition of commercial use, the definition of results or derivatives, permitted users (named researcher, institution, affiliates), termination and destruction obligations, and any clause about contacting the licensor for other uses. For licenses with standard Creative Commons NC terms rather than bespoke EULAs, the reading is different; see CC BY-NC and other non-commercial datasets.

Why retraining is not a clean fix

Retraining from scratch on commercial data is the safest remediation, but only if the research-only data and everything derived from it are fully removed from the pipeline. Contamination paths that teams miss include checkpoints initialized from the prototype, distilled or pseudo-labeled data generated by the prototype model, tokenizers or vocabularies fit on the research corpus, hyperparameters and data mixtures tuned against it, and eval sets carved out of it that now gate releases.

Licenses that define trained models as results [1] or exclude "derived data" [5] can plausibly reach these artifacts. Keep a lineage record (dataset ID, license version, git commit, checkpoint hash) for every run, and document the clean-room rebuild before launch. Where the research corpus contained personal data, the EDPB's Opinion 28/2024 also considers how unlawful processing during development can affect later deployment of a model [11], which is a separate question from copyright or contract.

Generating synthetic data from the research corpus is not a workaround either; whether outputs count as derivatives depends on the license text. See synthetic data from licensed data and derivative and successor model rights.

Route 1: license commercial rights from the rights holder

A direct commercial license is usually the most direct route when one identifiable party controls the data and the contributor consents allow it. Some publishers already run dual licensing, offering a research-only option alongside a separate commercial license that permits models derived from the dataset to be used commercially [3]. JParaCrawl's terms point commercial users to NTT directly [5].

The hard part is often finding who can grant rights. The Data Provenance Initiative found license information omitted for more than 70% of popular hosted datasets and license errors in more than 50% [8], so the license shown on a hub page may not be the operative one. Trace back to the original release, the paper's data statement and the institution's technology transfer office.

What to ask for: a training grant covering the specific model types and deployments you plan, rights to keep and commercialize models trained during the prototype phase (or an explicit release for past use), a representation that the licensor has authority to grant commercial rights, and the contributor consent basis. The AI training rights grant clause page covers wording.

Route 2: consortium and distributor commercial tiers

Language-resource consortia and distributors often separate research and commercial tiers, and membership can be a precondition for any commercial license. The Linguistic Data Consortium publishes a for-profit membership agreement, and membership is described as the gateway to commercial licenses for most LDC corpora [4]. Other distributors use their own tiers, so confirm current commercial terms directly with the distributor for each corpus.

Check three things before paying. First, whether the specific corpus is offered for commercial use at all, since some items are research-only regardless of membership. Second, whether "commercial use" in the distributor agreement covers model training and deployment, not only internal technology development. Third, whether membership-year timing affects which corpora you can license. Commercial speech corpora for ASR deserve extra scrutiny of speaker consent; see the open speech corpora license audit.

Route 3: re-collect equivalent data under your own consents

When contributors never consented to commercial use, which can happen with academic speech, dialogue, handwriting and behavioral datasets, re-collection (or replacement under Route 4) is usually the remaining option. You can reuse the published protocol (task design, prompts, annotation guidelines, metadata schema) as a specification, provided you do not copy protected text or recordings and the protocol documents themselves are not restricted.

Draft new consent forms that name commercial AI training, model release and data retention explicitly. For low-resource languages, community involvement and IP review matter from the start; researchers building African-language corpora describe licensing as a major bottleneck [12]. Re-collection takes longer, but you own the consent record, which is what reviewers will ask for. For conversational data specifically, see human dialogue corpora with commercial rights.

Route 4: replace with licensed operational data

If the prototype showed that a data type works, the commercial version may come from companies that hold similar records in their own systems: support tickets, call transcripts, engineering records or document workflows. This changes the rights question from "can the research licensor upgrade us" to "can this business license its own data, with personal details removed."

SourceX sources this kind of operational data from US companies on request and manages the licensing process; it is not held in stock, and a request does not guarantee a match. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe the data you need on the buyer request page. For a comparison of sourcing models, see licensed vs synthetic vs scraped data.

Choosing a route: decision table and request template

The right route depends on who controls the rights and what contributors agreed to. Use the table to shortlist, then send a structured request to the rights holder or distributor.

Illustrative example: invented to show structure; it does not describe an available dataset.

SituationLikely routeKey diligence questionMain failure mode
Single institution or company released the data; consents mention commercial use or are broadRoute 1: direct licenseDoes the licensor hold authority to grant commercial training rights?Hub license differs from original terms [8]
Corpus distributed by a consortium with tiered licensingRoute 2: consortium tierIs this corpus offered commercially, and does "commercial" cover model deployment?Membership bought, corpus still research-only
Contributors consented to research only, or consent forms are unavailableRoute 3: re-collectCan the protocol be reused without copying restricted materials?Prototype artifacts leak into the new pipeline
Prototype proved a data type, not a specific corpusRoute 4: licensed operational dataDo the supplier's ownership and consent records support training?Personal data removal not documented

Illustrative example: invented to show structure; it does not describe an available dataset.

Subject: Commercial license request – [corpus name, version, DOI]

1. Current use: research license accepted [date] by [named researcher, institution].
2. Prototype scope: [ASR / SFT / NLP classifier / eval]; checkpoints trained: [count, dates].
3. Requested grant: training, fine-tuning and evaluation of models deployed in
   [products, regions]; retention and commercialization of prototype models.
4. Derivatives: tokenizers, embeddings, synthetic or pseudo-labeled data.
5. Contributor basis: please share consent form version and any withdrawal process.
6. Authority: confirm you can grant commercial rights for all included materials.
7. Term, territory and termination: what happens to trained models if the license ends.

Budget and timeline expectations

Expect commercial upgrades to take longer than the original research download, because someone has to confirm authority, consent scope and pricing. Direct licenses depend on the rights holder's legal and tech-transfer capacity; consortium tiers depend on membership terms; re-collection depends on recruiting and annotation throughput. Pricing structures vary widely, so compare models using the pricing structures guide and run the negotiation checklist before signing.

If you provide a general-purpose AI model on the EU market and sign the GPAI Code of Practice, its copyright chapter commits you to draw up, keep up to date and implement a copyright policy covering your models [13]. Recording how research-only sources were excluded from commercial training is a sensible part of that record; keep the clean-room evidence regardless of jurisdiction.

Finding commercially licensed data after a research-only prototype

SourceX sources operational datasets from US companies on request and manages the commercial process, from assessing data and licensing permissions to agreeing allowed uses in a license. Every dataset is rights-reviewed for ownership and consents, and nothing is contracted until the supplier agrees. Describe the data your prototype needs on the SourceX buyer page.

For more on the cluster, start at the AI data licensing hub or the AI data hub.

Frequently asked questions

Can we keep a model trained on a research-only dataset if we later get a commercial license?

Only if the commercial license says so. Ask for an explicit grant or release covering models trained before the effective date, since licenses that treat trained models as results [1] otherwise leave those checkpoints restricted.

Does a gated download with "commercial use requires contact" mean a license is available?

No. It means the publisher may consider a request. The answer depends on contributor consent and the publisher's authority, and some will decline.

Is evaluation-only use of a research corpus safe for a commercial product?

Often not. Some EULAs define testing commercial systems as commercial use [2], so a benchmark run that gates a product release can breach research-only terms.

Sources

  1. Qualcomm, "Dataset Research License (February 25, 2025)" (2025). https://www.qualcomm.com/content/dam/qcomm-martech/dm-assets/documents/Dataset-Research-License-Feb-25-2025.pdf
  2. Delft University of Technology, "MatchNMingle End User License Agreement". https://matchmakers.ewi.tudelft.nl/matchnmingle/pmwiki/eula.pdf
  3. Hugging Face, "umbra dataset card (README.md)". https://huggingface.co/datasets/fedric95/umbra/blob/51c3525c1102b4cc2425a62c56c4cfc79bd92247/README.md
  4. Linguistic Data Consortium, "LDC For-Profit Membership Agreement". https://Catalog.Ldc.Upenn.Edu/license/ldc-for-profit-membership.pdf
  5. NTT Communication Science Laboratories, "JParaCrawl". https://www.rd.ntt/cs/team_project/icl/lirg/jparacrawl/
  6. ImageNet (Princeton University and Stanford University), "ImageNet download and terms of access". https://image-net.org/download.php
  7. University of Bern / FKI, "Download the IAM Handwriting Database". https://fki.tic.heia-fr.ch/databases/download-the-iam-handwriting-database
  8. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  9. UK Intellectual Property Office, "Copyright, Designs and Patents Act 1988 (consolidated), section 29A". https://assets.publishing.service.gov.uk/media/60180c2b8fa8f53fc62c5897/Copyright-designs-and-patents-act-1988.pdf
  10. UK Intellectual Property Office, "Exceptions to copyright: Research" (2014). https://assets.publishing.service.gov.uk/government/uploads/system/uploads/attachment_data/file/375954/Research.pdf
  11. European Data Protection Board, "Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models" (2024). https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
  12. arXiv, "Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo" (2025). https://arxiv.org/pdf/2501.11003
  13. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data