Data licensing for AI training
Prohibited-use clauses in AI data licenses: competitor, surveillance and sensitive-use bans
Quick answer
A prohibited-use clause is the negative list in a data license: the things you may never do with the records, the trained model or its outputs, even inside the licensed field. Typical bans cover re-identification, resale or redistribution, training models that compete with the licensor, surveillance, biometric identification, and releasing weights openly. Counsel's job is to make each ban objective, scoped to the right artifact (data, model, output) and survivable for the products your company actually ships.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why prohibited-use lists now decide deal value
Prohibited-use lists decide deal value because they bind the model you ship, not just the files you store. In November 2025 Fastcase sued Alexi Technologies, alleging that licensed data was used to train and power a commercial AI legal research product outside the license; Proskauer's analysis of the dispute recommends naming pre-training, fine-tuning, RAG, inference, weights and derived data explicitly rather than relying on general "internal use" language [1]. Licensors have responded with standalone AI policies: Nasdaq's Data AI Policy bars use of Nasdaq Information in open-source AI models without a written agreement and prohibits distributing derived data [2], and the Roper Center publishes an AI policy governing how its archive data may be put into AI tools [3].
The practical consequence is that a prohibited-use list drafted for a 2015 analytics license can quietly forbid a 2026 product. A ban on "derived databases" may reach embeddings in a vector store; a ban on "redistribution" may reach a model served through a public API. Read the negative list against your architecture, not against the data alone. The positive grant is covered in our guide to field-of-use restrictions; this page covers what sits on the other side of the line.
The seven bans counsel sees most often
The bans that recur in AI data licenses fall into seven families, and each needs a different redline. The table below maps each family to the artifact it really constrains and the failure mode to test.
| Ban family | What licensors usually intend | Artifact actually bound | Failure mode for buyers |
|---|---|---|---|
| Competitor training | Stop the data training a rival to the licensor's own product | Model and downstream customers | "Competitor" undefined, so any general-purpose model qualifies |
| Resale and redistribution | Stop the raw records being resold | Data, sometimes outputs | Outputs or weights treated as "redistribution" |
| Re-identification | Stop linkage back to people or companies | Data and joins with other data | Ban drafted so broadly it blocks entity resolution you need for dedup |
| Surveillance and tracking | Stop monitoring of individuals or employees | End use of model | No definition, so fraud detection or security analytics is caught |
| Biometric identification | Stop face, voice or gait ID | Model capability | Speaker diarization or liveness checks read as "identification" |
| Open release | Stop data leaking via public weights or datasets | Weights, checkpoints, synthetic data | Ban extends to research papers or eval sets |
| Benchmark publication | Stop unflattering public comparisons | Results and reports | DeWitt-style consent rights over your own model cards [8] |
For re-identification bans specifically, see re-identification prohibition clauses, which covers the language buyers are asked to sign and the joins it should still allow.
Defining "competitor" so the ban can be applied
A competitor ban is only workable when "competitor" is defined by an objective test the two parties can apply on the same facts. Undefined terms such as "any product that competes with Licensor" invite dispute every time either company launches something new, and they tend to expand as the licensor's roadmap grows.
Push for one of three objective structures. First, a named list in a schedule, updated only by mutual written amendment. Second, a product-category test tied to the licensor's revenue lines on the effective date (for example, "a commercially offered legal research database"), frozen at signing. Third, a purpose test that bans using the data to replicate the licensed dataset itself, which captures the real risk without banning general models. Also cap the reach: the ban should bind your own use and named customer agreements, not every downstream user of a general-purpose API.
Whether a supplier's data can train models that compete with someone else is a different question, answered in can licensed data be used to train competitors' models.
Surveillance and biometric bans need technical definitions
Surveillance and biometric bans are reasonable, but they must be defined by function so that legitimate security, fraud and accessibility features survive. Regulators already treat facial recognition for surveillance as a high-risk use: the FTC's Rite Aid action alleged untested facial recognition produced false matches, and the proposed order would ban the retailer from using it for security or surveillance for five years [7]. Licensors of voice, video or workplace data may cite cases like this when they insist on broad bans.
Draft the definitions around capability and purpose. "Biometric identification" should mean one-to-many matching of a person against a reference set to establish identity; it should not capture speaker diarization (who spoke when, without naming), liveness detection, or emotion-agnostic transcription. "Surveillance" should mean persistent monitoring of identified individuals' location, behavior or communications without their knowledge, so that aggregate fraud detection, network security monitoring and consented employee productivity analytics are expressly carved out or handled by a separate approval process.
Open release, derived data and the weights question
Open-release bans are where data licenses and model licenses collide, so decide early whether you will ever publish weights, checkpoints, synthetic datasets or evaluation sets. Nasdaq's policy is a clear example: open-source AI models are off-limits without a separate written agreement [2]. If your roadmap includes an open-weights release, a ban like this means the licensed data must be excluded from that training run, and you need lineage good enough to prove it.
Separate the artifacts in the clause. Weights trained on the data, synthetic data generated by a model trained on the data, and aggregate statistics carry different leakage risk and deserve different treatment. Behavioral-use licenses such as RAIL-D (data) and RAIL-M (model) attach use restrictions to AI artifacts and expect them to travel with the artifact [5]; if you plan to release under such a license, check whether the data licensor's bans are compatible with it. Successor models and distillation are covered in derivative and successor model rights.
Flow-down: making downstream customers carry the ban
A prohibited-use clause is only as strong as the flow-down that pushes it into your customer terms, so negotiate the flow-down obligation as carefully as the ban itself. Morgan Lewis notes that AI contracts commonly push data obligations down to vendors and suppliers and keep them alive after termination [4]. Licensors want the same chain running toward your customers.
Accept a flow-down obligation only in a form you can operate: a duty to include materially similar restrictions in your acceptable use policy, to act on credible reports of breach, and to suspend the offending account. Resist strict liability for every downstream user's conduct and any audit right over your customers. Check that your acceptable use policy already bans the same families (surveillance, biometric ID, re-identification) so you are not maintaining one version per licensor. Contractor and cloud-processor access is a separate clause, covered in affiliates, contractors and cloud processors.
Redline checklist for a licensor's prohibited-use list
Use this checklist when the licensor's paper arrives; it orders the review from highest to lowest product impact.
Illustrative example: invented to show structure; it does not describe an available dataset.
PROHIBITED-USE REVIEW — Dataset: [support-ticket corpus, 2019–2025]
Product uses mapped: pre-training [ ] fine-tuning [x] RAG index [x] eval [x] open weights [ ]
1. Competitor ban
[ ] "Competitor" defined by schedule, frozen category, or replication test
[ ] Applies to Licensee and named customers only, not all API users
2. Redistribution
[ ] Model outputs and weights expressly NOT "redistribution of Data"
[ ] Embeddings in internal vector store expressly permitted
3. Re-identification
[ ] Ban limited to persons/entities; dedup and entity resolution allowed
4. Surveillance / biometric ID
[ ] Functional definitions; diarization, liveness, fraud detection carved out
5. Open release
[ ] Weights, synthetic data and eval sets addressed separately
[ ] Research publication of aggregate metrics allowed
6. Benchmark publication
[ ] No vendor consent right over Licensee's own model results
7. Flow-down
[ ] "Materially similar" AUP terms; no strict liability for end users
8. Survival and remedy
[ ] Bans survive only as to models trained during term; cure period stated
Expect the provenance chain behind the list to be imperfect. The Data Provenance Initiative found license information missing for more than 70% and wrong for more than 50% of popular datasets on hosting sites [9], which is why a bespoke license with an explicit negative list is often safer than an inherited click-through. For warranties that back the list, see data warranties for AI training licenses.
When to accept a broad ban instead of fighting it
Accept a broad ban when the dataset itself makes the prohibited use implausible, because negotiating capital is better spent elsewhere. A healthcare contracting white paper makes the same point about field-of-use terms: the data often limits what can realistically be done, so aggressive negotiation can be worth less than parties assume [6]. A ban on biometric identification in a corpus of de-identified invoices costs you nothing.
Fight the ban when it reaches an artifact on your roadmap (weights, a public API, a customer-facing RAG feature) or when its trigger is subjective. Trade where you can: accept a wider surveillance ban in exchange for a narrow, scheduled competitor definition, or accept open-release limits in exchange for explicit permission to publish evaluation results. Track every accepted ban in your model registry so release engineering knows which runs carry which restrictions. If you would rather describe the data and required uses up front so a supplier can respond to them, you can submit a data request as a buyer.
Sourcing data with clear prohibited-use terms through SourceX
SourceX sources operational datasets from US companies on request and manages the commercial process, from assessing data and licensing permissions to agreeing pricing and allowed uses in a license; every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Nothing is contracted until a supplier agrees, and a request does not guarantee a match. To describe the data you need and the uses your product requires, start a buyer request.
Related reading
- AI training data licensing: the buyer's guide
- Field of use (glossary)
- What buyers are allowed to do with licensed data
- Internal-use-only data licenses
Sources
- Proskauer Rose LLP, "Data License Restrictions in the AI Spotlight: Careful Drafting Is More Important Than Ever". https://www.proskauer.com/blog/data-license-restrictions-in-the-ai-spotlight-careful-drafting-is-more-important-than-ever
- Nasdaq, "Data AI Policy". https://nasdaqtrader.com/content/AdministrationSupport/AgreementsData/Data_AI_Policy.pdf
- Roper Center for Public Opinion Research, Cornell University, "Roper Center Artificial Intelligence (AI) Policy". https://ropercenter.cornell.edu/policies/roper-center-artificial-intelligence-ai-policy
- Morgan Lewis, "Key Concepts in AI Contracting: Data Rights and Restrictions" (2025). https://www.morganlewis.com/blogs/sourcingatmorganlewis/2025/12/key-concepts-in-ai-contracting-data-rights-and-restrictions
- Responsible AI Licenses (RAIL), "From RAIL to Open RAIL: Topologies of RAIL Licenses" (2022). https://www.licenses.ai/blog/2022/8/18/naming-convention-of-responsible-ai-licenses
- Foley Hoag / Alliance for Artificial Intelligence in Healthcare, "Critical Considerations for AI Model Licensing Agreements in Healthcare (AAIH White Paper)". https://www.foleyhoag.com/getmedia/c7dfcada-a731-41e8-b7a8-a35b47faf77c/Critical-Considerations-for-AI-Model-Licensing-Agreements-in-Healthcare-AAIH-White-Paper.pdf
- Federal Trade Commission, "Coming face to face with Rite Aid's allegedly unfair use of facial recognition technology" (2023). https://www.ftc.gov/business-guidance/blog/2023/12/coming-face-face-rite-aids-allegedly-unfair-use-facial-recognition-technology
- David A. Wheeler, "The DeWitt clause's censorship should be illegal". https://dwheeler.com/essays/dewitt-clause.html
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.