Data licensing for AI training
Gated and custom-licensed datasets on model hubs: reading the terms before training
Quick answer
A gated dataset's click-through is a contract, not a download formality. The gate controls who gets the files; the attached terms decide whether you may train a commercial model, ship its weights, or keep using it after access is revoked. Before pulling a gated or "license: other" dataset into an SFT or eval pipeline, identify the governing text, confirm who in your company may accept it, check commercial use, model distribution, attribution and termination clauses, and record exactly what was accepted and when.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What a gate actually does, and what it does not
A gate is an access-control mechanism; it does not grant or certify any rights by itself. On Hugging Face, a gated repo shows a prompt such as "You need to agree to share your contact information to access this dataset," and the files become available only after the request is accepted [1]. The same Structured3D page then limits use to non-commercial research and education, which is the part that actually governs training [1].
The dataset card's YAML usually carries the gate configuration (extra_gated_prompt, extra_gated_fields, often a checkbox such as "I agree to use this dataset for non-commercial use only") and a license: field. When that field says other, the real terms live somewhere else: a LICENSE file in the repo, a paragraph in the README, a PDF agreement, or an external website. Those terms, not the green "access granted" banner, define your rights.
The gap between what a hub page says and what the data's terms actually are is well documented. The Data Provenance Initiative audited more than 1,800 text datasets and reported license omission rates above 70% and license error rates above 50% on popular hosting sites [2]. Treat the hub's license tag as a lead to verify, not a finding.
Common patterns in gated terms
Most gated research datasets restrict commercial use, and the restriction often reaches further than engineers expect. Four patterns recur across hub pages:
- Restated upstream terms. A gated ImageNet derivative restates the ImageNet terms: non-commercial research and educational use only, with the obligations extending to the researcher's employer [3][4]. A re-packaging does not loosen the original agreement.
- Contact-first commercial clauses. Some gates allow research access but require commercial users, explicitly including those training AI models, to contact the owner before any such use [5]. Accepting the gate does not cover the commercial case.
- Written-permission prohibitions. Challenge and benchmark datasets frequently bar commercial use without explicit written permission [6], which matters if your eval results feed product decisions.
- Custom model-style licenses. Bespoke licenses can reach downstream artifacts. The Llama 3.1 Community License, for example, attaches conditions to using outputs to create, train or fine-tune other AI models, including a naming requirement [7].
Not every custom license is restrictive. The Copernicus Sentinel data licence grants free, full and open access with broad reuse rights, subject to a liability waiver and an attribution requirement [9]. The point is that "custom" means you must read it; it does not signal either answer. For standard licenses such as CC BY or ODC-BY, use the open data license compatibility matrix instead.
Who in your company can accept the terms
Hub access requests are made from a personal account, so the person who clicks becomes the visible party to the terms [1]. Yet many gated agreements, like the ImageNet-style terms above, purport to bind the user's employer as well [3][4].
That creates two failure modes. An engineer with no signing authority may commit the company to non-commercial or attribution obligations it never reviewed. Conversely, a colleague who did not click may pull the same files from a shared bucket, outside any agreement the licensor recognizes.
A workable internal rule separates three cases:
- Engineer may accept alone: terms limited to research or evaluation inside a sandbox, with no commercial model, no redistribution, and no clauses binding the employer beyond confidentiality.
- Counsel or delegated reviewer approves first: any "license: other", any clause naming the employer, indemnities, governing-law or audit clauses, or any intended use in a model that will ship.
- Do not accept; escalate to procurement: commercial training is prohibited or requires owner contact, or the terms require a signed agreement. That path leads to a negotiated license; see research-only datasets: getting commercial rights.
The custom dataset license review checklist
Each gated dataset needs answers to the same ten questions before it reaches a training run. Fill this in per dataset and store it with the run's data manifest.
Illustrative example: invented to show structure; it does not describe an available dataset.
| # | Question | Where to look | Red flag |
|---|---|---|---|
| 1 | What text governs: card license: value, LICENSE file, gate prompt, external agreement? | Dataset card YAML, repo files, extra_gated_prompt | license: other with no linked text |
| 2 | Is commercial use permitted, and is model training named? | Permitted-use clause | "non-commercial research and educational purposes only" |
| 3 | Is training distinguished from evaluation or benchmarking? | Definitions, use clause | Eval allowed only for "the challenge" |
| 4 | May you distribute trained weights, checkpoints or adapters? | Derivatives, redistribution | Weights treated as derivative works |
| 5 | Do terms reach model outputs or downstream models? | Output, naming, attribution clauses | Naming or license-propagation conditions |
| 6 | What attribution or citation is required, and where? | Attribution clause, README | Attribution required in product UI |
| 7 | Who is bound: the user, the employer, or both? | Party definitions | Clause extending obligations to the user's employer |
| 8 | What happens on termination or revocation? | Termination, deletion | Deletion of copies and derived works |
| 9 | Which law and forum govern, and are there audit rights? | Boilerplate | Foreign forum, licensor audit access |
| 10 | Does upstream content carry its own terms? | Data sources section | Scraped or third-party content with unclear rights |
Questions 4, 5 and 8 decide most commercial outcomes. If termination requires deleting "derivative works," ask counsel whether the licensor could argue that includes trained weights; the companion page on what happens to trained models when a data license ends covers that negotiation.
Recording acceptance so it survives an audit
An acceptance you cannot reproduce is close to no acceptance at all. Access approved through a gate can later be withdrawn by the owner [10], and terms pages can change after you clicked, so capture the evidence at the moment of access.
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset_acceptance:
repo_id: example-org/example-gated-corpus
revision: 3f2c9e1 # commit hash pinned in the training config
license_field: other
terms_snapshot: s3://legal-evidence/datasets/example-gated-corpus/terms-2026-10-09.pdf
terms_sha256: 9b1d...e4
accepted_by: hub-user-jdoe # individual account that clicked
accepted_at: 2026-10-09T14:22:00Z
approver: counsel-ticket-LEGAL-4182
permitted_use: internal evaluation only
restrictions: [no commercial training, no redistribution, attribution in paper]
downstream_runs: [eval-suite-v7]
Pin the dataset revision rather than main, because a later commit can change both files and README terms. Link this record to your data lineage so any model card can trace back to the exact terms in force. The broader due diligence steps are in the AI training data due diligence checklist, and the open dataset license audit covers ungated sources.
When terms change after you have access
Assume the terms can move and plan a re-check. A dataset owner can edit the README, swap the LICENSE file, or tighten the gate prompt on a new commit, and an existing grant does not alert you. Your pinned snapshot is your evidence of what applied when you trained.
Regulators have also noticed unilateral terms changes, though from the other direction. FTC staff warned in February 2024 that companies quietly adopting more permissive data practices, such as AI training, through terms-of-service changes could be acting unfairly or deceptively [8]. That staff post is not a rule, but it reinforces a practical point: the version of the terms in force at collection and at use both matter.
Set a review trigger in your data catalog: when a pinned dataset's upstream repo gets a new commit touching README, LICENSE or the gate fields, flag every model trained on it for a terms diff. If the new terms are stricter, counsel decides whether your earlier acceptance still governs.
Gated research data versus a negotiated license
When the checklist turns up "non-commercial," "contact the owner," or "written permission required," the gate has done its job: it tells you that a separate commercial license is needed. Click-through terms rarely include warranties, indemnities or deletion terms you can negotiate; the ImageNet agreement, for instance, disclaims warranties outright [3]. For production SFT or eval data, buyers usually need a bilateral agreement that states records, uses, term and delivery.
At that point the options are to approach the owner, replace the dataset, or source comparable operational data under a negotiated license. The AI data license negotiation checklist and the licensing hub cover what to ask for. If you need operational data from US companies rather than research corpora, you can describe the data you need to SourceX: it sources such datasets on request, with each release approved by the supplying company, and a request does not guarantee a match.
Sourcing licensed data instead of a gated dataset
SourceX sources operational datasets from US companies, such as support and sales histories, engineering records and documents, and manages the licensing process. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and nothing is contracted until a supplier agrees. Describe the data you need at sourcex.si/buyers.
Sources
- Hugging Face (dataset page), "Pointcept/structured3d-compressed repository (gated access prompt)". https://www.huggingface.co/datasets/Pointcept/structured3d-compressed/tree/main
- Longpre et al. (Nature Machine Intelligence 6, 2024), "A large-scale audit of dataset licensing and attribution in AI". https://www.nature.com/articles/s42256-024-00878-8
- ImageNet (Princeton University and Stanford University), "ImageNet download and terms of access". https://image-net.org/download.php
- Hugging Face (dataset page), "timm/imagenet-12k-wds dataset repository". https://huggingface.co/datasets/timm/imagenet-12k-wds/tree/main
- Hugging Face (dataset page), "amphion/Emilia-Dataset repository". https://huggingface.co/datasets/amphion/Emilia-Dataset/tree/main/Emilia
- Hugging Face (dataset page), "ODELIA-AI/ODELIA-Challenge-2025 repository". https://huggingface.co/datasets/ODELIA-AI/ODELIA-Challenge-2025/tree/main/example-algorithm
- Meta, via Hugging Face, "Llama 3.1 Community License (LICENSE file)". https://huggingface.co/Mozilla/Meta-Llama-3.1-8B-Instruct-llamafile/blob/7d8b93e61bb828c11a25085aefc28ebf9a7952f2/LICENSE
- Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
- Copernicus / European Commission, "Copernicus Sentinel data licence (rev. 1)". https://ewds.climate.copernicus.eu/licences/ec-sentinel
- Luccioni et al. (2022), "A Framework for Deprecating Datasets: Standardizing Documentation, Identification, and Communication". https://huggingface.co/papers/2111.04424
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.