Provenance, rights and permitted use
Tracing Upstream Licenses in Aggregated and Derived Datasets
Quick answer
A dataset collection does not have one license; it has the license of its packaging plus every license of every component, and any model-output terms behind generated rows. To trace inheritance, explode the collection into components, resolve each to its original source and terms, classify each restriction (commercial use, share-alike, attribution, model-output limits), and apply the most restrictive term to any mixture that includes that component. Record per-component IDs so you can drop a component later without rebuilding the whole audit.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why a collection's headline license is not the answer
The license printed on a collection usually covers the compilation, not the rights in its parts. The Data Provenance Initiative audited 44 instruction-tuning collections built from 1,858 component datasets and found that license information on popular hosting sites was frequently missing or wrong, with omission above 70% and error rates above 50% [1]. An earlier analysis reached the same structural conclusion: datasets are often assembled from multiple sources with different licenses, and dataset licenses are written ad hoc rather than drawn from a small set of standardized licenses the way software licenses are [2].
The practical failure mode is a permissive label on top of restrictive parts. A collection repository tagged apache-2.0 can contain a component released under CC BY-NC 4.0, such as Alpaca, alongside one under CC BY-SA 3.0, such as Dolly [3]. If your SFT mixture inherits that component, the Apache tag on the wrapper does nothing to lift the non-commercial term. For the portfolio-level process of reviewing open datasets before commercial use, see auditing open dataset licenses before commercial training; this page covers the recursion inside one compilation.
The four layers you have to resolve
Each row in an aggregated dataset sits under up to four sets of terms, and you need all of them. Treat them as separate fields, not one "license" string.
- Compilation layer. The license or terms of use of the collection repository itself, including click-through gating. The Stack v2, for example, is distributed with its own terms of use and access gating on top of the licenses of the underlying source files [9].
- Component layer. The license the original dataset authors attached when they released the component, which may differ from what the collection's card restates [1].
- Underlying content layer. The rights in the raw text the component was built from: web pages, forum posts, code files, books, transcripts. A component author cannot grant more than they received. Tracing this layer for transcripts, translations and summaries is covered in derived records rights tracing.
- Generator layer. For synthetic or model-written rows, the terms of the model or API that produced them. LLaVA's instruction data, for instance, was produced by prompting language-only GPT-4 with image-text pairs [4], so any analysis of rows like these has to account for the generating service's output terms as they stood when the rows were generated. See provenance records for synthetic training data for what to capture.
How to trace a collection, step by step
Tracing works as a depth-first walk from the collection to every leaf source, recorded as a table rather than prose. Plan for the walk to stall at some leaves; an unresolved leaf is a finding, not a gap to paper over.
- Pin the version. Record the collection's repository ID and commit hash or release tag. Collection cards change, and components get added or removed between versions.
- Explode the components. Read the build scripts, not just the card. Mixture configs,
datasetsloading scripts and task registries reveal components the README omits. - Resolve each component to its origin. Follow it to the original paper, repository or release page, and record the license exactly as published there, with a URL and retrieval date. Use the hosting site's
licensemetadata only as a pointer, given the audit's error rates [1]. - Descend one more level where content was reused. If a component was built from Reddit, StackExchange, news, books or GitHub, record that upstream source and its terms separately.
- Flag generated rows. Note the generating model, provider and approximate generation date for any component written by a model.
- Classify restrictions. Normalize each license into flags: commercial use allowed, share-alike, attribution required, no-derivatives, model-output restrictions, custom or unknown. Element-by-element handling of BY, SA, NC and ND is in Creative Commons licenses and AI training.
- Propagate. For each training mixture, compute the union of restriction flags across included components. The mixture's effective terms are the most restrictive terms of anything in it.
- Decide per component. Keep, drop, replace or escalate to counsel. Removing one non-commercial component is usually cheaper than relicensing a whole model.
Propagation rules for mixed-license mixtures
The safe default is that restrictions accumulate and permissions do not. A mixture is only as usable as its most restrictive included component, so a single NC component makes the mixture non-commercial for planning purposes until counsel says otherwise.
A few rules of thumb keep the propagation table honest:
- Unknown is restrictive. A component whose license you could not find gets the "custom or unknown" flag and is treated as excluded from commercial mixtures.
- Share-alike is an open question for weights. Whether ShareAlike obligations reach a trained model is unsettled; record SA as its own flag and let counsel decide rather than collapsing it into "permissive."
- Attribution needs a carrier. If several components require attribution, decide where it lives (model card, documentation, NOTICE file) and list each attributed source.
- Filtering does not launder terms. Quality filtering, deduplication or reformatting a component into chat format changes the rows, not the rights. LIMA-style curation still inherits the terms of whatever it curated; see filtering instruction-tuning data for quality.
- Evaluation splits count. An eval set built from a restricted component still carries those terms, and test-set contamination makes it more likely the component influenced training.
A component-level provenance card you can adopt
A provenance card is a per-component record that sits beside the collection card and travels with each training mixture. Existing documentation formats give you the slots: Hugging Face dataset cards hold YAML metadata such as license [5], Croissant describes datasets and their file resources in schema.org-based JSON-LD [6], Data Cards call for upstream sources and collection methods [7], and Datasheets for Datasets cover composition and collection process [8]. The Data Provenance Initiative also released tooling to trace lineage and generate provenance cards for the collections it audited [1].
The record below extends those formats with the fields that propagation needs. Keep one row per component, keyed by a stable component ID, and reference those IDs from your AI training data register.
Illustrative example: invented to show structure; it does not describe an available dataset.
collection:
id: example-org/instruct-mix
version: "v2.1" # commit hash or release tag
compilation_license: apache-2.0
compilation_terms_url: https://example.org/instruct-mix/terms
components:
- component_id: IM-014
name: example-qa-pairs
origin_url: https://example.org/qa-pairs/release
license_as_published: CC-BY-NC-4.0
license_on_collection_card: apache-2.0 # mismatch -> finding
underlying_sources: [community-forum-dump-2022]
underlying_terms: "forum ToS; CC BY-SA per post"
generated: false
flags: {commercial: false, share_alike: true, attribution: true, no_derivs: false, output_terms: false, unknown: false}
rows_in_mix: 41230
decision: drop_from_commercial_mix
reviewed_by: j.doe
reviewed_on: 2026-10-01
- component_id: IM-022
name: example-synthetic-dialogs
origin_url: https://example.org/synthetic-dialogs
license_as_published: MIT
generated: true
generator: {model: "example-llm-v4", provider: "Example AI", generated_on: "2024-03"}
flags: {commercial: true, share_alike: false, attribution: true, no_derivs: false, output_terms: true, unknown: false}
decision: escalate_counsel
mixture_effective_terms:
mix_id: sft-2026-10-a
includes: [IM-022]
effective_flags: {commercial: true, attribution: true, output_terms: true}
The mismatch between license_as_published and license_on_collection_card is the single most useful field in the record, because it surfaces exactly the error pattern the large-scale audit found [1]. Record-level detail, when a component mixes rows with different terms, belongs in record-level provenance.
Common failure modes when inheriting a collection
Most license problems in fine-tuning mixtures come from a few repeatable mistakes. Check for each before a mixture is frozen.
| Failure mode | What it looks like | Check |
|---|---|---|
| Wrapper license trusted | Collection tagged permissive; NC component inside | Compare each component's published license to the card [1][3] |
| Stale card | Card lists 60 components; build script loads 64 | Explode components from code, pinned to a commit |
| Lost generator terms | Synthetic component listed as MIT with no generator noted | Record model, provider and generation date [4] |
| Gating terms ignored | Click-through terms accepted by one engineer, not recorded | Store accepted terms of use with the version [9] |
| Eval leakage | Restricted eval set also present in SFT mix | Hash-match eval and train rows by component ID |
| No exit path | Cannot identify which models used a component | Link component IDs to dataset-to-model traceability |
When to replace components with licensed data
Replacing a restricted or unresolved component with data licensed directly for training is often simpler than defending it. This applies most when the component is high-value for your domain, when its underlying content layer cannot be resolved, or when its generator terms conflict with your intended use. For open alternatives already screened for commercial fine-tuning, start with open instruction and preference datasets for commercial use.
Directly licensed operational records can shorten the trace to a single documented hop, though the underlying content layer still needs review. SourceX sources operational datasets, such as support and sales histories, engineering records, documents, and finance and legal workflows, from US companies on request, and each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Diligence materials on source, rights, preparation and allowed use are prepared per dataset, which can give your provenance card a single component row with documented terms. Buyers can describe the data they need at SourceX for AI data buyers.
For definitions behind this page, see the glossary entries on data lineage and instruction tuning, and the cluster overview in data provenance for AI training data.
Sourcing fine-tuning data with a short license chain
If your mixture audit shows components you cannot trace, SourceX can look for US businesses that hold the operational records you describe and manage the licensing process, with every release approved by the supplying company. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and nothing is contracted until a supplier agrees. Describe the data your SFT or eval work needs at https://sourcex.si/buyers.
Sources
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- arXiv (Rajbahadur et al.), "Can I use this publicly available dataset to build commercial AI software? Most likely not" (2021). https://arxiv.org/abs/2111.02374v4
- arXiv, "Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!" (2025). https://arxiv.org/pdf/2505.15656
- arXiv (Liu, Li, Wu, Lee), "Visual Instruction Tuning" (2023). https://arxiv.org/pdf/2304.08485
- Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
- arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- arXiv (Pushkarna, Zaldivar, Kjartansson), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- arXiv (Gebru et al.), "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
- Hugging Face (BigCode), "bigcode/the-stack-v2-dedup". https://huggingface.co/datasets/bigcode/the-stack-v2-dedup
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.