Fine-tuning and post-training data
Open instruction and preference datasets: which allow commercial fine-tuning
Quick answer
Some instruction tuning datasets allow commercial use, but the license tag answers only part of the question. Human-written or human-annotated sets under licenses that allow commercial use, such as databricks-dolly-15k (CC BY-SA 3.0) and HelpSteer2 (CC-BY-4.0), are published for commercial use, while Stanford Alpaca, licensed CC BY-NC 4.0, is not [1][2][3]. In between sit permissively tagged sets whose responses or preference labels came from a proprietary model, where that provider's terms can still restrict commercial use [2][14]. Check the license, the generator and every upstream subset before you train.
By SourceX Editorial · Updated
This guide is for teams doing instruction tuning and preference optimization on open data. It sits in the fine-tuning and post-training data guide; for license families across all data types, use the open data license compatibility matrix.
Three layers decide whether an SFT or preference set is commercially usable
An open dataset is usable for commercial fine-tuning only when three layers all permit it: the publisher's license, the terms that bound whoever generated the responses or labels, and the licenses of the upstream datasets its rows came from. Most screening stops at the first layer, so the other two surface late, often during model release review.
- The publisher's license. This is the grant on the compiled dataset: Apache-2.0, MIT, CC BY, CC BY-SA, CC BY-NC or a custom research-only text. Data licenses differ in how far they reach. CDLA-Permissive-2.0, for example, defines "Results" to include machine learning models and places no obligations on them [4]. Creative Commons says its licenses apply only where copyright permission is needed, and its deliberately conservative reading treats NonCommercial as covering every stage from copying during training to distributing the model [5].
- Generator and annotator terms. If a model wrote the responses, or an LLM judge picked the preferred answer, the provider's terms of use or the open-weight model's license governed that generation. These are mostly contract terms. Lemley and Henderson argue that model output lacks the human authorship copyright needs, so output restrictions rest on contract; commentators also question whether such terms bind someone who never accepted them [6].
- Upstream sources. Prompts are often borrowed from other datasets, and collections repackage dozens of subsets. Each subset keeps its own terms, so licenses have to be checked against the original repositories [7].
A permissive tag describes what the publisher grants. It cannot remove restrictions that bound the publisher when the data was made. A July 2026 law-firm diligence note warns that generator restrictions can leave a synthetic dataset unusable for its intended purpose [8].
Where well-known open SFT and preference sets land
Among well-known open sets, human-written or human-annotated data under CC BY-SA or CC BY is commercially usable with conditions, CC BY-NC data is not, and permissively tagged model-generated data depends on the generator's terms. The readings below reflect the cited sources as of October 2026. Cards change, so re-read the card at the exact revision you download.
| Dataset | Used for | License as published | Where responses and labels came from | Reading for commercial fine-tuning |
|---|---|---|---|---|
| databricks-dolly-15k | SFT | CC BY-SA 3.0 [1][9] | Written by Databricks employees [1] | Published for research and commercial use [1]; ShareAlike conditions apply when you share the data or adaptations of it |
| Stanford Alpaca (52K) | SFT | CC BY-NC 4.0 [3][9] | Generated with OpenAI's text-davinci-003 using Self-Instruct [3] | Not usable: the README says models trained on it should not be used outside research [3] |
| HelpSteer2 | Reward modeling, preference tuning | CC-BY-4.0 [2] | About 10,000 human-annotated response pairs, built without distilling from proprietary models [2] | Usable with attribution, which is the authors' stated aim [2] |
| UltraFeedback | DPO, reward modeling | Permissive tag on the release; read the current card | Responses from several models, with preference ratings by GPT-4 [2] | The HelpSteer2 authors note that GPT-4-distilled preference sets of this kind are often restricted to academic or non-commercial use [2] |
| ShareGPT 52K, LMSYS-Chat-1M | Multi-turn SFT | A third-party catalog lists ShareGPT 52K as CC BY-NC 4.0 and LMSYS-Chat-1M as non-commercial [10] | Real user conversations with chat assistants [10] | Treat as non-commercial; confirm on the original card |
| NVIDIA OpenCodeInstruct | Code SFT | Card marks it ready for commercial and non-commercial use [11] | Read the card's generation details | Possibly usable: the card permits commercial use but makes you responsible for checking the license fits your purpose [11]; check the generator models it names |
| NusaX-senti-LexC-Gen | Synthetic classification data | Read the card's tag; the card says its generator's BigScience RAIL License applies downstream [12] | Generated with BLOOMZ models [12] | Use restrictions travel to classifiers fine-tuned on it [12] |
Tag counts across the field look reassuring. A survey of 103 instruction fine-tuning datasets found Apache-2.0 the most common license (43 datasets), followed by GPL-3.0 and MIT [13]. But an Apache-2.0 dataset of distilled GPT-4 completions still raises the generator question, so an Apache-licensed instruction dataset is a lead, not a clearance. For the NonCommercial rows, CC BY-NC datasets and commercial training covers what counts as commercial use and whether NC terms follow the weights.
Model-written responses and AI-judged preferences carry provider terms
When a proprietary model wrote the responses or chose the preferred answer, its provider's terms become the first constraint to check, and many prohibit using outputs to build competing models. As of October 2026, Anthropic's help center states that its terms do not allow outputs to train models competitive with its own, giving general-purpose chatbots and models designed for open-ended text generation as examples [14]. The same page lists sentiment analysis, content categorization, summarization and information extraction tools as non-competing uses [14]. The HelpSteer2 authors describe preference data distilled from GPT-4 as carrying commercial-use restrictions imposed by model providers [2].
Preference data is exposed twice. In UltraFeedback-style sets, both the candidate responses and the ratings that decide "chosen" versus "rejected" are model outputs [2]. Direct Preference Optimization (DPO) fits the policy directly to those pairs without a separate reward model [15], so a DPO run on GPT-4-judged pairs is shaped by GPT-4's judgments even when every response came from an open model. The trade-offs of judge-labeled data are covered in AI feedback versus human preference data, and pair structure in preference datasets for DPO.
Open-weight generators are not automatically cleaner. The LexC-Gen card states that because its data came from BLOOMZ under the BigScience RAIL License, that license would apply to classifiers fine-tuned on it [12]. Mayer Brown advises asking which model generated a synthetic dataset and whether that model's license or acceptable use policy restricts commercial use of its outputs [8].
For each model-generated subset, record the generator model and version, the generation dates, an archived copy of the terms in force then, and whether your model would compete with the provider's products. Counsel then decides whether those terms bind you. When the synthetic data is bought rather than downloaded, apply the due diligence for purchased synthetic fine-tuning data.
Reading a Hugging Face dataset card for license evidence
The license shown on a Hugging Face dataset page is a metadata tag the uploader entered, so treat it as a lead to verify, not as the license itself. The Hub renders each repository's README.md as the dataset card and reads the license from a YAML block at the top, which also drives the license filter on huggingface.co/datasets [16]. Nothing in that mechanism checks the tag against the source terms.
Audits show how often tags are wrong. The Data Provenance Initiative traced more than 1,800 text datasets and reports license omission above 70% and error rates above 50% on popular hosting sites [17]. It also found Hugging Face tags often in a different use category from the original license, frequently a more permissive one, and non-commercial terms dominating newer synthetic data [17]. Commit histories on the Hub record license tags on Alpaca and an Alpaca-derived collection being changed to CC BY-NC after publication, in one case because the data contains OpenAI model outputs [18][19].
A dataset license check on Hugging Face reads these, in order:
- The YAML header's
licensevalue, and anyconfigsentries that split the data into subsets. - The card's licensing and source-data sections, which may contradict the tag in prose ("for research purposes only").
- Any LICENSE or DATA_LICENSE file, and the terms you accept to open a gated dataset.
- The paper or GitHub README, which usually names the generator model and the prompt sources.
- The commit history, to see whether the license changed and which revision you downloaded.
The Data Provenance Explorer, released with the audit, filters fine-tuning collections by traced license and source [17]. For a method that covers images, audio and code as well, see auditing open dataset licenses before commercial training.
Mixtures and collections: clear each subset, not the wrapper
A collection that bundles instruction or preference subsets is only as usable as the most restrictive subset you keep. The Agent Data Protocol authors, who unify many agent fine-tuning datasets, list per-dataset licenses spanning Apache 2.0, MIT, CC BY 4.0 and CC BY-SA 4.0, and note that some datasets restrict commercial use or redistribution [7]. The Data Provenance Initiative built its audit around the same problem, tracing the component datasets of widely used fine-tuning collections back to their original sources and licenses [17].
The practical routine is to build an allow-list of subsets, filter rows on the mixture's per-row source column before tokenization, and save the kept row IDs with the training config. A mixture with no per-row source field cannot be cleared this way. Keep prompt-only reuse in scope too: a prompt set lifted from a non-commercial dataset still comes from that dataset, even when you generate new responses.
A license register entry for every dataset revision you train on
Keep one register entry per dataset revision, linked to every checkpoint trained on it, so release reviewers can answer "what was this model tuned on, and under which terms" without redoing the research. The ML engineer fills in revision and origin fields, counsel signs the decision and its conditions, and the release owner links checkpoints. Record licenses as SPDX identifiers such as CC-BY-SA-3.0 or CDLA-Permissive-2.0 so tooling can read them [4].
Disclosure duties make the register more than housekeeping. Since 1 January 2026, California's AB 2013 has required developers of generative AI systems offered to Californians to post training-data documentation that includes a high-level summary of the datasets, and to do so again before each later release or substantial modification [20]. Whether a fine-tune makes you a developer under that law, or a provider under the EU AI Act, is covered in provider and developer duties when fine-tuning with acquired data.
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset: example-org/support-instruct-mix
revision: 3f9c2e1 # commit actually downloaded
license_tag: apache-2.0 # Hub YAML header
license_file: "LICENSE @ 3f9c2e1" # note any conflict with the tag
spdx_ids: [Apache-2.0, CC-BY-SA-4.0]
subsets:
- name: human_written_qa
rows_kept: 41210
response_origin: human annotators (per card section "Source data")
upstream_prompts: none
decision: allowed
- name: distilled_chat
rows_kept: 0
response_origin: model-generated (proprietary chat model, version per paper)
generator_terms: archived 2026-09-30; competing-model clause present
decision: excluded pending counsel review
preference_labels:
origin: not applicable (SFT only)
obligations: [attribution in model card, ShareAlike review before sharing subsets]
reviewed_by: [ml-lead, product-counsel]
review_date: 2026-10-09
checkpoints: [sft-v3-2026-10-12]
Red flags that should stop an open dataset before training
These signals mean a dataset needs counsel review or a replacement before it enters an SFT or DPO run:
- Responses in a commercial assistant's voice, or a dataset name that includes a proprietary model's name, with no generator disclosure on the card.
- A permissive tag beside prose limiting use to research, or a README that restricts models trained on the data, as Alpaca's does [3].
- A license value of
unknown,otherwith no linked text, or a tag that changed in the commit history. - Conversations collected from users' shared chatbot sessions, such as ShareGPT-derived sets [10].
- Preference labels from an unnamed "LLM judge".
- A mixture with no per-row source field.
- Gated access whose click-through terms differ from the tag.
Before you buy anything, the fine-tuning dataset evaluation guide covers quality checks that sit alongside these license checks.
When open data stops covering your fine-tuning target
Open instruction and preference sets cover general assistant behavior, but domain-specific SFT and DPO data built from real support resolutions, claims decisions or contract redlines rarely exists openly with clear commercial terms. The options are to write it in-house, commission collection, or license records that businesses already hold and convert them, as described in turning business records into instruction-response pairs. A license written for the purpose, such as a fine-tuning-only data license, states the permitted uses rather than leaving you to infer them from a tag; the general trade-offs are in public versus proprietary data.
SourceX sources operational datasets from US companies on request, rather than holding them in stock, and manages the commercial process, including the licensing agreement. Each dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place. It is then delivered under a license that defines the included records, their permitted uses, the license term and delivery; the clauses to look for are explained in AI data license terms. Buyers can describe the instruction or preference data they need without finding or approaching companies themselves.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Know which fine-tuning uses you need licensed?
Describe the records, tasks and uses you need licensed, from SFT pairs to preference judgments. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license and delivery. Submit your licensing requirements.
Sources
- Databricks, "Free Dolly: Introducing the World's First Truly Open Instruction-Tuned LLM (German-locale page)" (2023). https://www.databricks.com/de/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm
- Wang et al. (NVIDIA), "HelpSteer2: Open-source dataset for training top-performing reward models" (arXiv 2024). https://arxiv.org/pdf/2406.08673
- Taori et al., Stanford (mirrored on git.hackliberty.org), "stanford_alpaca README (commit on a mirror of the tatsu-lab/stanford_alpaca repository)" (2023). https://git.hackliberty.org/AI/stanford_alpaca/commit/65512697dc67779a6e53c267488aba0ec4d7c02a
- SPDX (Linux Foundation), "Community Data License Agreement Permissive 2.0 (SPDX License List)". https://spdx.org/licenses/CDLA-Permissive-2.0.html
- Creative Commons, "Using CC-licensed works for AI training". https://creativecommons.org/using-cc-licensed-works-for-ai-training-2/
- SpicyIP, "Discussing Lemley and Henderson's 'The Mirage of Artificial Intelligence Terms of Use Restrictions'" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
- arXiv:2510.24702, "Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents" (2025). https://arxiv.org/pdf/2510.24702
- Mayer Brown, "Synthetic Data as a Deal Asset: Ownership, Provenance and Diligence Considerations in AI Acquisitions" (2026). https://www.mayerbrown.com/en/insights/publications/2026/07/synthetic-data-as-a-deal-asset-ownership-provenance-and-diligence-considerations-in-ai-acquisitions
- arXiv:2505.15656, "Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!" (2025). https://arxiv.org/pdf/2505.15656
- LLM Configurator (third-party dataset catalog), "ShareGPT 52K: LLM Instruction / SFT Dataset". https://llmconfigurator.com/en/datasets/sharegpt
- NVIDIA (Hugging Face Hub), "OpenCodeInstruct dataset card (README.md)". https://huggingface.co/datasets/nvidia/OpenCodeInstruct/blob/eea66e883f681c64ff09cffbe82d73843fdfdaed/README.md
- BatsResearch (Hugging Face Hub), "BatsResearch/NusaX-senti-LexC-Gen dataset card (commit 260c323)". https://huggingface.co/datasets/BatsResearch/NusaX-senti-LexC-Gen/commit/260c3230882c4337c063ca2458bb67ba74667fa4
- Liu et al., "Datasets for Large Language Models: A Comprehensive Survey" (arXiv 2024). https://arxiv.org/pdf/2402.18041
- Anthropic Help Center, "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- Rafailov et al. (Stanford), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (arXiv 2023; NeurIPS 2023). https://arxiv.org/abs/2305.18290v1
- Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- Hugging Face Hub, "tatsu-lab/alpaca dataset repository commit b87f0d2". https://huggingface.co/datasets/tatsu-lab/alpaca/commit/b87f0d2ef96f6a01ca53a647e456287cacee75c0
- Hugging Face Hub, "QingyiSi/Alpaca-CoT dataset repository commit d665273". https://huggingface.co/datasets/QingyiSi/Alpaca-CoT/commit/d665273c8e52b35ddbc1eee85c91e1b04d3930bd
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.