Skip to content

Fine-tuning and post-training data

Due diligence for purchased synthetic fine-tuning data: generator terms, seed rights, quality and contract

Quick answer

A synthetic fine-tuning data license is only as sound as the inputs behind each generated record. Before buying synthetic SFT (supervised fine-tuning) or preference data, establish which models generated and judged it, under which terms and on what dates; which seed prompts or documents conditioned it and whether the vendor could use them; how outputs were filtered, deduplicated and decontaminated; and whether the data lifts your model on real held-out prompts. Then put warranties for generator terms and seed rights into the license.

By SourceX Editorial · Updated

This is the vendor-diligence step for synthetic data in the fine-tuning and post-training data guide. For data-type trade-offs, see licensed vs synthetic vs scraped training data; for any supplier, the AI training data due diligence checklist.

Four layers of a synthetic record, and where each one fails

Every synthetic SFT or preference record has up to four inputs: a seed, a generator model, a judge or filter model, and human edits or ratings. Each carries its own rights and quality risk, so diligence covers all four, not only the license on the dataset card.

LayerWhat it isRights questionQuality question
SeedPrompt pools, source documents, few-shot examples, personasCould the vendor use it for generation and resale? Any personal data?Does it match your task distribution, or a benchmark's?
GeneratorModel that wrote responses, or the chosen and rejected candidatesDo its terms allow outputs to train your model? Do conditions attach to that model?Correctness, refusals, verbosity, repeated phrasing
Judge or filterModel or rubric that scored, ranked or discarded outputsThe same terms questions as the generatorPosition and length bias; where the cut-off sits
Human layerWriters who edited outputs; raters who labeled pairsDid contributors assign or license their work? Did they use AI tools?Agreement rates, reviewer qualifications

In preference data a model often acts as the labeler. The RLAIF study by Lee et al. (2023) found that preferences labeled by an off-the-shelf LLM gave results comparable to human labels on summarization and dialogue [2], so the judge's identity and terms matter as much as the generator's (AI feedback vs human preference data). A July 2026 Mayer Brown article on AI acquisitions advises inventorying synthetic, licensed real-world and scraped data separately, because each carries different ownership and infringement risk [1].

The diligence sequence: who asks what, and in which order

Run diligence in a fixed order so legal review starts from facts. Each step has an owner and a document that closes it.

StepOwnerWhat to obtainExit condition
1. Intended usePost-training lead, counselUse statement: SFT, DPO (direct preference optimization), reward model or distillation; internal, API or released weights; whether your model competes with a generator's providerApproved before the first vendor call
2. LineageProcurementGeneration manifest per batch: models, versions, access route, terms version, dates, seed sources, filtersEvery record maps to a batch
3. Terms reviewCounselGenerator and judge terms in force on the generation datesEach term fits the use statement
4. Seed and privacy reviewCounsel, privacySeed inventory with licenses; personal-data (PII) scan resultsEvery seed source has a documented right to generate
5. Sample verificationML engineerRandom sample you draw, sized with sample sizes for estimating a dataset's error rateDuplicate, overlap and error rates within thresholds fixed in advance
6. PilotPost-training leadAblation on your held-out real promptsLift over baseline without regressions
7. ContractCounsel, procurementGrant, warranties, indemnity, replacement, disclosure cooperationSigned terms match step 1

Step 1 decides most of the rest: "commercial use allowed" on a dataset card means little until you know whether you will release weights or compete with a generator's provider. General supplier questions are in the data provider due diligence questionnaire; the sections below cover what is specific to generated data.

Which model generated this dataset? Questions that settle it

Ask the vendor to name every generator and judge model with its exact version, access route (provider API, cloud marketplace or self-hosted weights), account type, terms version and generation dates per batch. A vendor that cannot produce these cannot credibly warrant that the outputs may train your model.

Some terms restrict the data; others attach conditions to the model you train:

  • Competition restrictions. As of October 2026, Anthropic's help center states that its terms do not allow using outputs to train models that compete with Anthropic's, lists general-purpose chatbots among the prohibited cases, and names sentiment analysis, summarization and information extraction tools as permitted examples [3]. One dataset can be usable for an extraction model and unusable for a chat model.
  • Naming conditions. A copy of the Llama 3.1 Community License distributed on Hugging Face requires that a model created, trained or fine-tuned with Llama materials or their outputs, if distributed or made available, include "Llama" at the beginning of its name [4].
  • Derivative-model definitions. Google's Gemma Terms of Use (April 2024 version) define Model Derivatives to include other models trained to perform similarly to Gemma by transferring patterns from its Outputs, explicitly including methods based on Gemma-generated synthetic data; Outputs themselves are not Model Derivatives [5]. Your model can fall inside the definition even though the dataset does not.
  • Use restrictions that travel. One dataset card states that because its data was generated with BLOOMZ models under the BigScience RAIL License v1.0, that license would apply to classifiers fine-tuned on it [6].

Mayer Brown advises asking what the generator itself was trained on and whether an open-weight license or acceptable-use policy limits commercial use of outputs, warning that such restrictions could make a synthetic dataset unusable for its purpose [1]. Lemley and Henderson argue that model output lacks the human authorship copyright requires, so these restrictions rest on contract, and whether they bind parties that never accepted them is debated [7]. Assume the terms apply until counsel concludes otherwise.

Terms change between versions, so the manifest records the terms version and date. Field detail is in provenance records for synthetic training data; provider analysis is in checking provider output terms before training.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "batch_id": "gen-batch-0142",
  "record_count": 18250,
  "task": "dpo_pairs_claims_policy_questions",
  "seed": {
    "source": "licensed_policy_faq_corpus",
    "license_ref": "LIC-[ID]",
    "pii_scan": {"method": "[TOOL_AND_VERSION]", "result": "pass"},
    "benchmark_overlap_check": "none above threshold"
  },
  "generators": [
    {"role": "candidate_responses", "model": "[MODEL_NAME]", "version": "[API_SNAPSHOT]",
     "access": "provider_api", "account_type": "enterprise", "terms_ref": "[TERMS_URL]@2026-05-02",
     "generated_from": "2026-05-04", "generated_to": "2026-05-09", "temperature": 0.8, "candidates_per_prompt": 4},
    {"role": "preference_judge", "model": "[MODEL_NAME]", "version": "[CHECKPOINT]",
     "access": "self_hosted_open_weights", "license": "[LICENSE_NAME_AND_VERSION]",
     "judge_prompt_ref": "judge-v3", "position_swap": true, "ties_allowed": true}
  ],
  "filters": [
    {"step": "exact_and_minhash_dedup", "threshold": 0.85, "removed": 1240},
    {"step": "eval_set_ngram_and_embedding_overlap", "removed": 37},
    {"step": "judge_score_below_4_of_5", "removed": 5610}
  ],
  "human_layer": {"role": "spot_review", "contributor_agreement": "assignment", "ai_tool_use_disclosed": true}
}

Seed data: the rights question vendors skip

Generation does not clean the rights in the seed. Prompt pools, source documents the generator rewrote, few-shot examples and personas each need a documented right to use them for generation and a check for personal data.

  • Non-commercial prompt pools. A third-party catalog lists ShareGPT 52K, user-shared ChatGPT conversations, under CC BY-NC 4.0 for non-commercial use [8]; prompts drawn from it bring that restriction into question for the synthetic set.
  • Stacked restrictions. Alpaca-style sets combine model-generated responses with their own dataset licenses; one curated list notes that non-commercial clauses in Alpaca and Cleaned Alpaca restrict their use in commercial products [9]. See open instruction and preference datasets for commercial use.
  • Unreliable license fields. The Data Provenance Initiative audited more than 1,800 text datasets and reported license omission rates above 70% and error rates above 50% on popular hosting sites [10]. A seed's license field is a lead to verify, not evidence.
  • Licensed or client documents. Ask whether the source agreement permits generating and selling derived data (rights to synthetic data generated from licensed data).
  • Personal data. Seeds built from support transcripts, emails or clinical notes can pass names, account numbers or rare facts into generated text (whether synthetic data derived from records is still personal data).

Seed pool design is covered in seed data for synthetic instruction generation.

Synthetic SFT data quality checks to rerun yourself

Ask for the vendor's filtering pipeline with rejection counts per step, then rerun the decisive checks on a random sample you draw, because vendor-reported quality reflects the vendor's thresholds and judges.

Filtering changes outcomes. The AlpaGasus authors found that Alpaca's 52,000 model-generated instruction examples contained many incorrect or irrelevant responses, and a model trained on a filtered subset of about 9,000 significantly outperformed the original Alpaca [11]. A study of synthetic tool-use data found that validated data beat unvalidated data even at smaller volume [12].

CheckHow to run itRed-flag result
Exact and near-duplicatesHash normalized text; MinHash with LSH (locality-sensitive hashing) over word n-grams; embedding similarity for paraphrasesTemplate families that differ only in names or numbers
Overlap with your evaluation setsn-gram plus embedding or model-based paraphrase detection against every benchmark and private eval you reportParaphrased test items; prompts lifted from public benchmarks
CorrectnessExecute code, check math against references, validate JSON against its schema, expert review of a stratified sampleErrors concentrated in the hardest task families
DiversityTopic coverage against your taxonomy; length distribution; repeated openings and closingsMost responses opening with the same few phrases
Generator artifactsSearch for refusals, assistant self-references and policy boilerplateResponses that name another assistant or refuse benign requests
Preference pairsShare where the chosen response is longer; agreement after swapping candidate order; tie rateJudge prefers by position or length; no tie option

Lee et al. (2021) found that deduplicating training data cut the rate at which models emit memorized text by about ten times and removed train-test overlap affecting more than 4% of validation examples in standard datasets [13]. For contamination, n-gram overlap is the predominant technique, with thresholds that vary by lab [14], but LMSYS researchers showed that paraphrased or translated test items pass simple n-gram checks, reporting a 13B model trained on rephrased MMLU items that scored 85.9 [15]. A generator seeded with benchmark-style prompts can produce that kind of paraphrase, so require semantic checks too.

Methods: near-duplicate detection with MinHash and LSH, contamination through synthetic data and assessing synthetic data quality.

Prove the lift on held-out real data before scaling the order

Vendor benchmark results show the data helped the vendor's model on the vendor's tests. Your decision needs an ablation on your base model, scored on real prompts the vendor never saw, run on a purchased slice before you commit to full volume.

  1. Hold out real prompts with reference answers or rubric scores from your own traffic, tickets or documents, and never share them with the vendor.
  2. Fine-tune the same base model on your current mix, the mix plus the synthetic slice and, if possible, the mix plus an equal volume of real data, to separate content from volume effects.
  3. Score each run on the real held-out set by task family, plus a regression set for general capability and safety behavior; for DPO data, also compare response length and refusal rate.
  4. Accept only if the lift holds on real prompts without material regressions.

Synthetic data can raise scores on tests that resemble it while missing the phrasing, errors and edge cases of real users. See using a licensed real-data holdout to validate synthetic data and combining licensed and synthetic data.

If you lack real records in the target domain, that is a sourcing problem of its own. SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing agreement. Datasets are sourced on request, not held in stock, so a request does not guarantee a match. You can describe the real records you need for validation.

Fine-tuning data license checklist for synthetic sets

The license should name each training use and each artifact you will produce, and the vendor should warrant generator-terms compliance, seed rights and contributor rights, backed by an indemnity and a replacement remedy. Synthetic data needs these promises more than real data, because the vendor may hold little copyright in purely generated text.

Mayer Brown reports U.S. Copyright Office guidance that protection for generative AI outputs depends on sufficient human expressive contribution, and that prompts alone generally do not suffice [1]. You are mostly paying for curation, documentation and contractual promises; exclusivity means little if anyone can generate similar data.

  • The grant covers SFT, preference optimization, reward-model training, distillation and your release mode (internal, API or open weights) (fine-tuning-only data licenses).
  • The generation manifest is a schedule, warranted accurate.
  • Warranty that generator and judge terms in force on the generation dates permitted the licensed uses (data warranties for AI training licenses).
  • Warranty of seed rights, and of no personal data beyond what the license allows.
  • Writers, editors and raters assigned or licensed their contributions and disclosed AI-tool use (annotator agreements and AI assistance).
  • Indemnity for claims arising from generator terms or seed sources (IP indemnities for licensed training data).
  • Replacement or refund for non-compliant records; notice if generator terms change mid-delivery.
  • Disclosure cooperation: the vendor supplies the facts you need to report synthetic data use.

That last item has statutory weight. California's AB 2013 requires developers of generative AI systems made publicly available to Californians to post training-data documentation, including whether synthetic data generation was used, due on or before 1 January 2026 [16]. As of October 2026, providers that place general-purpose AI models on the EU market must publish a training-content summary using the AI Office template under Article 53(1)(d) of the AI Act [17]. For when these duties reach fine-tuning teams, see provider duties when fine-tuning with acquired data.

Illustrative example: invented to show structure; it does not describe an available dataset. Not legal advice; adapt with counsel.

Generation warranty. Supplier represents and warrants that (a) each Generated Record was produced only by the models identified in Schedule B (Generation Manifest), under the terms identified there for the stated dates; (b) those terms permitted Supplier to license the Generated Records to Licensee for the Licensed Uses in Schedule A, including training models that Licensee distributes; (c) Supplier had the right to use all Seed Materials for generation and licensing; and (d) every human contributor has assigned or licensed their contribution to Supplier. Supplier will replace any non-compliant Generated Record or refund its pro-rata fee, without limiting Licensee's indemnity rights.

Red flags that should pause a synthetic data purchase

Pause the purchase when the vendor's answers show any of these signs:

  • "Generated with frontier models" and no model names, versions or dates.
  • A "commercial use" grant that says nothing about generator or judge terms.
  • Quality reported only as a score from the same model that generated the data.
  • Refusal to let you draw your own sample, or hand-picked samples.
  • "Human-reviewed" claims with no contributor agreements or reviewer counts; for the reverse problem, see detecting model-generated content in purchased human data.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Need real records to validate or replace synthetic fine-tuning data?

Describe the domain, task families and volume of real records you need, as a held-out evaluation set or as training data alongside synthetic examples, and the uses the license must cover. SourceX looks for US businesses that hold matching data, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Start a data request with SourceX.

Sources

  1. Mayer Brown, "Synthetic Data as a Deal Asset: Ownership, Provenance and Diligence Considerations in AI Acquisitions" (2026). https://www.mayerbrown.com/en/insights/publications/2026/07/synthetic-data-as-a-deal-asset-ownership-provenance-and-diligence-considerations-in-ai-acquisitions
  2. Lee et al. (arXiv:2309.00267; ICML 2024), "RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback" (2023). https://arxiv.org/pdf/2309.00267
  3. Anthropic Help Center, "Can I use my outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
  4. Meta (license copy distributed by Mozilla on Hugging Face), "Llama 3.1 Community License Agreement" (2024). https://huggingface.co/Mozilla/Meta-Llama-3.1-8B-Instruct-llamafile/blob/7d8b93e61bb828c11a25085aefc28ebf9a7952f2/LICENSE
  5. Google AI for Developers, "Gemma Terms of Use (archived version dated April 1, 2024)" (2024). https://ai.google.dev/gemma/terms-archive/terms_04_01_24
  6. BatsResearch, Hugging Face, "NusaX-senti-LexC-Gen dataset card (commit 260c323)". https://huggingface.co/datasets/BatsResearch/NusaX-senti-LexC-Gen/commit/260c3230882c4337c063ca2458bb67ba74667fa4
  7. SpicyIP, "Discussing Lemley and Henderson's 'The Mirage of Artificial Intelligence Terms of Use Restrictions'" (2025). https://spicyip.com/2025/01/discussing-lemley-and-hendersons-the-mirage-of-artificial-intelligence-terms-of-use-restrictions.html
  8. LLM Configurator (third-party dataset catalog), "ShareGPT 52K: LLM Instruction / SFT Dataset". https://llmconfigurator.com/en/datasets/sharegpt
  9. GitHub (inimah), "awesome-instruction-dataset". https://github.com/inimah/awesome-instruction-dataset
  10. Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  11. Chen et al. (arXiv:2307.08701), "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
  12. arXiv:2409.16341 (EMNLP 2024), "Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs" (2024). https://arxiv.org/html/2409.16341v2
  13. Lee et al. (arXiv:2107.06499; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  14. arXiv:2406.04244, "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
  15. LMSYS Org, "LLM Decontaminator blog post (14 November 2023)" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
  16. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  17. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data