Skip to content

Procurement, samples and ongoing supply

Build, buy or synthesize training data: a decision framework

Quick answer

The build vs buy training data decision turns on where the behavior you need is already recorded and who has the right to use it. Build in-house when your own systems hold it, your terms allow training on it, and the task is narrow. License existing records when the behavior lives in other organizations' systems or you need it sooner than a labeling program can ramp up. Synthesize when a program can check each example, real seed and test sets exist, and the generator's terms allow your use.

By SourceX Editorial · Updated

What building, buying and synthesizing commit you to

Each path is a separate supply chain with its own fixed costs, lead time and rights record, so compare them as operations, not as prices per example.

  • Build: extract first-party records (help-desk tickets, CRM activity, code review threads, claims notes) or run your own collection program, then label under guidelines you write. Contracted annotators can do the work; you still own the protocol, quality bar and rights basis.
  • Buy: license records another organization holds, buy an off-the-shelf dataset, or commission a supplier to collect to your specification. For pre-training text it includes openly licensed corpora such as the 8 TB Common Pile v0.1 [1]. Commissioning is compared separately in custom data collection vs licensing existing records.
  • Synthesize: generate examples with a language model prompted from seeds, a simulator, or a tabular generator fit to real data, then filter and validate.

What each data origin is good at is covered in licensed vs synthetic vs scraped training data.

Seven questions that settle most build-vs-buy calls

Answer in order; an early answer can rule out a path before you cost it.

QuestionPoints towardWhy it matters
1. Do systems you control record the behavior with outcomes attached (resolution codes, approvals, reverted commits)?BuildYou hold the raw material; cost shifts to extraction, cleaning and labeling.
2. Do your privacy notices and customer contracts permit model training on it?Build only if yesFTC staff wrote in 2024 that adopting AI training through a surreptitious, retroactive terms change may be unfair or deceptive [2]. The 2021 Everalbum order required deletion of models developed from users' photos and videos [3].
3. Does the behavior live mainly in other organizations' records, or do you need variation across many of them?LicenseNo internal program or generator can produce outcomes only other companies observed.
4. Can a program verify each example (unit tests, a JSON Schema, a solver, a rules engine)?Synthesize at scaleWithout a verifier, every synthetic example needs human review.
5. Is the set for evaluation?Build or license held-out real dataA test set made by the generator behind your training data shares its blind spots, so scores can overstate performance.
6. Must labeled data arrive before you could hire and calibrate annotators?LicenseTime to the first usable batch is often the binding constraint.
7. Will volume, language count or the quality bar rise?Buy, or outsource labelingA language-services vendor argues in-house annotation is usually cheaper at low, steady volume in one language and a narrow domain, with outsourcing pulling ahead as those grow [4].

Cost per usable example: the number to compare

Compare paths on cost per usable example over the period you will use the data, not on a license fee set against an annotation hourly rate.

Cost per usable example = (fixed setup + variable cost per raw example × raw volume + QA and rework + legal and privacy review + refresh over the term) ÷ (raw volume × yield), where yield is the share of raw examples that survive deduplication, de-identification, QA and filtering.

Cost lineBuildBuy or licenseSynthesize
AcquisitionEngineering to extract and join source recordsLicense fee or collection statement of workInference or API fees; seed data
LabelingGuidelines, annotator time, tooling, ramp-upIncluded, or a separate annotation layerFrom the generator; must be verified
QAGold items, double annotation, adjudicationSample inspection, acceptance testingVerifiers, model scoring, human spot checks
Legal and privacyNotice and contract review, de-identificationRights diligence, negotiation, de-identification checksGenerator terms, seed-data license, leakage tests
Yield lossRejected and reworked labelsRecords failing acceptanceGenerations discarded by filters
RefreshRe-extraction and relabeling after changesRenewal or supply agreementRegeneration when generator or task changes

Yield is where comparisons go wrong. AlpaGasus scored the 52,000 model-generated Alpaca instruction examples with an LLM grader, kept about 9,000, and reported a better model from the smaller set [5]. Measure yield on a pilot for every path. The licensed column is broken down in total cost of ownership for licensed training data.

Time to the first usable batch

Lead time depends on the steps each path must finish before one batch passes acceptance; synthesis usually waits on real data from the other paths.

PathSteps before the first accepted batchUsual gate
BuildRights check; extraction; guidelines; annotator recruiting and calibration; pilot with agreement measuredGuidelines stable enough for agreement to hold
BuySpecification; supplier discovery; sample under an evaluation license; rights review; negotiation; de-identification and deliverySupplier approval and contract review
SynthesizeSeed set and rubric; generator and terms check; pipeline; generation; filtering; validation on held-out real dataA real held-out set to validate against

Licensing stages are covered in how long it takes to license training data.

When building in-house is the right call

Build when the data sits in your systems, your right to train on it is clear, and the task is narrow enough for a small team to hold one labeling standard.

Small curated sets go far in post-training: LIMA fine-tuned a 65-billion-parameter LLaMA model on 1,000 curated prompt-response pairs, and its authors note such curation is labor-intensive [6]. Building costs are mostly this quality work:

  • Versioned guidelines with decision rules for ambiguous cases, so each label traces to the rules in force.
  • Gold items with known answers mixed into queues to track each annotator's accuracy.
  • Double annotation on a sample, scored with Krippendorff's alpha (1 is perfect reliability, 0 is agreement no better than chance) [7], with adjudication that updates the guidelines.
  • Label audits on evaluation sets. An audit of test sets from 10 widely used vision, language and audio datasets estimated an average label error rate of at least 3.3% [8].

ISO/IEC 5259-4 sets out a data quality process framework for ML covering acquisition, preparation, labeling and evaluation [9], usable as a checklist for an internal program. In the EU, training on customers' personal data needs a lawful basis; the EDPB's Opinion 28/2024 says legitimate interest is not a default basis for AI models and must pass a three-step test [10].

Evaluation is the strongest case for building: held-out production records match your real inputs and were never published. Researchers argue that private evaluation by outside curators gives up the transparency of open benchmarks [11]; if you buy test data, ask how it is kept apart from training data the supplier sells. See the evaluation datasets hub.

When licensing existing records beats collecting them

License when the behavior is recorded in other organizations' operational systems, when you need variation across many of them, or when collection cannot produce real outcomes in time.

Multi-year support histories with resolutions, claims with adjudication outcomes, and reverted code changes record consequences a new collection program would have to wait for and a generator cannot supply.

Buying does not remove rights work. The Data Provenance Initiative audited more than 1,800 text datasets and reported license omission above 70% and license error rates above 50% on popular hosting sites [12], so trace the license chain. Collecting third-party content yourself is not free either: as of October 2026, the Third Circuit's precedential opinion of 29 September 2026 in Thomson Reuters v. Ross held that ROSS's use of Westlaw headnotes to train its non-generative legal research tool was not fair use [13].

SourceX works on this path: it sources operational datasets from US companies and manages the licensing agreement and ongoing purchases. Each dataset goes through rights review and is delivered under a license defining the included records, permitted uses, term and delivery. Datasets are sourced on request, so a request does not guarantee a match. If licensing is your route, describe the records you need to SourceX, then estimate the dataset's value on a sample before committing to volume.

What synthetic data needs before it enters the plan

Synthesize when you have a real seed set, a real held-out set, a verifier or review budget, and a generator whose terms allow your use.

  1. Real reference data. SDMetrics, for example, evaluates synthetic data against the real data it was modeled on with statistical, detection, efficacy and privacy metrics [14]. Similarity scores alone do not show the data serves its task.
  2. Generator terms on record. Some providers restrict training competing models on their outputs; as of October 2026, Anthropic's help center states its terms do not allow outputs to train models competitive with Anthropic's, while naming permitted non-competing uses [15]. Record model, version and terms date per batch.
  3. Seed-data rights. If seeds are licensed records, confirm the license permits derived data and says whether it survives termination; see derivative and successor model rights and combining licensed and synthetic data.
  4. A real-data stage. TabPFN was pre-trained solely on synthetic tables; the Real-TabPFN authors report that continued pre-training on a small, curated set of real datasets improved accuracy on 29 benchmark datasets [16].
  5. A privacy test, not an assumption. Synthetic records derived from personal data fall outside the GDPR only if they meet Recital 26's anonymity standard, judged by the means reasonably likely to be used to identify someone [17]. The EDPB counts synthetic data among measures that can weigh in the legitimate-interest balancing test [10].
  6. License checks on public synthetic sets. The same provenance audit found non-commercial or closed licenses dominating newer synthetic datasets [12].

See also synthetic data vs real business data.

The provenance records each path leaves behind

Each path produces different evidence of where a dataset came from, and disclosure rules now ask for it. As of October 2026:

  • California AB 2013 requires developers of generative AI systems offered to Californians to post documentation covering dataset sources or owners, whether datasets include copyrighted, licensed or personal information, collection periods, and whether synthetic data was used, due by 1 January 2026 and before each new or substantially modified release [18]. See the AB 2013 records to collect from suppliers.
  • EU AI Act Article 53(1)(d) requires providers of general-purpose AI models to publish a training-content summary on the AI Office template [19].
  • EU AI Act Article 10(2) requires data governance for high-risk systems covering collection processes, data origin, annotation, labeling and cleaning, bias examination and data gaps [20]. Regulation (EU) 2026/1744 amended the AI Act; secondary sources report that it moved Annex III high-risk obligations to 2 December 2027 [21].
PathRecords to keep for each dataset
BuildSource system and extraction query; privacy notice or contract terms at collection; guideline version; agreement statistics; QA logs
BuyLicense and supplier rights representations; datasheet; de-identification method and sample check; delivery manifest
SynthesizeGenerator model, version and terms date; seed sources and licenses; prompts; filter rules and rejection rates; validation results

A datasheet covering motivation, composition, collection process and recommended uses gives all three paths one format [22]; the training data provenance guide goes further. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

A path decision record to adapt

Record the decision per data need so finance, legal and the model team review the same reasoning.

Illustrative example: invented to show structure; it does not describe an available dataset.

data_need: DN-014
capability: "Draft first replies to B2B software support tickets"
uses: [sft, eval]
build:
  source: "own help-desk export, three product lines"
  rights_basis: "customer terms cover service improvement, not model training; counsel to confirm"
  gap: "few escalations with recorded resolutions"
buy:
  target: "resolved multi-turn tickets with resolution codes from several companies"
  uses_to_license: [training, evaluation, retain trained weights after term]
  first_step: "sample under an evaluation license, then rights review"
synthesize:
  generator: "<model, version, terms date>"
  verifier: "none for reply quality; human review required"
  role: "augment rare escalation types only"
decision:
  train: "license a foundation set; synthesize rare-case augmentation"
  eval: "build from held-out own tickets; no generator output"
  compare_on: "cost per usable example over the license term; yield from pilot"
  revisit: "after the pilot learning curve"

Then size the purchase with how much training data to buy.

Mistakes that skew the comparison

Most poor build-vs-buy calls come from comparing the wrong quantities.

  • Treating first-party data as free when the privacy notice at collection never mentioned model training [2].
  • Accepting synthetic data on fidelity scores without the efficacy test of training on it and scoring on held-out real data [14].
  • Judging a licensed dataset on a hand-picked sample instead of a random draw from the records you would receive.
  • Costing a build as one-off when product, policy or label changes will force a re-run.
  • Building the evaluation set last, which leaves no way to measure pilot yield or a purchase's value.

For the full buying sequence, see the AI training data procurement guide.

Weighing a license against building it yourself?

If the records you need sit in other companies' systems, describe them on the SourceX buyer page: what the records contain, the history and volume you need, and the uses you need licensed. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery. Start a data request with SourceX.

Sources

  1. Kandpal et al. (arXiv:2506.05209; NeurIPS 2025), "The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text" (2025). https://arxiv.org/html/2506.05209v1
  2. Federal Trade Commission, Office of Technology (Tech@FTC staff blog), "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  3. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  4. Acolad (language and data services vendor), "Data annotation cost". https://www.acolad.com/en/services/data-services/data-annotation-cost
  5. Chen et al. (arXiv:2307.08701), "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
  6. Zhou et al. (arXiv:2305.11206; NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  7. Klaus Krippendorff, Annenberg School for Communication, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  8. Northcutt, Athalye and Mueller (arXiv:2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  9. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  10. CMS (law-firm summary of European Data Protection Board Opinion 28/2024), "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models". https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  11. arXiv preprint 2503.04756, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
  12. Longpre et al. (arXiv; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  13. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (precedential opinion)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  14. Synthetic Data Vault (SDV) project documentation, "SDMetrics". https://docs.sdv.dev/sdmetrics
  15. Anthropic (Claude Help Center), "Can I use my Outputs to train an AI model?". https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
  16. arXiv:2507.03971 (ICML 2025 workshop), "Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data" (2025). https://arxiv.org/pdf/2507.03971
  17. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  18. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  19. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  20. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  21. European Union, Official Journal (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en
  22. Gebru et al. (arXiv:1803.09010; Communications of the ACM, 2021), "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data