Skip to content

For AI data buyers

AI data for buyers: sourcing, licensing and evaluating datasets

Quick answer

To source and license AI training data, start from the model use and work backward. Name the training stage, record type, fields, volume and permitted uses, then pick a channel: open datasets, licensed business records, commissioned collection or synthetic generation. Verify the supplier's chain of title and de-identification, test a sample against your own held-out evaluation set, and sign a license that names each permitted use and the delivery terms. This hub maps each of those decisions to the guide that covers it.

By SourceX Editorial · Updated

The six decisions behind every AI data purchase

Every dataset purchase settles the same six questions, in roughly this order: what the data is for, where it can come from, whether the supplier can grant the rights, whether privacy and disclosure duties are met, whether the data is good enough, and how it reaches your pipeline. Deals stall when a team negotiates price before the third or fourth question has an answer.

DecisionWhat you should hold before moving onGo deeper
1. Use and specificationA written specification of the record and the uses to license (example below)Training data RFP template
2. ChannelA reasoned choice between open data, licensed records, commissioned collection and synthetic generationBuild, buy or synthesize
3. RightsChain-of-title documents, consent and notice records, and a rights grant that names each useChain of title for training data; pre-training rights grant
4. Privacy and disclosureThe de-identification standard applied, its evidence, and the per-dataset facts your disclosure duties requireDe-identified vs anonymized definitions; disclosure duties compared
5. Quality and fitSample test results against the checklist belowAcceptance criteria for training data
6. DeliveryFormat, schema, data dictionary, manifest with checksums, transfer channelManifests and checksums

Decision 3 cannot rest on a license tag in a public repository. The Data Provenance Initiative's audit of more than 1,800 text datasets reported license omission rates above 70% and license error rates above 50% on popular dataset hosting sites [1]. Re-check open datasets with an open dataset license audit before they enter a commercial training mix.

When the data you need sits in US companies' operational systems (support and sales histories, engineering records, documents, finance and legal workflows), SourceX sources it on request for AI teams and manages the licensing and ongoing purchases. These are kinds of data it sources, not inventory under contract; it does not source scraped web content and does not train models.

Write the specification before contacting suppliers

A specification that describes the record, not the vendor, lets suppliers answer yes or no quickly and gives counsel the scope of the license. The guide to writing a data request for suppliers explains each field below.

Illustrative example: invented to show structure; it does not describe an available dataset.

use:
  stage: supervised_fine_tuning   # pre_training | continued_pretraining | preference | evaluation | retrieval | agent_trajectories
  target_behavior: "draft replies to B2B support tickets that cite the correct help-center article"
record:
  type: support_ticket_thread
  source_systems: ["helpdesk export", "CRM account table"]
  unit: "one ticket with customer messages, agent replies, internal notes and the resolving article ID"
  required_fields: [ticket_id, created_at, product_area, customer_tier, messages, resolution_code, kb_article_id, csat_score]
volume:
  minimum_records: 20000
  time_range: "2022-01-01 to 2025-12-31"
  languages: [en-US]
exclusions: ["tickets from payment-card or health workflows", "file attachments"]
privacy:
  treatment: "names, emails, phone and account numbers removed or replaced before delivery"
  evidence: "method description and post-processing sample check"
rights_requested:
  uses: [fine_tuning, internal_evaluation]
  model_retention_after_term: required
  exclusivity: non_exclusive
holdout: "10% of accounts reserved for evaluation, never trained on"
delivery:
  format: "JSONL, one thread per line"
  manifest: "file list with SHA-256 checksums and record counts"

Match the dataset to the model stage you are buying for

The model stage decides both what a usable record looks like and which uses the license must name. A dataset licensed for evaluation only does not cover a fine-tuning run, and a license to ground answers in retrieved content is a different grant from a training license.

StageWhat a usable record containsLicense point to settleStart here
Pre-training and continued pre-trainingLong-form domain text or code, with source, date and license per documentTraining rights that extend to derivative and successor modelsDomain corpora for continued pre-training
Supervised fine-tuningPrompt-response pairs or worked solutions checked by domain expertsDeployment of the fine-tuned model to your customersSourcing SFT data
Preference dataA prompt, two or more responses, a preference label and rater metadataOwnership of labels and rater consentPreference datasets for DPO
EvaluationUnpublished items with gold answers or rubricsEvaluation-only use, no training, access controlsPrivate eval sets vs public benchmarks
Retrieval and groundingDocuments with stable IDs, timestamps, section structure and relevance judgmentsIndexing, caching, excerpt display and index deletion at term endGrounding license vs training license
AgentsTrajectories of observations, actions, tool calls and outcomes, or linked event logsReplay, derived tasks and benchmark creationProcess mining event logs

Record count matters less than curation at the post-training stages. LIMA fine-tuned a 65B-parameter LLaMA model on 1,000 curated prompt-response pairs [2], and Direct Preference Optimization (DPO) trains on preference pairs without a separate reward model [3]. Evaluation data loses its value once it leaks into training: OpenAI stopped reporting SWE-bench Verified, stating that score gains increasingly reflected training-time exposure to the benchmark [4]. For agents, process-mining logs are one structured source; the OCEL 2.0 specification defines SQLite, XML and JSON exchange formats for object-centric event logs [5].

Rights, privacy and disclosure: what reviewers will check

Counsel, privacy and security reviewers approve a data deal when three things are documented: the supplier's right to license each record for your named uses, the de-identification standard applied, and the facts your own disclosure duties require. The training data due diligence checklist lists the underlying documents.

Rights to each record

Business records carry the holder's promises to its own customers. FTC staff warned AI companies in January 2024 that breaking promises not to use customer data for model training can violate laws the FTC enforces [6], and the FTC's 2021 Everalbum order required deletion of models and algorithms built from users' photos and videos, not only the data [7]. Check customer contracts and DPAs for training use before relying on a supplier's ownership claim.

Copyright questions attach to the training copies themselves, not only to redistribution. The US Copyright Office's Part 3 report on generative AI training, a pre-publication version released in May 2025, concluded that many acts in training may be prima facie infringing unless an exception such as fair use applies, and it examined licensing [8]. On 29 September 2026 the Third Circuit held, in a precedential opinion, that ROSS's use of Westlaw material to train a non-generative legal-research tool was not fair use [9]. Commissioned collection has its own consent rules: California Penal Code section 632 bars recording a confidential communication without the consent of all parties [10].

What "de-identified" means in each regime

The same dataset can be de-identified under one law and personal data under another. HIPAA recognizes Safe Harbor, which removes 18 listed identifiers and requires no actual knowledge that the rest could identify someone, and Expert Determination, in which a qualified expert finds the identification risk very small [11]. The CCPA treats information as deidentified only if, among other conditions, the business holding it contractually binds recipients to the definition's terms, so a buyer should expect contract terms that bar re-identification [12]. Under the GDPR, pseudonymised data that can be re-attributed with additional information remains personal data [13].

Disclosure duties as of October 2026

Three regimes require developers to describe their training data, one of them only from 2027:

  • EU AI Act, Article 53. General-purpose AI model providers must keep a copyright policy that identifies and honors text-and-data-mining rights reservations, and must publish a summary of training content using the AI Office template [15]. The template, published 24 July 2025, asks providers to list the main data collections and explain other sources [16]. These duties have applied since 2 August 2025; EUR-Lex lists a consolidated text reflecting the July 2026 amendments in Regulation (EU) 2026/1744 [14]. See EU AI Act summaries: what buyers need from suppliers.
  • California AB 2013. Developers of generative AI systems offered to Californians had to post training-data documentation by 1 January 2026, covering dataset sources or owners, copyrighted or licensed material, personal information, collection periods and synthetic data [17].
  • Colorado SB26-189. Signed in May 2026, it requires developers of automated decision-making technology that materially influences consequential decisions to give deployers documentation that includes training data categories from 1 January 2027 [18].

Collect source, owner, license status, personal-information status and collection dates for every dataset at intake, so a single intake record can feed all three disclosures.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Test a sample before negotiating volume

A sample is worth more than a data sheet: run it through the checks you will apply at acceptance, and against your own evaluation sets, before price and volume are agreed.

  • Duplicates. Measure exact and near-duplicate rates. Lee et al. found one sentence repeated over 60,000 times in C4, and deduplicated training cut memorized output roughly tenfold [19].
  • Label errors. Hand-audit a stratified slice. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets [20].
  • PII residue. Run a detector, then review by hand. Presidio's maintainers caution that, because it relies on trained models, it cannot guarantee finding all sensitive information [21].
  • Contamination. Check overlap with your held-out sets and the public benchmarks you report; see decontaminating against benchmarks.
  • Coverage. Compare time range, languages, rare classes and edge cases with your deployment traffic.
  • Schema conformance. Check null rates, value ranges and field types against the data dictionary.

ISO/IEC 5259-2 defines a data quality model and measurable quality characteristics for machine learning data [22], which gives supplier and buyer a shared vocabulary for acceptance criteria.

What a complete delivery includes

A delivery is complete when files, documentation and integrity evidence arrive together. JSON Lines suits text records: UTF-8 without a byte order mark, one JSON value per line [23]. Parquet suits tables because its footer metadata lets readers fetch only the columns they need [24]. A zero-copy warehouse share is another route: Snowflake's documentation describes shares in which no data is copied and the consumer's access is read-only [25], which shifts the end-of-term question from deleting copies to revoking the share and tracing any copies made from it.

Documentation should travel with the files. Datasheets for Datasets covers motivation, composition, collection process and recommended uses [26]. Croissant expresses dataset metadata as schema.org-based JSON-LD [27], and NeurIPS 2026 requires Croissant-RAI responsible-AI metadata for its Evaluations and Datasets Track [28]. Ask for both a human-readable card and a machine-readable record.

The 21 guide collections in this hub

The hub has 21 collections: five on the deal, five on how the data will be used, eight on data types, and three on context.

CollectionRead it when you need to
AI training data licensingNegotiate the rights grant, field of use, exclusivity, term and pricing structure
Data provenance and permitted useProve a supplier may license each record: chain of title, consent records, opt-outs
De-identified and privacy-sensitive dataBuy records that contain personal, health, financial, biometric or children's data
AI training data complianceMap the EU AI Act, state laws and copyright guidance to the records you keep
AI training data procurementRun RFPs, samples, scorecards, security reviews, acceptance and renewals
Fine-tuning and post-training dataBuy SFT pairs, preference data, reasoning traces or continued pre-training corpora
LLM evaluation datasetsBuild private held-out sets, golden sets or agent task suites
RAG and retrieval dataLicense content for grounding, or buy relevance judgments, query logs and reranker data
AI agent training dataSource trajectories, tool-call traces, event logs and decision records
Training data quality assessmentMeasure duplicates, label quality, coverage and contamination before acceptance
Text datasets for LLM trainingLicense long-form, multilingual, dialogue, forum or review text
Document AI datasetsSource invoices, forms, contracts and scans with OCR, layout or key-value labels
Tabular, time-series and transactional dataBuy ledgers, telemetry, CRM or ERP tables, or text-to-SQL data
Code and software engineering datasetsLicense proprietary repositories, issue-to-fix pairs, code review and CI history
Image datasets for computer visionSource inspection, product, aerial or medical images with releases
Speech and audio datasetsBuy contact-center or conversational audio and check recording consent
Video datasetsSource procedural, operations, driving or screen-recording video
Multimodal and embodied dataBuy aligned image, text, audio and sensor data or robotics demonstrations
Industry-specific operational dataFind one record type inside one industry, such as claim notes or maintenance logs
AI data sourcing by teamSee what your role should require of a supplier
Dataset delivery formats and transferSpecify formats, manifests, transfer channels, versioning and destruction certificates

Mistakes that stall AI data deals

Four gaps delay deals and are cheap to close early:

  • Asking for a category instead of a record. "Healthcare data" cannot be matched or scoped; name the record type, source system and required fields, as in the specification above.
  • Leaving model retention open. Agree what happens to trained models when the license ends before signature, not at renewal.
  • Splitting the evaluation holdout after training starts. Reserve held-out accounts or time periods before the first training run, or contamination is hard to rule out.
  • Accepting files without a manifest. Without per-file checksums and record counts you cannot prove what was delivered, or later what was deleted.

License operational data from US businesses

If your specification calls for real records from US companies' operational systems, describe the data, the model stage and the uses you need licensed. SourceX looks for US businesses that hold it, checks the data and each supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. You can also browse the dataset catalog of record types it sources. Start a data request with SourceX.

Guides by topic

Sources

  1. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  2. Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  3. Rafailov et al., "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  4. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  5. OCEL standard authors, "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
  6. Federal Trade Commission, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  7. Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
  8. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  9. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  10. California Legislature, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632
  11. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  12. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  13. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)". https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  14. European Parliament and Council (EUR-Lex), "Regulation (EU) 2024/1689 (Artificial Intelligence Act)" (2024, consolidated 2026). https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
  15. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  16. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  17. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  18. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
  19. Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  20. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  21. Microsoft (Presidio project), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  22. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and machine learning, Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  23. jsonlines.org, "JSON Lines". https://jsonlines.org/
  24. The Apache Software Foundation, "Parquet File Format". https://parquet.apache.org/docs/file-format/
  25. Snowflake Inc. (vendor documentation), "About Secure Data Sharing". https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
  26. Gebru et al., "Datasheets for Datasets" (2018, revised 2021). https://arxiv.org/pdf/1803.09010
  27. Akhtar et al., "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  28. NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data