For AI data buyers
AI data for buyers: sourcing, licensing and evaluating datasets
Quick answer
To source and license AI training data, start from the model use and work backward. Name the training stage, record type, fields, volume and permitted uses, then pick a channel: open datasets, licensed business records, commissioned collection or synthetic generation. Verify the supplier's chain of title and de-identification, test a sample against your own held-out evaluation set, and sign a license that names each permitted use and the delivery terms. This hub maps each of those decisions to the guide that covers it.
By SourceX Editorial · Updated
The six decisions behind every AI data purchase
Every dataset purchase settles the same six questions, in roughly this order: what the data is for, where it can come from, whether the supplier can grant the rights, whether privacy and disclosure duties are met, whether the data is good enough, and how it reaches your pipeline. Deals stall when a team negotiates price before the third or fourth question has an answer.
| Decision | What you should hold before moving on | Go deeper |
|---|---|---|
| 1. Use and specification | A written specification of the record and the uses to license (example below) | Training data RFP template |
| 2. Channel | A reasoned choice between open data, licensed records, commissioned collection and synthetic generation | Build, buy or synthesize |
| 3. Rights | Chain-of-title documents, consent and notice records, and a rights grant that names each use | Chain of title for training data; pre-training rights grant |
| 4. Privacy and disclosure | The de-identification standard applied, its evidence, and the per-dataset facts your disclosure duties require | De-identified vs anonymized definitions; disclosure duties compared |
| 5. Quality and fit | Sample test results against the checklist below | Acceptance criteria for training data |
| 6. Delivery | Format, schema, data dictionary, manifest with checksums, transfer channel | Manifests and checksums |
Decision 3 cannot rest on a license tag in a public repository. The Data Provenance Initiative's audit of more than 1,800 text datasets reported license omission rates above 70% and license error rates above 50% on popular dataset hosting sites [1]. Re-check open datasets with an open dataset license audit before they enter a commercial training mix.
When the data you need sits in US companies' operational systems (support and sales histories, engineering records, documents, finance and legal workflows), SourceX sources it on request for AI teams and manages the licensing and ongoing purchases. These are kinds of data it sources, not inventory under contract; it does not source scraped web content and does not train models.
Write the specification before contacting suppliers
A specification that describes the record, not the vendor, lets suppliers answer yes or no quickly and gives counsel the scope of the license. The guide to writing a data request for suppliers explains each field below.
Illustrative example: invented to show structure; it does not describe an available dataset.
use:
stage: supervised_fine_tuning # pre_training | continued_pretraining | preference | evaluation | retrieval | agent_trajectories
target_behavior: "draft replies to B2B support tickets that cite the correct help-center article"
record:
type: support_ticket_thread
source_systems: ["helpdesk export", "CRM account table"]
unit: "one ticket with customer messages, agent replies, internal notes and the resolving article ID"
required_fields: [ticket_id, created_at, product_area, customer_tier, messages, resolution_code, kb_article_id, csat_score]
volume:
minimum_records: 20000
time_range: "2022-01-01 to 2025-12-31"
languages: [en-US]
exclusions: ["tickets from payment-card or health workflows", "file attachments"]
privacy:
treatment: "names, emails, phone and account numbers removed or replaced before delivery"
evidence: "method description and post-processing sample check"
rights_requested:
uses: [fine_tuning, internal_evaluation]
model_retention_after_term: required
exclusivity: non_exclusive
holdout: "10% of accounts reserved for evaluation, never trained on"
delivery:
format: "JSONL, one thread per line"
manifest: "file list with SHA-256 checksums and record counts"
Match the dataset to the model stage you are buying for
The model stage decides both what a usable record looks like and which uses the license must name. A dataset licensed for evaluation only does not cover a fine-tuning run, and a license to ground answers in retrieved content is a different grant from a training license.
| Stage | What a usable record contains | License point to settle | Start here |
|---|---|---|---|
| Pre-training and continued pre-training | Long-form domain text or code, with source, date and license per document | Training rights that extend to derivative and successor models | Domain corpora for continued pre-training |
| Supervised fine-tuning | Prompt-response pairs or worked solutions checked by domain experts | Deployment of the fine-tuned model to your customers | Sourcing SFT data |
| Preference data | A prompt, two or more responses, a preference label and rater metadata | Ownership of labels and rater consent | Preference datasets for DPO |
| Evaluation | Unpublished items with gold answers or rubrics | Evaluation-only use, no training, access controls | Private eval sets vs public benchmarks |
| Retrieval and grounding | Documents with stable IDs, timestamps, section structure and relevance judgments | Indexing, caching, excerpt display and index deletion at term end | Grounding license vs training license |
| Agents | Trajectories of observations, actions, tool calls and outcomes, or linked event logs | Replay, derived tasks and benchmark creation | Process mining event logs |
Record count matters less than curation at the post-training stages. LIMA fine-tuned a 65B-parameter LLaMA model on 1,000 curated prompt-response pairs [2], and Direct Preference Optimization (DPO) trains on preference pairs without a separate reward model [3]. Evaluation data loses its value once it leaks into training: OpenAI stopped reporting SWE-bench Verified, stating that score gains increasingly reflected training-time exposure to the benchmark [4]. For agents, process-mining logs are one structured source; the OCEL 2.0 specification defines SQLite, XML and JSON exchange formats for object-centric event logs [5].
Rights, privacy and disclosure: what reviewers will check
Counsel, privacy and security reviewers approve a data deal when three things are documented: the supplier's right to license each record for your named uses, the de-identification standard applied, and the facts your own disclosure duties require. The training data due diligence checklist lists the underlying documents.
Rights to each record
Business records carry the holder's promises to its own customers. FTC staff warned AI companies in January 2024 that breaking promises not to use customer data for model training can violate laws the FTC enforces [6], and the FTC's 2021 Everalbum order required deletion of models and algorithms built from users' photos and videos, not only the data [7]. Check customer contracts and DPAs for training use before relying on a supplier's ownership claim.
Copyright questions attach to the training copies themselves, not only to redistribution. The US Copyright Office's Part 3 report on generative AI training, a pre-publication version released in May 2025, concluded that many acts in training may be prima facie infringing unless an exception such as fair use applies, and it examined licensing [8]. On 29 September 2026 the Third Circuit held, in a precedential opinion, that ROSS's use of Westlaw material to train a non-generative legal-research tool was not fair use [9]. Commissioned collection has its own consent rules: California Penal Code section 632 bars recording a confidential communication without the consent of all parties [10].
What "de-identified" means in each regime
The same dataset can be de-identified under one law and personal data under another. HIPAA recognizes Safe Harbor, which removes 18 listed identifiers and requires no actual knowledge that the rest could identify someone, and Expert Determination, in which a qualified expert finds the identification risk very small [11]. The CCPA treats information as deidentified only if, among other conditions, the business holding it contractually binds recipients to the definition's terms, so a buyer should expect contract terms that bar re-identification [12]. Under the GDPR, pseudonymised data that can be re-attributed with additional information remains personal data [13].
Disclosure duties as of October 2026
Three regimes require developers to describe their training data, one of them only from 2027:
- EU AI Act, Article 53. General-purpose AI model providers must keep a copyright policy that identifies and honors text-and-data-mining rights reservations, and must publish a summary of training content using the AI Office template [15]. The template, published 24 July 2025, asks providers to list the main data collections and explain other sources [16]. These duties have applied since 2 August 2025; EUR-Lex lists a consolidated text reflecting the July 2026 amendments in Regulation (EU) 2026/1744 [14]. See EU AI Act summaries: what buyers need from suppliers.
- California AB 2013. Developers of generative AI systems offered to Californians had to post training-data documentation by 1 January 2026, covering dataset sources or owners, copyrighted or licensed material, personal information, collection periods and synthetic data [17].
- Colorado SB26-189. Signed in May 2026, it requires developers of automated decision-making technology that materially influences consequential decisions to give deployers documentation that includes training data categories from 1 January 2027 [18].
Collect source, owner, license status, personal-information status and collection dates for every dataset at intake, so a single intake record can feed all three disclosures.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Test a sample before negotiating volume
A sample is worth more than a data sheet: run it through the checks you will apply at acceptance, and against your own evaluation sets, before price and volume are agreed.
- Duplicates. Measure exact and near-duplicate rates. Lee et al. found one sentence repeated over 60,000 times in C4, and deduplicated training cut memorized output roughly tenfold [19].
- Label errors. Hand-audit a stratified slice. Northcutt et al. estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets [20].
- PII residue. Run a detector, then review by hand. Presidio's maintainers caution that, because it relies on trained models, it cannot guarantee finding all sensitive information [21].
- Contamination. Check overlap with your held-out sets and the public benchmarks you report; see decontaminating against benchmarks.
- Coverage. Compare time range, languages, rare classes and edge cases with your deployment traffic.
- Schema conformance. Check null rates, value ranges and field types against the data dictionary.
ISO/IEC 5259-2 defines a data quality model and measurable quality characteristics for machine learning data [22], which gives supplier and buyer a shared vocabulary for acceptance criteria.
What a complete delivery includes
A delivery is complete when files, documentation and integrity evidence arrive together. JSON Lines suits text records: UTF-8 without a byte order mark, one JSON value per line [23]. Parquet suits tables because its footer metadata lets readers fetch only the columns they need [24]. A zero-copy warehouse share is another route: Snowflake's documentation describes shares in which no data is copied and the consumer's access is read-only [25], which shifts the end-of-term question from deleting copies to revoking the share and tracing any copies made from it.
Documentation should travel with the files. Datasheets for Datasets covers motivation, composition, collection process and recommended uses [26]. Croissant expresses dataset metadata as schema.org-based JSON-LD [27], and NeurIPS 2026 requires Croissant-RAI responsible-AI metadata for its Evaluations and Datasets Track [28]. Ask for both a human-readable card and a machine-readable record.
The 21 guide collections in this hub
The hub has 21 collections: five on the deal, five on how the data will be used, eight on data types, and three on context.
| Collection | Read it when you need to |
|---|---|
| AI training data licensing | Negotiate the rights grant, field of use, exclusivity, term and pricing structure |
| Data provenance and permitted use | Prove a supplier may license each record: chain of title, consent records, opt-outs |
| De-identified and privacy-sensitive data | Buy records that contain personal, health, financial, biometric or children's data |
| AI training data compliance | Map the EU AI Act, state laws and copyright guidance to the records you keep |
| AI training data procurement | Run RFPs, samples, scorecards, security reviews, acceptance and renewals |
| Fine-tuning and post-training data | Buy SFT pairs, preference data, reasoning traces or continued pre-training corpora |
| LLM evaluation datasets | Build private held-out sets, golden sets or agent task suites |
| RAG and retrieval data | License content for grounding, or buy relevance judgments, query logs and reranker data |
| AI agent training data | Source trajectories, tool-call traces, event logs and decision records |
| Training data quality assessment | Measure duplicates, label quality, coverage and contamination before acceptance |
| Text datasets for LLM training | License long-form, multilingual, dialogue, forum or review text |
| Document AI datasets | Source invoices, forms, contracts and scans with OCR, layout or key-value labels |
| Tabular, time-series and transactional data | Buy ledgers, telemetry, CRM or ERP tables, or text-to-SQL data |
| Code and software engineering datasets | License proprietary repositories, issue-to-fix pairs, code review and CI history |
| Image datasets for computer vision | Source inspection, product, aerial or medical images with releases |
| Speech and audio datasets | Buy contact-center or conversational audio and check recording consent |
| Video datasets | Source procedural, operations, driving or screen-recording video |
| Multimodal and embodied data | Buy aligned image, text, audio and sensor data or robotics demonstrations |
| Industry-specific operational data | Find one record type inside one industry, such as claim notes or maintenance logs |
| AI data sourcing by team | See what your role should require of a supplier |
| Dataset delivery formats and transfer | Specify formats, manifests, transfer channels, versioning and destruction certificates |
Mistakes that stall AI data deals
Four gaps delay deals and are cheap to close early:
- Asking for a category instead of a record. "Healthcare data" cannot be matched or scoped; name the record type, source system and required fields, as in the specification above.
- Leaving model retention open. Agree what happens to trained models when the license ends before signature, not at renewal.
- Splitting the evaluation holdout after training starts. Reserve held-out accounts or time periods before the first training run, or contamination is hard to rule out.
- Accepting files without a manifest. Without per-file checksums and record counts you cannot prove what was delivered, or later what was deleted.
License operational data from US businesses
If your specification calls for real records from US companies' operational systems, describe the data, the model stage and the uses you need licensed. SourceX looks for US businesses that hold it, checks the data and each supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. You can also browse the dataset catalog of record types it sources. Start a data request with SourceX.
Guides by topic
- AI Agent Training Data: Trajectories, Logs and DecisionsWhat AI agent training data is made of: GUI trajectories, tool-call traces, event logs, decision records and seed data, and where to source each kind.
- AI Data Sourcing by Team: Who Owns Each DecisionWho owns each AI data sourcing decision, from model teams to counsel, privacy, governance and data engineering, and what each must get from a supplier.
- AI Training Data Compliance: Laws, Standards and RecordsWhich AI laws, copyright rules, privacy regimes and standards reach acquired training data as of October 2026, and the record each one requires.
- AI Training Data Licensing: Rights, Terms and PricingHow AI training data licensing works for buyers: rights to train, keep models and use outputs, risk terms, pricing structures and open-license limits.
- AI Training Data Procurement: A Stage-by-Stage Buyer MapAI training data procurement mapped stage by stage: specs, sourcing routes, RFPs, samples, DDQs, approvals, contracts, acceptance, supply and renewal.
- Code Datasets for LLM Training: Types, Uses and RisksA buyer's map of code datasets for LLM training: repositories, issue-to-fix tasks, reviews, CI logs and legacy code, what each trains, and what to check.
- Data Provenance for AI Training: A Buyer's Evidence MapProvenance evidence to require for licensed AI training data: source, chain of title, consent, opt-outs, licenses, synthetic origin and permitted use.
- Dataset Delivery Formats and Transfer for Licensed AI DataDataset delivery formats by data type, plus the manifests, metadata, transfer channels, encryption and update rules to specify for licensed AI data.
- De-identified Data for AI Training: Standards and EvidenceWhat de-identified means under HIPAA, GDPR and US state law, where identifiers hide in each data type, and the evidence to request before licensing.
- Document AI Datasets by Task: Labels, Sources and RightsWhich annotation layer each document AI task needs, which public datasets allow commercial use, and when to license real business documents instead.
- Fine-Tuning Datasets: SFT, Preference and RL Data GuideCompare seven kinds of fine-tuning and post-training data, from SFT pairs to preference and RLVR tasks, and four routes to source and license each one.
- Image Datasets for Computer Vision: Sources and RightsCompare public benchmarks, scraped images, licensed company photo archives and commissioned capture for computer vision, and the checks each route needs.
- Industry-Specific Data for AI Training: Systems and RulesWhich industry operational records AI teams license, the systems they come from, and the HIPAA, GLBA, consent and export rules that limit each one.
- LLM Evaluation Datasets: Benchmarks vs Private Test DataCompare five sources of LLM evaluation data: public benchmarks, commissioned sets, licensed business records, synthetic items and golden sets.
- Multimodal Training Data: Types, Sources and LicensingHow AI teams source multimodal training data: paired, interleaved, synchronized and robot records, where each comes from, and how to license and check it.
- RAG Content Licensing: Grounding Rights and Retrieval DataHow to license content for RAG and AI grounding, and how to source retriever training and evaluation data: rights, pricing, corpora, qrels and pitfalls.
- Speech Datasets for AI: Types, Sources and Voice RightsMap speech and audio data types to sourcing routes, then check license terms, call-recording consent and voiceprint risk before licensing training data.
- Tabular Data for AI: Tables, Time Series and TransactionsWhich tabular, relational, SQL, spreadsheet, ledger, time-series and log data trains which AI models, and what to check before licensing real records.
- Text Datasets for LLM Training: Sources and Rights TiersA buyer's map of text datasets for LLM training: five rights tiers, what open corpora lack, needs by training stage and checks before commercial use.
- Training Data Quality Assessment for Licensed DatasetsAssess training data quality before you license it: five check families, ISO/IEC 5259 metrics, contamination tests, and when quality beats volume.
- Video Datasets for AI Training: Sources, Specs and RightsWhere AI training video comes from: benchmarks, stock, company recordings and commissioned capture, plus the spec, rights and privacy checks to run.
Sources
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- Rafailov et al., "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
- OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
- OCEL standard authors, "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
- Federal Trade Commission, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
- California Legislature, "California Penal Code section 632". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)". https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- European Parliament and Council (EUR-Lex), "Regulation (EU) 2024/1689 (Artificial Intelligence Act)" (2024, consolidated 2026). https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Microsoft (Presidio project), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Data quality for analytics and machine learning, Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- The Apache Software Foundation, "Parquet File Format". https://parquet.apache.org/docs/file-format/
- Snowflake Inc. (vendor documentation), "About Secure Data Sharing". https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
- Gebru et al., "Datasheets for Datasets" (2018, revised 2021). https://arxiv.org/pdf/1803.09010
- Akhtar et al., "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.