Procurement, samples and ongoing supply
AI Training Data RFP Template: Sections, Response Forms and Scoring
Quick answer
An AI training data RFP template sets up a formal request that makes every supplier of licensed or custom-collected data answer the same questions in the same format, so you can compare proposals on fit, rights, privacy, quality and cost. Unlike a software RFP, it sets a record-level specification and asks for evidence of each source's licensing basis, the de-identification standard applied, a randomly drawn sample under evaluation terms, and prices on one common unit.
By SourceX Editorial · Updated
When a formal RFP beats a data request or an RFI
Issue a formal request for proposal (RFP) when you can write the record-level specification, expect several credible suppliers, and need a documented competitive award; otherwise use a lighter instrument. The RFP is one stage in the AI training data procurement lifecycle.
| Instrument | Use it when | What comes back |
|---|---|---|
| Data request | One route is likely | A feasibility answer (how to write a data request for suppliers, data request builder) |
| Request for information (RFI) | You do not know who holds the data | A market map (data RFI guide) |
| RFP | The specification is fixed; several suppliers can bid | Comparable, priced proposals with samples |
Generic AI-vendor RFP templates cover scope, technical requirements, privacy and security, total cost of ownership, a proof of concept, service levels and legal terms; one also flags training data governance [1]. None of those headings asks what each record contains, who may license it, or what was removed.
The eleven sections of a training data RFP
A training data RFP needs eleven sections, and four of them (data specification, rights and provenance, de-identification of the offered records, and the sample) have no real counterpart in a software RFP.
| # | Section | You specify | Each bidder returns |
|---|---|---|---|
| 1 | Context and intended use | Training stage (pre-training, supervised fine-tuning, evaluation, retrieval), products, territories | Uses it can license and uses it excludes |
| 2 | Data specification | Record unit, mandatory fields, volume floor, time window, languages, exclusions | Fact sheet with population counts and field fill rates |
| 3 | Rights and provenance | Uses to license, contractor access, model retention after the term | Legal owner, licensing basis with evidence, conflicting exclusive grants |
| 4 | Privacy | Required standard, such as HIPAA Expert Determination or Safe Harbor [2] or CCPA "deidentified" [3] | Method, who applied it, residual quasi-identifiers, re-identification testing |
| 5 | Security | Minimum controls in transit, at rest and for samples | Questionnaire; subprocessors with access |
| 6 | Sample and pilot | Draw rule, size, evaluation license, measurements | Sample per the rule |
| 7 | Pricing | One pricing unit; options priced separately | Completed price form |
| 8 | Delivery | Formats, data dictionary, transfer method, refresh cadence | Proposed method and schedule |
| 9 | Legal terms | Draft license or term sheet | Markup or exceptions list |
| 10 | Evaluation | Gates, weights, scoring panel | None; publishing them shapes bids |
| 11 | Timeline and rules | Dates, Q&A channel, submission format, bid validity | Intent to bid, questions, proposal |
Sections 2 to 4 also feed your own disclosures. As of October 2026, California AB 2013 requires developers of generative AI systems offered to Californians to post training-data documentation (due since 1 January 2026) on items such as dataset sources or owners, copyrighted or licensed material, personal information, collection periods and synthetic data [4]. The European Commission's training-content summary template (24 July 2025) asks general-purpose model providers to summarise the data used for training [5], and AI Act Article 10 will require data governance for high-risk systems, covering origin, preparation, bias examination and gaps, once the high-risk obligations apply [6].
Write the specification around the record, not the topic
Bids become comparable only when the RFP defines the record unit, mandatory fields, volume floor and time window tightly enough that two suppliers cannot read them differently. One data-preparation RFP guide warns that a vague RFP draws vague bids, while an overly rigid one does not help selection [7].
Illustrative example: invented to show structure; it does not describe an available dataset.
- Vague: "Customer support conversations for LLM fine-tuning, large volume."
- Specific: "Resolved support ticket threads from US business software companies, 2021 to 2025: every message in order, agent or customer role per turn, product area, resolution code, satisfaction score where captured. English. At least 200,000 threads. Contact and account identifiers removed before sampling."
Three drafting rules:
- Tag each requirement must, should or could. One AI RFP template cites advice to keep to 8 to 20 pages plus attachments and 25 or fewer categorized functional requirements [1]; put field lists in an attached data dictionary.
- Ask for quality as named measures with a method. ISO/IEC 5259-2 sets out data quality measures [8]; ask for each measure's property and the method used to quantify it. Require fill rates on the full offered population, not the sample.
- Ask for the near-duplicate rate and method (exact hashing, MinHash or none). One study found a sentence repeated more than 60,000 times in C4, and deduplicated training cut memorized output about tenfold [9].
Turning model goals into data requirements covers deriving the must-haves.
Licensed records and custom collection go in separate lots
If you will consider both existing records and new collection, split the RFP into lots: the routes are evidenced, priced and contracted differently, so one scoring sheet cannot judge both.
| Lot A: license existing records | Lot B: custom collection | |
|---|---|---|
| What exists at bid time | The data: ask for counts, fill rates, a sample | A protocol: ask for the plan, recruitment method, a pilot batch |
| Rights evidence | Customer terms, notices or agreements permitting licensing | Contributor consent and release forms; rights in deliverables |
| Privacy question | Which de-identification was applied, and how | Consent scope and what is captured |
| Pricing unit | Per record or slice, plus refresh | Per collected hour, session or item, plus setup, QA and rework |
| Contract | License | Statement of work plus license or assignment |
Custom collection versus licensing explains when each fits.
Response forms: the questions every data provider answers
Fixed forms stop bidders answering in marketing prose: require fact sheet, rights and provenance, privacy, security and price forms per offered dataset, with evidence attached. Labels alone are unreliable: an audit of more than 1,800 text datasets reported license omission above 70% and license error rates above 50% on popular hosting sites [10].
Base the fact sheet on Datasheets for Datasets (motivation, composition, collection process, recommended uses) [11] and Data Cards (upstream sources, annotation methods, intended use) [12]. Ask for a machine-readable copy too: Croissant-RAI extends the Croissant JSON-LD vocabulary with responsible-AI fields [13], and NeurIPS 2026 requires responsible-AI metadata built on them for its Evaluations and Datasets Track [14]. Borrow two questions from the FISD Alternative Data Council questionnaire: consent terms for data about individuals, and contract terms permitting resale of data bought from others [15].
Illustrative example: invented to show structure; it does not describe an available dataset.
bid_id: "BID-07" # one response per offered dataset
lot: A # A = existing records, B = custom collection
fact_sheet:
dataset_name: "Resolved support threads, B2B software"
source_systems: ["Zendesk", "Salesforce Service Cloud"]
legal_owner: "entity that controls the records"
record_unit: "resolved ticket thread"
population_count: 0 # full offered population, not the sample
collection_period: {start: "2021-01", end: "2025-12"}
languages: ["en"]
fields: # full list in the attached data dictionary
- {name: "message_text", type: "string", fill_rate_pct: 0, measured_on: "full population"}
- {name: "resolution_code", type: "enum", fill_rate_pct: 0, measured_on: "full population"}
near_duplicate_rate_pct: 0
dedup_method: "MinHash; threshold stated"
synthetic_share_pct: 0
known_gaps: []
machine_readable_card: "Croissant JSON-LD with RAI fields, or 'not available'"
rights_and_provenance:
licensing_basis: "owned records; customer terms permit licensing"
evidence_attached: ["customer terms excerpt, version date", "privacy notice, version date"]
third_party_content: "customer attachments excluded"
web_scraped_content: false
conflicting_exclusive_grants: "none"
uses_offered: ["pre-training", "fine-tuning", "evaluation"] # must map to RFP section 1
uses_excluded: ["retrieval with verbatim display"]
consent_terms_for_individuals: "attached"
resale_rights_if_acquired_from_third_party: "not applicable"
pending_claims: "none"
privacy:
personal_data_categories: ["names", "emails", "phone numbers", "account numbers"]
standard_met: "CCPA 1798.140(m) deidentified"
method: "rules plus named-entity recognition; surrogate replacement; post-processing sample check"
applied_by: "supplier"
residual_quasi_identifiers: ["company size band", "US state"]
reidentification_testing: "method and date"
security:
questionnaire: "attached"
subprocessors_with_access: []
delivery:
formats: ["Parquet"]
transfer_method: "cloud bucket with cross-account access"
data_dictionary: "attached"
refresh: {cadence: "quarterly", incremental: true}
On the privacy form, ask which standard was met, never whether data is "anonymized"; HIPAA Safe Harbor, for example, requires removing 18 listed identifiers [2]. CCPA "deidentified" status obliges the business to bind recipients by contract, so expect those terms in the license [3], and NIST SP 800-188 recommends measurable de-identification performance levels plus re-identification studies [16]. Pair the forms with the data provider due diligence questionnaire, a data rights attestation and a supplier security review.
Sample and pilot rules to publish with the RFP
Publish the sample rules (draw method, size, form, evaluation terms and measurements) so no bidder can submit a hand-picked showcase.
- Draw and form. Random or stratified from the offered population, processed exactly as the full delivery would be, including de-identification.
- Terms. An evaluation-only license with an NDA. Market examples include NVIDIA's revocable, non-transferable sample data license limited to evaluation and testing, which bars distribution [17], and a template combining a trial data license with a mutual NDA [18].
- Size. Enough to measure fill rates and label accuracy on critical fields. FISD's questionnaire asks for under 100 rows over three months old [15]: fine for diligence, too small for a training test.
- Measurements. Your own tests in your own pipeline. One AI RFP guide argues for proof on the buyer's data rather than vendor presentations [19]; another notes AI performance claims are harder to verify than traditional software features [20].
- Evaluation data. Ask who has accessed the held-out split and whether the bidder also supplies training or annotation data to model developers, an overlap researchers flag as a risk of private evaluation curators [21].
See requesting a sample, evaluation licenses and NDAs, running a data pilot and contamination-resistant evaluation design.
A price form that converts every bid to cost per usable record
Fix the pricing unit and cost components in the RFP, then convert each bid to cost per usable record (records passing your acceptance checks). Every price form shows unit price and volume tiers, one-time preparation fees, minimum commitment, refresh price, separately priced options (exclusivity, longer term, added uses), replacement of rejected records, and price validity.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Bid A | Bid B | |
|---|---|---|
| Quoted total (price index) | 100 | 130 |
| Records offered | 1,000,000 | 1,000,000 |
| Share passing your pilot checks | 60% | 90% |
| Usable records | 600,000 | 900,000 |
| Price index per 1,000 usable records | 0.167 | 0.144 |
Bid A looks 23% cheaper but costs about 15% more per usable record. Price the transfer method too: Snowflake Secure Data Sharing copies no data between accounts and is read-only for the consumer [22], Delta Sharing, documented by the open-source Delta Lake project, is another share-based option [23], and Parquet or JSON Lines files put a full copy, and its storage cost, on your side. See comparing vendor quotes and total cost of ownership.
Scoring: pass/fail gates first, weights second
Publish gates and weights in the RFP. Rights evidence, the required de-identification standard and minimum security controls should be pass/fail gates, because no price or volume compensates for data you cannot lawfully use.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Criterion | Type | Weight | Scored on |
|---|---|---|---|
| Rights and provenance evidence | Gate | n/a | Evidence; no conflicting exclusive grant |
| Required de-identification standard | Gate | n/a | Privacy form and sample inspection |
| Minimum security controls | Gate | n/a | Security questionnaire |
| Fit to specification | Weighted | 30% | Fields, volume, time window, languages |
| Sample or pilot results | Weighted | 25% | Measured fill rates, label accuracy, duplicates |
| Cost per usable record | Weighted | 20% | Price form adjusted by pilot pass rate |
| Documentation depth | Weighted | 15% | Data dictionary, datasheet or Croissant card |
| Delivery and refresh | Weighted | 10% | Format, transfer method, refresh terms |
Score with a panel (ML lead on fit and pilot, counsel on rights, privacy and security on their forms, procurement on price), scoring documents before opening prices if policy allows. The data vendor evaluation scorecard covers scales and tie-breaks; internal approvals for a data purchase covers sign-off.
A data RFP timeline with a Q&A window and a sample stage
Data RFPs need two stages software RFPs often compress: time for suppliers to check their customer contracts before answering rights questions, and a sample or pilot stage before award.
| Stage | What happens |
|---|---|
| 1. Issue | RFP, data dictionary, draft license and NDA go out together |
| 2. Intent to bid | Bidders confirm lots and sign the NDA |
| 3. Written Q&A | Answers go to all bidders without naming who asked; the specification freezes at close |
| 4. Proposals due | Forms, evidence and prices, with a validity period |
| 5. Gate review and shortlist | Gates applied, documents scored |
| 6. Samples or pilot | Random draw under evaluation terms; paid pilot terms where needed |
| 7. Best and final offers | Price and license markup refined |
| 8. Award | License, acceptance criteria and delivery plan |
Supplier counsel review and sample extraction drive elapsed time; see the training data licensing timeline.
Mistakes that make data proposals impossible to compare
Three drafting gaps survive even good response forms.
- Software service levels. Uptime does not describe data; acceptance needs record counts, schema conformance, field completeness, and rejection and replacement terms.
- Unnamed uses. If the RFP never mentions retrieval or evaluation, the license draft may cover training only.
- No indemnity question. Ask each bidder's position on indemnity for third-party infringement up front.
Where to send a training data RFP
Send it to every route that could hold the data: operating companies that own the records, brokers and resellers, collection vendors, and managed sourcing services that approach data owners for you.
SourceX works in that last route. It sources operational datasets, such as support histories, engineering records and finance workflows, from US companies on request and manages the licensing agreement. Buyers describe the data, not the businesses; every dataset goes through rights review, and scraped public web content is out of scope. You can submit your RFP's data specification to SourceX as a data request; the categories it describes are kinds of data it sources, not inventory under contract.
Running a data RFP? Share the specification with SourceX
Describe the records, fields, volume, time window and uses from your RFP. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, agrees allowed uses in a license and coordinates delivery; nothing is contracted until a supplier agrees, and a request does not guarantee a matching dataset. Send SourceX your data specification.
Sources
- Dan Cumberland Labs, "AI Vendor RFP Template" (consultancy blog). https://dancumberlandlabs.com/blog/ai-vendor-rfp-template/
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- Ertas, "How to Scope an AI Data Preparation Project (RFP Template)" (vendor blog). https://www.ertas.ai/blog/ai-data-preparation-rfp-template
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- Gebru et al., "Datasheets for Datasets" (2018; Communications of the ACM 2021). https://arxiv.org/pdf/1803.09010
- Pushkarna, Zaldivar and Kjartansson, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- Jain et al., "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- NeurIPS, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
- FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (DDQ) with GenAI Questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- NVIDIA, "NVIDIA Sample Data License for Evaluation" (2026, license document). https://developer.download.nvidia.com/licenses/nvidia-sample-data-license-for-evaluation-2026.01.19.pdf
- New Constructs, "Trial Data License and Mutual Non-Disclosure Agreement" (2024, template). https://www.newconstructs.com/wp-content/uploads/2024/10/New-Constructs-TDLA-Mutual-NDA-General.pdf
- Amit Kothari, "AI RFP template" (practitioner blog; title from URL). https://amitkoth.com/ai-rfp-template/
- Pertama Partners, "AI RFP template: key sections and questions" (consultancy; title from URL). https://www.pertamapartners.com/insights/ai-rfp-template-key-sections-questions
- arXiv:2503.04756, "Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators" (2025). https://arxiv.org/html/2503.04756v1
- Snowflake, "About Secure Data Sharing" (Snowflake Documentation). https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
- Delta Lake project, "Read Delta Sharing Tables" (documentation). https://docs.delta.io/delta-sharing/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.