Skip to content

Procurement, samples and ongoing supply

Data RFI: Mapping the Supplier Market Before an RFP

Quick answer

A data RFI is a short, non-binding questionnaire sent to possible training data suppliers to find out whether a category can be sourced at all. It asks what source systems they hold, what rights basis covers AI use, how many years of history exist, rough volume ranges, how personal data is handled, whether samples are possible and how long preparation takes. It does not ask for prices or binding proposals. Those belong in the RFP, which you write once the RFI answers show what the market can realistically deliver.

By SourceX Editorial · Updated

RFI vs RFP for data procurement: what each one is for

Use an RFI to learn what the market can do. Use an RFP to buy. An RFI asks whether the market can supply something, while an RFP asks a supplier to commit to scope and price. Mixing the two produces either long questionnaires that holders of good data never answer, or price quotes based on guesses about rights and history. Settle feasibility first, then solicit binding proposals.

The difference matters more for data than for software. A missing data category often looks like a pricing problem. In practice the supplier lacks consent language that covers model training, keeps only 18 months of history in its ticketing system, or cannot release free text without a de-identification project. An RFP written before you know these limits gets either no bids or bids that quietly rewrite your requirements. As one vendor's RFP guide puts it, the quality of your scoping sets the quality of the bids you get back [6].

DimensionData RFIData RFP
Question it answersCan this category be sourced, and on what basis?Who will supply it, on what terms and at what price?
CommitmentNone on either sideBinding proposal, often with validity period
Length10-15 questions, about one pageFull specification, schema, acceptance criteria, contract draft
PricingNot requested (optional order-of-magnitude only)Required, normalized per usable record
SamplesAsks whether a sample is possible and under what conditionsRequests the sample itself, under NDA or evaluation terms
OutputShortlist plus realistic requirementsAward decision

When you are ready for the second column, use the AI training data RFP template. The full lifecycle from requirement to renewal is covered in the AI training data procurement hub.

Data sourcing RFI questions that actually predict feasibility

The questions that predict feasibility are the ones about source systems, rights and history. Questions about capacity and volume predict much less. Ask in the order an assessor would check them, so a supplier who fails early does not spend effort on later answers.

  1. Source system and record type. Which system produces the records: Zendesk or Salesforce Service Cloud tickets, Jira issues and linked commits, NetSuite journal entries, contract repositories, call recordings? Ask for the native export format (JSON API export, CSV, Parquet, EML/MBOX, WAV plus transcript).
  2. Rights basis. Does the supplier own the records outright, or do they belong to the supplier's customers? Do the terms of service, employee agreements or consent notices permit licensing for model training? Is there any third-party content inside, such as customer attachments or vendor documents?
  3. History depth and continuity. How many years exist, are there gaps from system migrations, and was the schema stable throughout? A help desk migration, for example from Freshdesk to Zendesk, can break thread linkage.
  4. Volume range. Ask for bands (under 10,000; 10,000-100,000; over 100,000 records), not exact counts. Bands are easier to give before internal review.
  5. Personal and sensitive data. Which fields carry names, emails, phone numbers, account numbers or free-text identifiers? Has the supplier de-identified a dataset before, and with what method? For protected health information, ask whether it would use Safe Harbor or Expert Determination under 45 CFR 164.514 [3].
  6. Sample policy. Can a redacted sample of 50-200 records be shared under NDA, or only viewed in a secure environment?
  7. Preparation lead time. How long are legal review, extraction and de-identification likely to take, stated as a range?
  8. Ongoing supply. Is the data still being generated, and could refreshes be delivered on a cadence?
  9. Documentation. Can the supplier describe provenance in a structured form, for example the Croissant RAI vocabulary for collection and provenance attributes [1]?
  10. Approval path. Who inside the supplier must sign off on releasing data (legal, security, privacy, the data owner), and has anyone there approved a similar release before?

Keep the instrument short. Operating companies that hold valuable records rarely have bid teams, so a 40-question form gets no answer from exactly the suppliers you most want to reach. Save the deep diligence for the data provider due diligence questionnaire and for chain-of-title documents once a supplier is shortlisted.

A one-page RFI template for training data suppliers

A usable RFI fits on one page. It has a need statement, the questions above, a response format and a non-binding notice. The template below is written for a hypothetical agent-training request.

Illustrative example: invented to show structure; it does not describe an available dataset.

REQUEST FOR INFORMATION (non-binding) - Ref: RFI-2026-07
Respond by: [date, ~3 weeks out]   Format: reply inline or attach a 1-2 page memo

1. NEED STATEMENT
   We are exploring licensed records of multi-step IT service desk work:
   ticket text, agent actions (status changes, assignments, macro use),
   linked knowledge-base articles, and resolution outcomes.
   Intended use: supervised fine-tuning and evaluation of task agents.
   Preferred: 3+ years of history; English; US operations.

2. QUESTIONS
   Q1  Source system(s) and native export format
   Q2  Ownership of records and any third-party content inside them
   Q3  Do existing terms/consents permit licensing for model training? (Y/N/Unknown)
   Q4  Years of history; known gaps or migrations
   Q5  Volume band: <10k | 10k-100k | >100k tickets
   Q6  Personal-data fields present; prior de-identification experience
   Q7  Sample possible? redacted file under NDA | secure viewing | not before contract
   Q8  Indicative preparation lead time (range)
   Q9  Ongoing generation; refresh cadence possible
   Q10 Internal approvers required for a release

3. RESPONSE NOTICE
   This RFI is for planning only. It is not a solicitation, and responses
   are not offers. No pricing is requested. Do not send any records or
   personal data with your response.

Two design choices matter. The need statement describes the data and the use. It does not name the model, the capability being targeted or the launch timeline. And the "Unknown" option on Q3 is deliberate. A truthful "Unknown" is worth more than a confident "Yes" that falls apart during rights review.

Protecting your roadmap during market sounding for datasets

Write the need around the record type and the behavior you need, not the product you are building. "Resolved IT tickets with agent action logs" tells a supplier enough to check its systems. "Data for our autonomous IT agent launching in Q2" tells competitors too much if the RFI is forwarded.

Use a mutual NDA only if you have to share schemas or evaluation criteria that give away strategy. Most first-round RFIs need none, and asking for one discourages smaller holders from replying. If you send the RFI through an intermediary, decide in advance whether suppliers will see your company name. Also tell suppliers not to attach real records, because unsolicited personal data in an RFI response becomes your compliance problem.

Reading RFI responses: from answers to RFP requirements

RFI answers are for adjusting requirements and building a shortlist. They do not choose a winner. Sort responses into three groups:

  • Feasible as specified: rights basis is clear, history meets your floor, and a sample is possible.
  • Feasible with changes: for example, only two years of history, free text available only after redaction, or audio without screen recordings. Decide whether the change is acceptable before the RFP and not during negotiation.
  • Not feasible: the rights basis is "Unknown" with no internal owner, or the data belongs to the supplier's customers and cannot be relicensed.

Then rewrite the RFP using what you learned. If every credible responder says de-identified free text needs eight weeks or more, an RFP that demands delivery in four weeks will select the least careful supplier. If history depth clusters at two to three years, set the floor there and treat more history as a scoring bonus. Check whether the shortlist covers your deployment distribution using a coverage gap analysis. For quality terms that will go into acceptance criteria, ISO/IEC 5259-2 provides defined data quality measures you can reference [2].

Downstream disclosure questions to plant in the RFI

Ask early whether a supplier can support the disclosures you will have to make later. Retrofitting provenance after the deal is expensive. As of October 2026, a provider placing a general-purpose AI model on the EU market must publish a training-content summary under AI Act Article 53(1)(d), using the Commission template dated 24 July 2025 [4]. A developer of generative AI offered in California had to post training data documentation under AB 2013, with a posting deadline of 1 January 2026 [5]. Both obligations fall on the model provider or developer rather than the data supplier, but both depend on the supplier being able to describe the source, the collection period and whether personal data is present.

One extra RFI question covers this: "Can you provide a written description of data source, collection period, and presence of personal or copyrighted third-party content suitable for public disclosure?" A supplier who hesitates here will probably struggle with consent and notice records later.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Where an RFI stops and a sourced request starts

An RFI only reaches suppliers you can already name. Many of the most useful operational records sit with ordinary businesses that do not sell data and will never see your RFI. For those categories, the alternative is to describe the data and let a sourcing partner find holders, as covered in how to write a data request for suppliers.

SourceX works this way. It sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, finance and legal workflows, and new recordings of hands-on work. It does not hold stock, and a request does not guarantee a match. Buyers describe the data, not the businesses, and every release is approved by the supplying company. You can describe the category your RFI is testing alongside your own market sounding.

Running a training data RFI with SourceX in the loop

If your RFI shows that a category is real but hard to reach, SourceX can look for US businesses that hold it. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, and nothing is contracted until a supplier agrees. Start a buyer request at sourcex.si/buyers.

Sources

  1. arXiv (MLCommons Croissant RAI authors), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  2. ISO, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  3. eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  4. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  5. California Legislative Information, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  6. Ertas, "AI Data Preparation RFP Template". https://www.ertas.ai/blog/ai-data-preparation-rfp-template

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data