Procurement, samples and ongoing supply
Data provider due diligence questionnaire (DDQ) for AI training data
Quick answer
A data provider due diligence questionnaire (DDQ) is a written, vendor-level questionnaire a buyer sends before approving a data supplier. For AI training data it covers the company, how and from whom it sources data, the consents and notices behind personal data, its right to resell or sublicense, litigation history, security, and generative AI questions such as copyright opt-outs and model-generated content. Start from the FISD alternative-data template, demand evidence for each answer, and score answers before approval.
By SourceX Editorial · Updated
What a vendor DDQ covers that other diligence does not
A vendor DDQ assesses the supplier as a company across everything it sells; it does not replace the dataset-level chain-of-title check or the security review. The vendor level matters because purchased data is a supply-chain component: NIST's Generative AI Profile (AI 600-1) lists intellectual property, data privacy, and value chain and component integration among 12 generative AI risks [1].
| Instrument | Question it answers | Usual owner |
|---|---|---|
| Vendor DDQ (this page) | Should we do business with this supplier at all? | Vendor-risk or procurement lead, with legal |
| Dataset due diligence checklist | May this licensor license this dataset for our use? | Counsel |
| Security review of a data supplier | Are data and our information protected in transit and at rest? | Security |
| Shortlist questions | Is this supplier worth a formal process? | Data partnerships lead |
The DDQ usually follows an RFP for training data and precedes any delivery beyond a small sample; the AI training data procurement guide shows the full sequence.
Adapting the FISD alternative-data DDQ to AI training data
A ready starting point is the illustrative Data Provider DDQ from FISD's Alternative Data Council, whose 2024 edition adds generative AI questions; FISD does not endorse any party using it [2]. It was written for investment managers worried about material non-public information (MNPI), so AI buyers should swap those items for copyright, consent and permitted-use questions.
Several requests carry over: a data dictionary, a small sample (fewer than 100 rows, from more than three months ago), copies of consent terms agreed with individuals, and, if the vendor buys data from others, the contract terms that permit resale [2]. A Lowenstein Sandler guide adds litigation and enforcement history, how data is sourced, processed and transmitted, and detailed provenance, and advises vendors to keep a form DDQ ready with redacted agreement excerpts or privacy notices [3]. Neudata's 2025 review benchmarks vendor DDQ responses across alternative data [4]: evidence of routine market practice, not a standard.
| FISD-style item | For AI training data |
|---|---|
| MNPI screens | Replace with third-party confidential information, trade secrets and copyright |
| Data dictionary | Keep; one per product line, with null rates and date ranges |
| Sample under 100 rows, over three months old | Keep the request, but ask for a representative sample instead; see how to request a training data sample |
| Consent terms with individuals | Keep; add each notice version in force at collection and when AI use was first disclosed |
| Resale terms | Keep; ask for sublicensing rights covering training, evaluation and deployment |
| Litigation and enforcement | Keep; add rightsholder complaints, opt-outs and takedown notices |
| Generative AI items, where present | Extend to model-generated content, crawling, the vendor's use of your data, and disclosure support |
The ten sections of an AI training data DDQ
Ten sections cover vendor-level training-data risk, and every question should demand a specific answer and its evidence. Ask for answers per product line where sourcing differs, because a company-level answer hides the riskiest source.
| Section | Core questions | Evidence to attach |
|---|---|---|
| 1. Company and control | Entity, owners, affiliates, recent changes of control, license signatories, state data-broker registrations | Organization chart, registrations |
| 2. Sourcing model | Per product: share from first-party operations, commissioned collection, customer data, purchased data, public web, model generation | Source-class breakdown, provenance metadata |
| 3. Upstream rights | Named upstream suppliers; resale or sublicensing rights for AI training; term and flow-down limits | Redacted contract excerpts [2][3] |
| 4. Consent and personal data | Notices and consents at collection; de-identification method and legal standard; sensitive categories | Dated notices, consent text, method statement |
| 5. Copyright and web content | Crawling, opt-out handling, how books, articles, code and media were acquired | Crawl policy, opt-out log |
| 6. Generative AI | Model-generated records; the vendor's own training on customer data | Generator model, version, terms |
| 7. Litigation and complaints | Past, current or threatened litigation, regulator inquiries, rightsholder and data-subject complaints | Summary under NDA |
| 8. Security | Summary only; detail goes to the security review | Attestation reports |
| 9. Documentation | Data dictionary, dataset card, fields for your own disclosures | Datasheet or machine-readable card |
| 10. Attestation | Officer signature; duty to notify new upstream sources, lawsuits, breaches, changes of control | Signed attestation |
For section 2, a fixed structure makes answers comparable, and the Data & Trust Alliance's proposed data provenance standards supply one: a dataset's source, legal rights and privacy protections, a standardized timestamp, data type, generation method, and intended uses and restrictions [5]. SourceX's data source documentation check shows which of these records your intake already requests.
Section 4 needs sector prompts, because the statute often binds the buyer too; the privacy review of a training data vendor covers the rest:
- CCPA "deidentified" claims. The definition requires the business to contractually bind recipients, including against re-identification [6]; ask for that language, since you will be bound by it.
- Consumer health data. Washington's My Health My Data Act requires a separate authorization to sell consumer health data, and seller and purchaser must keep copies for six years [7].
- Financial data. Nonpublic personal information from a financial institution carries Regulation P reuse and redisclosure limits that bind the recipient [8]; ask how the vendor received it, and under which exception if any.
- Changed terms. Ask when AI use first appeared in the source's notice or terms: FTC staff warn that adding AI training by surreptitious, retroactive amendment may be unfair or deceptive [9], and that breaking promises not to train on customer data may create liability [10].
Generative AI questions to add to any data vendor DDQ
Whatever generative AI items a template already has, an AI training data DDQ needs detailed questions on copyright opt-outs, acquisition method, model-generated data and disclosure support. Ask every supplier, even those claiming only first-party data.
Crawling and opt-outs. If any content came from the web, ask whether the crawler honored robots.txt and other machine-readable reservations, and when each source was checked. The EU AI Act requires general-purpose AI model providers to keep a copyright policy that honors rights reservations under Article 4(3) of the DSM Directive [11]. The GPAI Code of Practice's copyright chapter commits signatories to crawl only lawfully accessible content, not to circumvent paywalls, and to follow robots.txt as specified in IETF RFC 9309 [12]. Opt-outs also change: an audit of 14,000 web domains found that between 2023 and 2024 about 5% of C4 tokens, and over 28% of its most actively maintained critical sources, became fully restricted [13].
Acquisition method. Ask whether any books, articles, code or media came from shadow libraries or torrents. In Bartz v. Anthropic, the class settlement received final approval in July 2026 [14].
Repackaged open datasets. Ask how the vendor verified licenses of public datasets it redistributes. The Data Provenance Initiative's audit of 1,800+ text datasets found license omission above 70% and error rates above 50% on popular hosting sites [15], so a hosting-site license field is not enough evidence on its own.
Model-generated content and the vendor's own use. Ask which records a model generated or rewrote, with model, version and the terms in force; some providers bar using outputs to train competing models, as Anthropic's help center states for its own terms as of October 2026 [16]. Ask too whether the vendor trains its own models on your specifications or feedback. Due diligence for purchased synthetic fine-tuning data goes deeper.
Disclosure support. Ask whether the vendor will supply the fields you must disclose. As of October 2026, California's AB 2013 has required developers of generative AI systems offered to Californians to post training-data documentation since 1 January 2026, covering dataset sources or owners, copyrighted or licensed material, personal information, collection periods and synthetic data [17]. Providers of general-purpose AI models on the EU market publish training-content summaries on the AI Office template of 24 July 2025 [18]; see what EU AI Act training data summaries need from suppliers.
Colorado's SB26-189, signed 14 May 2026, will require developers of automated decision-making technology used in consequential decisions to give deployers documentation including training data categories from 1 January 2027 [19]. Machine-readable Croissant-RAI metadata, designed partly for regulatory compliance, is one way to receive such fields [20].
Record withdrawal. Ask whether the vendor can withdraw individual records after delivery; see record-level takedown obligations.
Scoring DDQ answers and setting knockout questions
Score each answer on the evidence behind it, not its tone, and treat a few gating questions as knockouts whatever the total. Weighted supplier comparisons belong in the data vendor evaluation scorecard; the DDQ decides eligibility.
| Score | Answer pattern | Action |
|---|---|---|
| 3 | Specific, evidence attached, consistent with the sample | Accept |
| 2 | Specific, evidence available under NDA before signature | Accept; make the evidence a license condition |
| 1 | Generic, such as "we hold all necessary rights" | Written follow-up; no approval yet |
| 0 | Refused, blank, or contradicted by evidence | Knockout on gating questions |
A 0, or an unresolved 1, stops approval when the vendor cannot name source classes and, under NDA, upstream suppliers; lacks a written right to sublicense upstream data for AI training; used pirate sources or circumvented access controls; leaves litigation unexplained; relies on notices that do not cover AI development; or refuses the attestation.
Illustrative example: invented to show structure; it does not describe an available dataset.
Question 3.2: "Do you acquire data in this product from third parties? If so, attach the provisions letting you sublicense it for machine-learning training."
| Response | Score | Reason |
|---|---|---|
| "About 40% of records come from two named partners. Redacted excerpts attached; clause 4.1 of each grants sublicensing for training and evaluation through 2028." | 3 | Sources, scope and term, with evidence |
| "Our agreements permit resale of derived products; excerpts available under NDA." | 2 | "Derived products" may not cover training; confirm wording |
| "We have all necessary rights to the data we provide." | 1 | Boilerplate; ask for the breakdown and excerpts |
| "Our sources are proprietary and confidential." | 0 | Knockout unless resolved under NDA |
Running the questionnaire: sequence and reviewers
A DDQ works when it is sized to the supplier's risk, routed to the reviewer who owns each section, followed up in writing and carried into the license.
- Tier the supplier. Send the full DDQ when the data holds personal information, web-derived or purchased content, or will train a deployed model; a short form can suit low-risk evaluation data.
- Issue under NDA. Require evidence for each answer, "not applicable" with a reason, and a named respondent.
- Route sections. Legal takes sections 1, 3, 5, 7 and 10; privacy takes 4; security takes 8; the ML team takes 2, 6 and 9 and checks answers against the sample. See internal approvals for a training data purchase.
- Follow up and record. Turn every score of 1 into a written question; record knockouts, conditions and accepted risks with the approver's name.
- Carry answers into the license. Attach the DDQ as an exhibit the supplier warrants as accurate; see data warranties in AI training licenses.
- Refresh it. Re-issue at renewal and after a change of control, a new upstream source or a lawsuit; see ongoing due diligence of data vendors.
Mistakes that make DDQ answers worthless
Most DDQ failures come from accepting answers nobody can check.
- Accepting the vendor's form DDQ. Vendors are advised to keep a pre-filled form ready [3]; it answers questions the vendor chose. Append your AI sections.
- Not reconciling answers with the sample. "No personal data" means little if free-text fields contain names; check answers against the sample and the evidence to request before you sign.
- Stopping at the reseller. A reseller answers for its own conduct; ask how the original data holder's answers reach you, as covered in buying through a broker or reseller.
Using a DDQ when SourceX sources the data
SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases; every release is approved by the supplying company. Every dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and diligence materials on source, rights, preparation and allowed use are prepared per dataset for the buyer's review.
Those materials relate to the subjects of sections 2 to 4, not to your whole questionnaire. In any intermediated purchase, separate the intermediary's process from the supplying business's practices, and let your vendor-risk policy decide. To test whether your requirements can be sourced this way, describe the data you need.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Evaluating a data provider for procurement?
Describe the data you need and the diligence your reviewers require. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. See how SourceX works with data buyers.
Sources
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (DDQ) with Generative AI Questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
- Lowenstein Sandler LLP, "Key considerations for alternative data and AI vendors to investment firms". https://www.lowenstein.com/media/iyrpwxij/key-considerations-for-alternative-data-and-ai-vendors-to-investment-firms.pdf
- Neudata, "2025 DDQ wrap-up: Key factors in data vendor risk" (2025). https://www.neudata.co/sentry-intelligence/2025-ddq-wrap-up-key-factors-in-data-vendor-risk
- IAPP, "Leading corporations' proposed data provenance standards aim to enhance quality of AI training data" (2023). https://iapp.org/news/a/leading-corporations-proposed-data-provenance-standards-aims-to-enhance-quality-of-ai-training-data
- California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- Washington State Legislature, "Chapter 19.373 RCW - Washington My Health My Data Act". https://app.leg.wa.gov/RCW/default.aspx?cite=19.373&full=true
- Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
- Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- Longpre et al., "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
- Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- Anthropic, "Can I use my outputs to train an AI model?" (Claude Help Center). https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
- Jain et al., "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (arXiv 2024). https://arxiv.org/pdf/2407.16883
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.