Skip to content

Procurement, samples and ongoing supply

Data provider due diligence questionnaire (DDQ) for AI training data

Quick answer

A data provider due diligence questionnaire (DDQ) is a written, vendor-level questionnaire a buyer sends before approving a data supplier. For AI training data it covers the company, how and from whom it sources data, the consents and notices behind personal data, its right to resell or sublicense, litigation history, security, and generative AI questions such as copyright opt-outs and model-generated content. Start from the FISD alternative-data template, demand evidence for each answer, and score answers before approval.

By SourceX Editorial · Updated

What a vendor DDQ covers that other diligence does not

A vendor DDQ assesses the supplier as a company across everything it sells; it does not replace the dataset-level chain-of-title check or the security review. The vendor level matters because purchased data is a supply-chain component: NIST's Generative AI Profile (AI 600-1) lists intellectual property, data privacy, and value chain and component integration among 12 generative AI risks [1].

InstrumentQuestion it answersUsual owner
Vendor DDQ (this page)Should we do business with this supplier at all?Vendor-risk or procurement lead, with legal
Dataset due diligence checklistMay this licensor license this dataset for our use?Counsel
Security review of a data supplierAre data and our information protected in transit and at rest?Security
Shortlist questionsIs this supplier worth a formal process?Data partnerships lead

The DDQ usually follows an RFP for training data and precedes any delivery beyond a small sample; the AI training data procurement guide shows the full sequence.

Adapting the FISD alternative-data DDQ to AI training data

A ready starting point is the illustrative Data Provider DDQ from FISD's Alternative Data Council, whose 2024 edition adds generative AI questions; FISD does not endorse any party using it [2]. It was written for investment managers worried about material non-public information (MNPI), so AI buyers should swap those items for copyright, consent and permitted-use questions.

Several requests carry over: a data dictionary, a small sample (fewer than 100 rows, from more than three months ago), copies of consent terms agreed with individuals, and, if the vendor buys data from others, the contract terms that permit resale [2]. A Lowenstein Sandler guide adds litigation and enforcement history, how data is sourced, processed and transmitted, and detailed provenance, and advises vendors to keep a form DDQ ready with redacted agreement excerpts or privacy notices [3]. Neudata's 2025 review benchmarks vendor DDQ responses across alternative data [4]: evidence of routine market practice, not a standard.

FISD-style itemFor AI training data
MNPI screensReplace with third-party confidential information, trade secrets and copyright
Data dictionaryKeep; one per product line, with null rates and date ranges
Sample under 100 rows, over three months oldKeep the request, but ask for a representative sample instead; see how to request a training data sample
Consent terms with individualsKeep; add each notice version in force at collection and when AI use was first disclosed
Resale termsKeep; ask for sublicensing rights covering training, evaluation and deployment
Litigation and enforcementKeep; add rightsholder complaints, opt-outs and takedown notices
Generative AI items, where presentExtend to model-generated content, crawling, the vendor's use of your data, and disclosure support

The ten sections of an AI training data DDQ

Ten sections cover vendor-level training-data risk, and every question should demand a specific answer and its evidence. Ask for answers per product line where sourcing differs, because a company-level answer hides the riskiest source.

SectionCore questionsEvidence to attach
1. Company and controlEntity, owners, affiliates, recent changes of control, license signatories, state data-broker registrationsOrganization chart, registrations
2. Sourcing modelPer product: share from first-party operations, commissioned collection, customer data, purchased data, public web, model generationSource-class breakdown, provenance metadata
3. Upstream rightsNamed upstream suppliers; resale or sublicensing rights for AI training; term and flow-down limitsRedacted contract excerpts [2][3]
4. Consent and personal dataNotices and consents at collection; de-identification method and legal standard; sensitive categoriesDated notices, consent text, method statement
5. Copyright and web contentCrawling, opt-out handling, how books, articles, code and media were acquiredCrawl policy, opt-out log
6. Generative AIModel-generated records; the vendor's own training on customer dataGenerator model, version, terms
7. Litigation and complaintsPast, current or threatened litigation, regulator inquiries, rightsholder and data-subject complaintsSummary under NDA
8. SecuritySummary only; detail goes to the security reviewAttestation reports
9. DocumentationData dictionary, dataset card, fields for your own disclosuresDatasheet or machine-readable card
10. AttestationOfficer signature; duty to notify new upstream sources, lawsuits, breaches, changes of controlSigned attestation

For section 2, a fixed structure makes answers comparable, and the Data & Trust Alliance's proposed data provenance standards supply one: a dataset's source, legal rights and privacy protections, a standardized timestamp, data type, generation method, and intended uses and restrictions [5]. SourceX's data source documentation check shows which of these records your intake already requests.

Section 4 needs sector prompts, because the statute often binds the buyer too; the privacy review of a training data vendor covers the rest:

  • CCPA "deidentified" claims. The definition requires the business to contractually bind recipients, including against re-identification [6]; ask for that language, since you will be bound by it.
  • Consumer health data. Washington's My Health My Data Act requires a separate authorization to sell consumer health data, and seller and purchaser must keep copies for six years [7].
  • Financial data. Nonpublic personal information from a financial institution carries Regulation P reuse and redisclosure limits that bind the recipient [8]; ask how the vendor received it, and under which exception if any.
  • Changed terms. Ask when AI use first appeared in the source's notice or terms: FTC staff warn that adding AI training by surreptitious, retroactive amendment may be unfair or deceptive [9], and that breaking promises not to train on customer data may create liability [10].

Generative AI questions to add to any data vendor DDQ

Whatever generative AI items a template already has, an AI training data DDQ needs detailed questions on copyright opt-outs, acquisition method, model-generated data and disclosure support. Ask every supplier, even those claiming only first-party data.

Crawling and opt-outs. If any content came from the web, ask whether the crawler honored robots.txt and other machine-readable reservations, and when each source was checked. The EU AI Act requires general-purpose AI model providers to keep a copyright policy that honors rights reservations under Article 4(3) of the DSM Directive [11]. The GPAI Code of Practice's copyright chapter commits signatories to crawl only lawfully accessible content, not to circumvent paywalls, and to follow robots.txt as specified in IETF RFC 9309 [12]. Opt-outs also change: an audit of 14,000 web domains found that between 2023 and 2024 about 5% of C4 tokens, and over 28% of its most actively maintained critical sources, became fully restricted [13].

Acquisition method. Ask whether any books, articles, code or media came from shadow libraries or torrents. In Bartz v. Anthropic, the class settlement received final approval in July 2026 [14].

Repackaged open datasets. Ask how the vendor verified licenses of public datasets it redistributes. The Data Provenance Initiative's audit of 1,800+ text datasets found license omission above 70% and error rates above 50% on popular hosting sites [15], so a hosting-site license field is not enough evidence on its own.

Model-generated content and the vendor's own use. Ask which records a model generated or rewrote, with model, version and the terms in force; some providers bar using outputs to train competing models, as Anthropic's help center states for its own terms as of October 2026 [16]. Ask too whether the vendor trains its own models on your specifications or feedback. Due diligence for purchased synthetic fine-tuning data goes deeper.

Disclosure support. Ask whether the vendor will supply the fields you must disclose. As of October 2026, California's AB 2013 has required developers of generative AI systems offered to Californians to post training-data documentation since 1 January 2026, covering dataset sources or owners, copyrighted or licensed material, personal information, collection periods and synthetic data [17]. Providers of general-purpose AI models on the EU market publish training-content summaries on the AI Office template of 24 July 2025 [18]; see what EU AI Act training data summaries need from suppliers.

Colorado's SB26-189, signed 14 May 2026, will require developers of automated decision-making technology used in consequential decisions to give deployers documentation including training data categories from 1 January 2027 [19]. Machine-readable Croissant-RAI metadata, designed partly for regulatory compliance, is one way to receive such fields [20].

Record withdrawal. Ask whether the vendor can withdraw individual records after delivery; see record-level takedown obligations.

Scoring DDQ answers and setting knockout questions

Score each answer on the evidence behind it, not its tone, and treat a few gating questions as knockouts whatever the total. Weighted supplier comparisons belong in the data vendor evaluation scorecard; the DDQ decides eligibility.

ScoreAnswer patternAction
3Specific, evidence attached, consistent with the sampleAccept
2Specific, evidence available under NDA before signatureAccept; make the evidence a license condition
1Generic, such as "we hold all necessary rights"Written follow-up; no approval yet
0Refused, blank, or contradicted by evidenceKnockout on gating questions

A 0, or an unresolved 1, stops approval when the vendor cannot name source classes and, under NDA, upstream suppliers; lacks a written right to sublicense upstream data for AI training; used pirate sources or circumvented access controls; leaves litigation unexplained; relies on notices that do not cover AI development; or refuses the attestation.

Illustrative example: invented to show structure; it does not describe an available dataset.

Question 3.2: "Do you acquire data in this product from third parties? If so, attach the provisions letting you sublicense it for machine-learning training."

ResponseScoreReason
"About 40% of records come from two named partners. Redacted excerpts attached; clause 4.1 of each grants sublicensing for training and evaluation through 2028."3Sources, scope and term, with evidence
"Our agreements permit resale of derived products; excerpts available under NDA."2"Derived products" may not cover training; confirm wording
"We have all necessary rights to the data we provide."1Boilerplate; ask for the breakdown and excerpts
"Our sources are proprietary and confidential."0Knockout unless resolved under NDA

Running the questionnaire: sequence and reviewers

A DDQ works when it is sized to the supplier's risk, routed to the reviewer who owns each section, followed up in writing and carried into the license.

  1. Tier the supplier. Send the full DDQ when the data holds personal information, web-derived or purchased content, or will train a deployed model; a short form can suit low-risk evaluation data.
  2. Issue under NDA. Require evidence for each answer, "not applicable" with a reason, and a named respondent.
  3. Route sections. Legal takes sections 1, 3, 5, 7 and 10; privacy takes 4; security takes 8; the ML team takes 2, 6 and 9 and checks answers against the sample. See internal approvals for a training data purchase.
  4. Follow up and record. Turn every score of 1 into a written question; record knockouts, conditions and accepted risks with the approver's name.
  5. Carry answers into the license. Attach the DDQ as an exhibit the supplier warrants as accurate; see data warranties in AI training licenses.
  6. Refresh it. Re-issue at renewal and after a change of control, a new upstream source or a lawsuit; see ongoing due diligence of data vendors.

Mistakes that make DDQ answers worthless

Most DDQ failures come from accepting answers nobody can check.

  • Accepting the vendor's form DDQ. Vendors are advised to keep a pre-filled form ready [3]; it answers questions the vendor chose. Append your AI sections.
  • Not reconciling answers with the sample. "No personal data" means little if free-text fields contain names; check answers against the sample and the evidence to request before you sign.
  • Stopping at the reseller. A reseller answers for its own conduct; ask how the original data holder's answers reach you, as covered in buying through a broker or reseller.

Using a DDQ when SourceX sources the data

SourceX sources operational datasets from US companies and manages the commercial process, including licensing agreements and ongoing purchases; every release is approved by the supplying company. Every dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place, and diligence materials on source, rights, preparation and allowed use are prepared per dataset for the buyer's review.

Those materials relate to the subjects of sections 2 to 4, not to your whole questionnaire. In any intermediated purchase, separate the intermediary's process from the supplying business's practices, and let your vendor-risk policy decide. To test whether your requirements can be sourced this way, describe the data you need.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Evaluating a data provider for procurement?

Describe the data you need and the diligence your reviewers require. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. See how SourceX works with data buyers.

Sources

  1. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
  2. FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (DDQ) with Generative AI Questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
  3. Lowenstein Sandler LLP, "Key considerations for alternative data and AI vendors to investment firms". https://www.lowenstein.com/media/iyrpwxij/key-considerations-for-alternative-data-and-ai-vendors-to-investment-firms.pdf
  4. Neudata, "2025 DDQ wrap-up: Key factors in data vendor risk" (2025). https://www.neudata.co/sentry-intelligence/2025-ddq-wrap-up-key-factors-in-data-vendor-risk
  5. IAPP, "Leading corporations' proposed data provenance standards aim to enhance quality of AI training data" (2023). https://iapp.org/news/a/leading-corporations-proposed-data-provenance-standards-aims-to-enhance-quality-of-ai-training-data
  6. California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  7. Washington State Legislature, "Chapter 19.373 RCW - Washington My Health My Data Act". https://app.leg.wa.gov/RCW/default.aspx?cite=19.373&full=true
  8. Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  9. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  10. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  11. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  12. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  13. Longpre et al., "Consent in Crisis: The Rapid Decline of the AI Data Commons" (2024). https://arxiv.org/pdf/2407.14933
  14. Authors Alliance, "Bartz v. Anthropic Settlement Receives Final Approval" (2026). https://www.authorsalliance.org/2026/07/21/bartz-v-anthropic-settlement-receives-final-approval/
  15. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  16. Anthropic, "Can I use my outputs to train an AI model?" (Claude Help Center). https://support.claude.com/en/articles/12326764-can-i-use-my-outputs-to-train-an-ai-model
  17. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  18. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  19. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
  20. Jain et al., "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (arXiv 2024). https://arxiv.org/pdf/2407.16883

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data