Skip to content

Procurement, samples and ongoing supply

Privacy Review of a Training Data Vendor

Quick answer

A privacy review of a training data vendor confirms four things before purchase: the lawful basis or consent under which the records were collected and can be licensed for model training, what notices individuals received, how personal data was de-identified and what evidence shows it worked, and whether any special categories such as health, biometric, children's or financial data require escalation. The output is a signed memo with conditions, not a questionnaire score, and it should rest on documents the vendor produces rather than assertions.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Where the privacy review sits in the purchase

The privacy review runs after the dataset is scoped and a sample exists, and before the license is signed. It is distinct from the security review of a training data supplier, which covers transfer, storage and access controls, and from the broader internal approvals for buying training data, where privacy is one of several sign-offs. The privacy team's question is narrower: may this organization lawfully receive and train on these records, and is the residual identification risk acceptable?

Starting from a structured questionnaire saves time. The FISD Alternative Data Council's data provider DDQ, updated in 2024 with generative AI questions, already asks providers about consent terms where data concerns individuals or their devices and about how personal data is handled [1]. Law-firm guidance for buyers of alternative data adds questions about how data is sourced, processed and transmitted, how PII is treated at each stage, and any past or threatened enforcement [2]. Use these as the intake layer, alongside the training data due diligence checklist, then go deeper on the evidence below. Our data provider due diligence questionnaire adapts that structure for training data.

The vendor should show the original collection context, the notice text in force at the time, and the legal basis that allows disclosure to you for model training. A data processing agreement alone does not answer this; it governs the vendor's processing, not whether the upstream collection permits your downstream use.

Request these specific items:

  • Notice and terms snapshots. The privacy notice and terms of service versions in effect during each collection window, with effective dates, so you can map records to the notice that covered them.
  • Consent records where consent is the basis. For opt-in collections (recorded calls, device telemetry, panel data), a sample of consent logs with fields such as subject pseudonym, consent text version, timestamp, channel and withdrawal status. Some licensors include consent-record checks in their final QA before delivery [3]; ask to see that output rather than accept the claim.
  • Basis under GDPR where EU residents are in scope. EDPB Opinion 28/2024 addresses legitimate interest as a basis for developing AI models and the consequences when development relied on unlawfully processed data [6]. Ask whether the vendor ran a legitimate interest assessment and whether it covers disclosure to third parties for training.
  • Sector redisclosure limits. Nonpublic personal information that a vendor received from a financial institution under a GLBA exception can only be used for the purpose it was received for, which usually rules out onward licensing for training [11]. Our page on financial records and GLBA covers the analysis.

Withdrawn consent is the most common gap. Ask how withdrawals after delivery propagate, and whether the license requires the vendor to notify you of records to suppress.

De-identification method and evidence

A vendor's statement that data is "anonymized" is not evidence; ask for the method, the parameters, the residual-risk assessment and test results on the delivered files. Legal standards differ: GDPR Recital 26 treats data as anonymous only when individuals are no longer identifiable by means reasonably likely to be used [7], while California's CCPA definition of deidentified information also requires the holder to publicly commit not to re-identify and to contractually bind recipients to the same [5]. That second condition means your license must carry re-identification prohibitions downstream.

For health data, HIPAA allows two routes: Safe Harbor, removing 18 listed identifiers with no actual knowledge that the remainder identifies someone, or Expert Determination, where a qualified expert documents that the risk is very small [4]. If the vendor relies on Expert Determination, request the report itself and review it with our guide to reviewing a HIPAA Expert Determination report.

For other data, ask for evidence at three layers:

  1. Direct identifiers. Which detector was used (for example, Microsoft Presidio recognizers, regex rules, or a fine-tuned NER model), at what confidence threshold, and what was done with hits: deletion, masking, or consistent surrogate replacement. Presidio's own documentation warns that ML-based detection cannot guarantee it finds all sensitive information [12], so a measured recall on a labeled sample matters more than the tool name.
  2. Quasi-identifiers. How dates, ZIP codes, job titles, free-text mentions and rare categories were generalized. NIST SP 800-188 describes both traditional techniques and formal methods such as differential privacy, and cautions about the limits of traditional de-identification [8].
  3. Structural uniqueness. Event logs, support tickets and workflow records can identify people through sequences and timestamps even after names are removed. Research on process-mining logs measured how unique individual cases become from these attributes [9]. Ask whether the vendor measured uniqueness on sequences, not just on columns.

Free text deserves its own test. Signatures, email footers, ticket bodies quoting account numbers and names inside filenames routinely survive column-level scrubbing. See the de-identification playbook and the data anonymization glossary entry for method detail.

Sensitive categories that trigger escalation

Health, biometric, genetic, children's, precise location and financial account data should route to counsel and, where applicable, a formal impact assessment before approval. These categories carry their own statutes and consent rules that ordinary vendor terms rarely satisfy.

Washington's My Health My Data Act is a useful test of scope: its definition of consumer health data reaches biometric and genetic data and precise location that could indicate an attempt to obtain health services, well beyond HIPAA-covered records [10]. Recorded voice and video of workers can carry biometric identifiers; call recordings can carry payment card numbers; support histories can carry diagnoses typed by customers. Ask the vendor for a category scan of the sample, then run your own.

Where GDPR applies, training on personal data from a third party will often meet the conditions for a data protection impact assessment under the Regulation [7]. Request the vendor's own DPIA or risk assessment as an input to yours, not a substitute. For California, see our page on CCPA risk assessments for AI training.

Privacy review evidence checklist

Use this as the request list you send the vendor and the record of what came back. Each row should end in "received and acceptable", "received with conditions", or "not received".

Illustrative example: invented to show structure; it does not describe an available dataset.

#Evidence itemWhat good looks likeRed flag
1Data inventory and data dictionaryField list with personal data flag per column and per free-text field"No personal data" with free-text columns present
2Collection notices and termsVersioned snapshots with effective dates covering every collection windowOnly the current notice
3Legal basis or consent recordsBasis per source; consent log sample with text version and withdrawal statusBasis stated as "customer agreement" with no text
4Upstream rights chainContracts or terms permitting disclosure for trainingData received from a third party under a narrower purpose
5De-identification method specTools, rules, thresholds, surrogate strategy, version and run dateMethod described only as "anonymized"
6Residual-risk assessmentUniqueness or k-anonymity measures on quasi-identifiers and sequences; Expert Determination report for PHINo measurement after redaction
7Detection test resultsRecall and precision on a hand-labeled sample, by entity typeSpot check with no counts
8Sensitive category scanResults for health, biometric, children's, location, financial termsScan not run on free text
9Withdrawal and deletion handlingDefined process to notify buyer of records to suppressNo mechanism after delivery
10Enforcement and complaint historyDisclosure of regulator inquiries or litigation on the data [2]Refusal to answer

Pair the checklist with an independent test: draw a few hundred records from the sample you requested, have two reviewers label residual identifiers, and compare against the vendor's claimed recall. If the sample is small or curated, check whether it is representative of the full dataset.

Writing the sign-off memo

The approval should be a short memo that states the decision, the conditions, and the evidence relied on, so a later audit or model-deletion inquiry can reconstruct why the purchase was approved. Regulators have required deletion of models trained on improperly obtained data, which is why the paper trail matters; see algorithmic disgorgement.

A workable memo has five parts:

  • Scope: dataset, record count range, fields, jurisdictions of data subjects, and intended model uses.
  • Findings: one line per checklist row, with the document reference.
  • Residual risk: what the testing showed and what risk remains; no method is perfect.
  • Conditions: license terms you require, such as re-identification prohibition, suppression on withdrawal, restriction on joining with other datasets, and memorization testing before release (see training-data extraction risk).
  • Owner and review date: who re-reviews when the vendor delivers a new tranche or changes its method.

For ongoing supply, treat each new delivery as a delta review: confirm the method version and rerun detection tests rather than restarting from zero. The procurement hub covers how this fits renewals and acceptance criteria.

How SourceX supports a buyer's privacy review

SourceX sources operational datasets from US companies on request and manages the licensing process, so privacy evidence is assembled per dataset rather than pulled from a stock catalog. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect; health records require HIPAA de-identification by Safe Harbor or Expert Determination. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which gives your privacy team documents to review. Buyers can describe the data they need to SourceX.

Request training data that can go through privacy review

SourceX finds US businesses that hold the data you describe, assesses data and licensing permissions, and agrees allowed uses in a license; nothing is contracted until the supplier agrees, and a request does not guarantee a match. Delivery happens through private, access-controlled workflows only after an executed agreement and supplier approval. Start a buyer request at SourceX.

Frequently asked questions

Can we rely on the vendor's DPIA instead of running our own?

No. A vendor's DPIA covers the vendor's processing. Your organization becomes a controller for its own training use and needs its own assessment, which can incorporate the vendor's document as evidence [7].

Is pseudonymized data out of scope for privacy review?

No. Consistent surrogates and hashed IDs keep records linkable, so under GDPR the data usually remains personal data [7]. Review it as personal data with reduced risk, not as anonymous data.

What if the vendor will not share its de-identification report?

Treat the evidence item as not received. You can approve with conditions, such as independent testing on a sample under an evaluation license or NDA, but record that the method was not verified.

Sources

  1. FISD Alternative Data Council, "Data Provider Due Diligence Questionnaire (DDQ) with GenAI Questions" (2024). https://fisd.net/wp-content/uploads/2024/02/FISD-Alternative-Data-Council-Due-Diligence-Questionnaire-with-GenAI-Questions-022824.docx
  2. Lowenstein Sandler, "Key Considerations for Alternative Data and AI Vendors to Investment Firms". https://www.lowenstein.com/media/iyrpwxij/key-considerations-for-alternative-data-and-ai-vendors-to-investment-firms.pdf
  3. Pocstock, "The Dataset Licensing Process: From Inquiry to Delivery". https://support.pocstock.com/en/articles/14772883-the-dataset-licensing-process-from-inquiry-to-delivery
  4. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  5. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  6. European Data Protection Board, "Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models" (2024). https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
  7. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  8. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  9. arXiv (Nuñez von Voigt et al.), "Quantifying the Re-identification Risk of Event Logs for Process Mining" (2020). https://arxiv.org/pdf/2003.10707
  10. Washington State Legislature, "Chapter 19.373 RCW - Washington My Health My Data Act". https://app.leg.wa.gov/RCW/default.aspx?cite=19.373&full=true
  11. Federal Trade Commission, "How To Comply with the Privacy of Consumer Financial Information Rule of the Gramm-Leach-Bliley Act". https://www.ecfr.gov/current/title-16/chapter-I/subchapter-C/part-313/section-313.11
  12. Microsoft (presidio project, indexed on pkg.go.dev), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data