Privacy, de-identification and sensitive data
De-identified data for AI training: a buyer's guide to privacy-safe licensed datasets
Quick answer
De-identified data for AI training is personal data transformed so that it no longer identifies people under a specific legal test, and the test changes with the regime. HIPAA accepts Safe Harbor or Expert Determination, the GDPR still treats pseudonymised data as personal data, and the CCPA requires the holder to bind every recipient by contract. Before licensing, ask which standard was met, by what method, with what measured miss rate and what you must sign, then plan for risk left in the trained model.
By SourceX Editorial · Updated
One label, three legal tests: HIPAA, GDPR and the CCPA
A dataset described as "de-identified" can pass one regime's test and fail another's, because HIPAA, the GDPR and the CCPA ask different questions about the same records. The table states each test as of October 2026.
| Regime | Test the data must meet | Still regulated? | What the recipient takes on | Read next |
|---|---|---|---|---|
| HIPAA Safe Harbor, 45 CFR 164.514(b)(2) | Remove 18 identifiers of the person and of relatives, employers and household members; no actual knowledge that the rest identifies anyone [1][2] | No; not protected health information (PHI) [1] | Nothing under HIPAA; license terms apply | Safe Harbor vs Expert Determination |
| HIPAA Expert Determination, 164.514(b)(1) | An expert finds a "very small" risk that an anticipated recipient could identify anyone and documents the analysis [2]; no numeric threshold [1] | No, while its assumptions hold | Conditions in the report, such as access limits | Reviewing an Expert Determination report |
| HIPAA limited data set, 164.514(e) | Direct identifiers removed; dates, town or city, state and ZIP code may remain [2] | Yes; still PHI, only for research, public health or health care operations [2] | A data use agreement, including no identifying or contacting individuals [2] | Limited data sets and DUAs |
| GDPR anonymous information, Recital 26 | No one identifiable by means reasonably likely to be used [3] | No | Little, if the test genuinely holds | Anonymised or pseudonymised? |
| GDPR pseudonymised data, Article 4(5) | Attributable to a person only with separately held additional information [3] | Yes for the controller; for recipients, see EDPS v SRB [4] | Lawful basis, Article 9 condition for health or biometric data, transfer safeguards [3] | Pseudonymised data from the recipient's side |
| CCPA "deidentified", Cal. Civ. Code 1798.140(m) | Not reasonably linkable to a consumer; the holder takes reasonable measures, publicly commits not to re-identify and contractually binds recipients [5] | No, while all three conditions hold | The same commitments, by contract [5] | CCPA obligations a buyer inherits |
In the EU, the Court of Justice held in EDPS v SRB (C-413/23 P, 4 September 2025) that sufficiently strong pseudonymisation can leave data personal for the controller holding the key but not for a recipient with no means to reverse it or identify people otherwise; the assessment is case by case [4]. As of September 2026, Digital Omnibus proposals to narrow the definition of personal data and to confirm legitimate interest for AI development were still proposals, not law [6].
In the UK, the ICO's anonymisation guidance (published March 2025; under review) recommends the "motivated intruder" test [7][8] and warns that re-identifying people from pseudonymised or ineffectively anonymised data can be a criminal offence [9]. Compare the four labels in de-identified, anonymised, pseudonymised and aggregated; SourceX's glossary defines data anonymization.
Sector and state rules that follow particular records
Some records carry rules that a general de-identification claim does not settle, so identify the source system and the people in it first.
- Education records. FERPA allows release without consent once all personally identifiable information is removed and the releaser reasonably determines, across multiple releases and other available information, that no student is identifiable [10] (FERPA guide).
- Financial records. Nonpublic personal information received under a Regulation P exception may be used only for the purpose it was received, even by a recipient that is not a financial institution [11]; ask counsel whether de-identified transactions still qualify (GLBA guide).
- Substance use disorder records. HHS's 2024 final rule aligned parts of 42 CFR Part 2 with HIPAA, with a compliance date of 16 February 2026 [12]; ask whether any records came from a Part 2 program (Part 2 records).
- Biometrics. Illinois' BIPA treats voiceprints and face geometry scans as biometric identifiers, requires a written release before collection and bars profiting from them [13]. Removing names does not help when the voice or face is the identifier (biometric data guide).
- Washington consumer health data. The My Health My Data Act counts inferences drawn from non-health data as consumer health data, and selling it needs a separate authorization that seller and purchaser must both keep for six years [14].
Why stripping identifiers does not make data anonymous
Removing names and account numbers leaves quasi-identifiers, such as dates, places, job titles and rare events, that outside data can link back to a person. Sweeney's k-anonymity requires each person to be indistinguishable from at least k-1 others on those attributes, and even then attacks can succeed when release policies are ignored [15]. Rocher et al. estimated with a generative model, not from actual re-identifications, that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes, so releasing only a sample does not make data anonymous [16].
NIST documents real re-identification cases and notes that medical text, photographs and genetic data are harder to de-identify than tables [17]. SP 800-188 (2023) cautions about the limits of traditional methods compared with formal ones such as differential privacy [18]. Watch for indirect identifiers in business text and linkage risk when combining licensed sets; SourceX's explainer is de-identified data truly anonymous? gives the short version.
Where identifiers hide in text, audio, images and tables
The method has to match the modality, because identifiers sit in different places in a support ticket, a call recording, a video frame and a claims table. Test each type's weak point in a sample.
| Data type | Where identifiers hide | Typical method | Known weak point | Read next |
|---|---|---|---|---|
| Free text: tickets, email, chat, clinical notes | Names, contacts and account numbers in prose, signatures, quoted headers | Entity detection, then masks or consistent surrogates | Detectors miss entities; Presidio says it cannot guarantee finding all sensitive data [19] | PII redaction for LLM data; masking vs surrogates |
| Call and meeting audio | Spoken names and card numbers; the voice itself | Audio masking aligned to transcripts; speaker anonymization | In one benchmark, a naive attacker's equal error rate (higher is more private) reached about 52.12%; an attacker trained on anonymized speech cut it to about 10.7% [20] | Spoken PII in call recordings; speaker anonymization |
| Images and video | Faces, plates, screens and papers in frame; EXIF GPS tags | Blur, pixelation or synthetic replacement | Missed small or occluded faces; Gaussian blur can be partly reversed [21] | Video anonymization; faces: consent or anonymization |
| Medical images (DICOM) | Header attributes, private tags, burned-in pixel text | PS3.15 Annex E confidentiality profiles | Profiles do not guarantee removal of all identifying information [22] | HIPAA-compliant training data |
| Scanned documents | Pixels, the OCR text layer, file metadata | Redact all three layers | Boxes over the image while the OCR layer stays searchable | PII in scanned documents |
| Tables and transactions | Quasi-identifier and free-text columns, stable keys | Generalization, suppression, keyed tokens | Rare value combinations stay unique [16] | De-identifying tabular data |
| Screen recordings and agent trajectories | On-screen records, typed input, URLs, tool-call arguments | Frame- and event-level redaction | Text visible in frames but absent from event logs | Screens and computer-use trajectories |
Privacy risk that carries into the trained model
Training can memorize personal data that survived de-identification, so the model itself can become the disclosure. Carlini et al. extracted hundreds of verbatim training sequences from GPT-2, including contact details, using only queries [23], and Nasr et al. (ICLR 2025) recovered thousands of training examples from aligned production models [24]. Alignment is not a privacy control.
A law-firm summary of EDPB Opinion 28/2024 (17 December 2024) reports that a model trained on personal data cannot be presumed anonymous; it is anonymous only if extracting personal data from it, directly or through queries, is insignificant [25] (EDPB Opinion 28/2024 for data buyers). In the US, the FTC's 2021 Everalbum order required deletion of models and algorithms built from users' photos and videos [26].
Residual personal data therefore decides whether you can ship weights or open a public API. Differential privacy bounds what any single record can reveal, usually at some cost in accuracy [18]; see memorization risk, differential privacy for LLM fine-tuning and releasing weights trained on personal data.
The privacy evidence to request before you sign
Ask for evidence that a named standard was met on this dataset, not a description of the supplier's general process.
- Standard and jurisdiction. The regime claimed and where the data subjects live.
- Method. Detector and version, entity types, masking or surrogate policy, date and location rules, and dropped fields.
- Measured misses. Recall per entity type on a labeled sample from this dataset, with sample size and labeler.
- Expert Determination report, if used. Methods, results, recipient assumptions and date [2]; HHS notes that because technology and information availability change over time, re-examination may be appropriate [1].
- Key custody. Who holds any re-identification code or pseudonymisation key; HIPAA permits such codes only under stated conditions [2].
- Residual risk. Quasi-identifiers left in and any re-identification test results; SP 800-188 recommends such studies and a disclosure review board [18].
- Your obligations. No re-identification, no linkage, onward-transfer limits and a duty to report found personal data (re-identification prohibition clauses).
- Source-side basis. Notice and consent records, and any Article 9 condition for special categories (consent and notice records).
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset_id: support-tickets-r1
source_system: helpdesk export (ticket bodies and replies; attachments excluded)
standard_claimed: CCPA 1798.140(m) deidentified # GDPR status: not claimed
data_subject_locations: [US-CA, US-TX, US-NY]
method:
detection: NER model + regex rules (card numbers, SSNs)
entity_types: [PERSON, EMAIL, PHONE, ACCOUNT_ID, STREET_ADDRESS, CARD_NUMBER]
replacement: consistent surrogates within each ticket thread
dates: shifted per customer; intervals preserved
dropped_fields: [attachments, agent_email, ip_address]
qa_sample:
records_reviewed: 2000
labeling: two annotators, adjudicated
recall_by_entity: {PERSON: 0.97, EMAIL: 1.00, PHONE: 0.99, ACCOUNT_ID: 0.95}
residual_risk_notes: product names and small-town names retained
reidentification_key: none retained by supplier
recipient_obligations: [no re-identification, no linkage to customer records, report found personal data]
For datasets sourced through SourceX, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded per dataset, and a sample is checked after processing; health records must meet HIPAA Safe Harbor or Expert Determination before a license is considered. Diligence materials on source, rights, preparation and allowed use are prepared per dataset. No method is perfect, so keep your own checks and describe the records and privacy standard you need. See the full de-identification evidence package checklist; suppliers can follow SourceX's playbook for de-identifying company data.
Controls that stay with the buyer after delivery
Supplier de-identification lowers risk at the source; your controls decide whether it stays low once the data meets other data and a model.
- Scan before training. Check inputs and context fields, not only targets (scanning a corpus before fine-tuning).
- Block linkage. Keep licensed sets out of joins with CRM data, logs or other licensed sets unless a fresh risk assessment covers the combination.
- Contain access. SP 800-188 lists protected, non-public enclaves among its data-sharing models [18]; train inside one where you can.
- Test the model. Run extraction probes or planted canary strings before any external release.
- Rehearse the incident. Agree how found personal data is quarantined, reported and deleted (response playbook).
What de-identification costs the model, and how to specify around it
Every method removes signal, so list the fields your model needs before choosing a standard. Safe Harbor removes all elements of dates tied to the person except the year and geographic units smaller than a state (three-digit ZIP codes for geographic units containing more than 20,000 people may be retained), and groups ages over 89 [2]. That can break a readmission or seasonal-demand model; Expert Determination has no fixed list, so an expert may allow more detail if the risk stays very small [1][2] (what Safe Harbor dates and ZIP codes cost).
In vision, a CVPR 2023 workshop study found that face-only anonymization of detection datasets caused a minimal accuracy drop, while traditional anonymization hurt noticeably, especially when whole bodies were masked [27] (blurring evidence). Privacy constraints also narrow coverage: a paper in Stanford's GRACE journal argues that de-identified health datasets used for foundation models are often outdated and demographically limited [28]. For high-risk systems, EU AI Act Article 10(3) requires data that is sufficiently representative of the people the system serves [29], so check coverage after de-identification; the compliance hub gives application dates as amended in 2026.
Start here: privacy guides by question
Each question below has its own guide.
| Your question | Read |
|---|---|
| Business associate agreement or data license? | BAA vs data license for AI developers |
| How is re-identification risk measured? | Risk assessment methods to request |
| Is synthetic data from licensed records still personal data? | Synthetic data privacy risk |
| US-sourced data with EU or UK subjects? | Transfer mechanisms for AI teams |
| Which US state standards apply? | State deidentification standards compared |
| Can anonymized business data be licensed at all? | Can anonymized data be licensed? |
For the rest of the purchase, start at the AI data buyer's hub.
Mistakes that turn a de-identified dataset into a liability
Most privacy failures in data deals come from trusting a label instead of the test behind it.
- Treating HIPAA de-identification as GDPR anonymisation. Safe Harbor answers a US question; any records in the same file that fall under the GDPR still face the Recital 26 test [3].
- Assuming pseudonymised data is non-personal for you. After EDPS v SRB the question is whether you could re-identify, including with data you already hold [4].
- Signing flow-down terms you cannot honor. Joining records to customer data puts the CCPA no-re-identification commitment at risk [5].
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Sourcing data that contains personal information?
Describe the records you need, where the people in them are located, and the privacy standard your reviewers require. SourceX looks for US companies that hold that data, checks the data and the supplier's licensing permissions, and manages the license and delivery through private, access-controlled workflows. Share your data and privacy requirements.
Guides in this section
- Anonymised vs Pseudonymised Training Data Under GDPRHow to classify licensed AI training data under GDPR: the Recital 26 test, EDPB pseudonymisation guidance, EDPS v SRB, and a classification worksheet.
- BAA or Data License? Health Data Contracts for AI TeamsWhen an AI developer needs a HIPAA business associate agreement, a data use agreement or a data license for health data, and why training changes it.
- Biometric Data in AI Training Sets: BIPA, CUBI, WashingtonBuyer checklist for faces, voices and other biometric identifiers in AI training data: BIPA, Texas CUBI and Washington consent, retention and evidence.
- CCPA Deidentified Data: Obligations AI Buyers InheritWhat CCPA/CPRA section 1798.140(m) requires of deidentified data, which conditions flow to AI data buyers by contract, and how to evidence compliance.
- De-identification Documentation Checklist for Data BuyersDocuments to request before licensing de-identified data: field inventory, method statement, QA results, expert report, residual risk and attestation.
- De-identified vs Anonymized Data: US and EU Legal StatusHow HIPAA, CCPA and GDPR classify the same licensed dataset as de-identified, pseudonymized, anonymous or aggregated, and which obligations follow it.
- Does Face and Plate Blurring Hurt Vision Model Training?What studies show about face, plate and full-body anonymization in computer vision training data, and which method and checks to request from a supplier.
- EDPB Opinion 28/2024: When Is an AI Model Anonymous?What EDPB Opinion 28/2024 requires before a model trained on personal data counts as anonymous, and which training-data records buyers should keep.
- HIPAA Limited Data Sets and DUAs for AI Model DevelopmentWhen a HIPAA limited data set under a data use agreement can support AI model development, what the DUA must contain, and when to use de-identified data.
- HIPAA-Compliant AI Training Data: What Buyers Must VerifyWhat a HIPAA-compliant label must mean when you license health data for AI: the de-identification method, expert evidence, and checks before signing.
- PII Redaction for LLM Training Data: Pipeline and MetricsHow to build and validate a PII redaction pipeline for LLM training text: detector layers, replacement choices, per-entity recall and residual audits.
- Pseudonymised data as a recipient after EDPS v SRBHow C-413/23 P lets a recipient treat pseudonymised data as non-personal, and the key, contract and auxiliary-data evidence an AI data buyer should keep.
- Re-identification Prohibition Clauses in Data LicensesWhat a re-identification prohibition clause asks AI data buyers to sign: no re-identification, linkage or contact, plus notice, flow-down and controls.
- Re-identification Risk Assessment for Licensed DatasetsWhich re-identification risk assessment methods to request from a data supplier, what the report should contain, and how to read its metrics.
- Redacting PII in Screen Recordings and Agent TrajectoriesWhere personal data hides in computer-use trajectories (pixels, DOM, keystrokes, URLs, tool-call arguments) and how to redact every channel consistently.
- Reviewing a HIPAA Expert Determination Report: Buyer GuideHow AI buyers test a supplier's HIPAA Expert Determination report: recipient, environment, fields, expiry, method evidence and model-training scope.
- Safe Harbor vs Expert Determination for AI Training DataWhich HIPAA de-identification method to require for AI training data: what Safe Harbor and Expert Determination do to dates, geography and clinical notes.
- Training-Data Extraction and Memorization Risk in LLMsHow extraction attacks recover personal data from LLMs trained on licensed sensitive text, and which mitigations and tests buyers should require.
- 42 CFR Part 2 Records in AI Training Data After 2024 RuleHow substance use disorder records under 42 CFR Part 2 can enter AI training data after the 2024 rule: HIPAA de-identification, segmentation and checks.
- Annotation Vendor Access to PHI and PII: BAAs and ControlsConditions for letting labeling vendors touch licensed PHI or PII: BAAs and DPAs, license onward-sharing, offshore annotators and least-privilege tooling.
- Combining De-identified Datasets: Linkage and Mosaic RiskWhen joining licensed, public or internal datasets voids de-identification: how linkage and mosaic risk arise, and the controls AI data buyers should run.
- Data Clean Rooms for AI Model Training: A Buyer's TestWhen a data clean room can replace delivery of sensitive records for AI training or evaluation, what it handles poorly, and the license terms to negotiate.
- Data Minimization for AI Training Data: A Field Spec GuideHow ML leads write a need-to-train data specification: which fields, history and granularity to request, and which fields should never leave the supplier.
- De-identifying Clinical Notes for LLM Training: PHI RiskWhat PHI detection in clinical free text must achieve before SFT: entity coverage, surrogates vs tags, recall evidence and residual leakage risk.
- De-Identifying Tabular and Transaction Data for MLHow to de-identify tabular and transactional data for ML: quasi-identifiers, generalization, suppression, keyed pseudonyms and free-text columns.
- Deidentified Data Under US State Privacy Laws ComparedHow California, Virginia, Colorado, Texas and Nebraska define deidentified and pseudonymous data, and one contract standard for a multi-state AI dataset.
- Differential Privacy for LLM Fine-Tuning: DP-SGD in PracticeHow DP-SGD fine-tuning works on sensitive licensed text: clipping, noise, epsilon choices, utility loss, user-level units and canary tests to verify it.
- Differentially Private Synthetic Text: Methods and LimitsHow DP synthetic text is built from a licensed sensitive corpus: DP-SGD generators vs API-based Private Evolution, epsilon budgets, utility, licensing.
- DOJ Bulk Data Rule (28 CFR 202) for AI Training Data BuyersHow 28 CFR Part 202 treats licensed US training data: bulk thresholds, de-identified data, countries of concern and the onward-transfer clause buyers sign.
- DPIA for AI Training Data: Buyer Template for Licensed SetsA buyer-side DPIA template for licensed AI training data: when GDPR Article 35 applies, what to document, supplier evidence to attach and residual risk.
- Employee Email, Chat and Meeting Data: AI Privacy RisksBuyer-side privacy guide to employee email, Slack and meeting transcripts in AI training data: notice laws, PII sources, exclusions and pseudonymization.
- Erasure and Opt-Out Requests After Training Data DeliveryHow GDPR erasure, objection and CCPA deletion requests reach a buyer of licensed training data, and what to do with the dataset, derivatives and models.
- EU and UK Personal Data in US Datasets: Transfer RulesWhen EU or UK personal data inside licensed US company records triggers transfer rules, and how DPF, SCCs or the UK IDTA apply to AI training buyers.
- Face Images in AI Training Data: Consent or AnonymizeHow CV teams decide between consented face data and anonymized footage: BIPA and Texas releases, bystanders, blur vs. synthesis, and what to request.
- FERPA De-identified Student Data for AI TrainingWhen de-identified education records can be licensed for AI training under FERPA, which school contracts block it, and how to check the chain of rights.
- Found PII in a Licensed Training Dataset: Response PlaybookResidual personal data in a delivered dataset: how to quarantine, scope, notify the supplier, get a corrected delivery and decide whether to retrain.
- GLBA and De-identified Financial Data for AI TrainingHow GLBA reuse and redisclosure limits apply to licensed bank, lending and payments records for AI training, and when de-identified data falls outside NPI.
- Indirect Identifiers in Business Text: Find and Treat ThemHow to find and treat indirect identifiers in tickets, emails and incident notes: job titles, rare events, small places, dates and B2B account names.
- Is Synthetic Data Personal Data? Privacy Tests for BuyersWhen synthetic records generated from licensed personal data remain personal data, how they leak, and which privacy tests to require.
- Legitimate Interest for AI Training on Licensed DataHow a DPO documents a legitimate interests assessment for training on licensed personal data: purpose, necessity, balancing and buyer mitigations.
- LLM PII Redaction vs NER and Regex: Accuracy and RiskWhen LLM-based PII redaction beats NER and regex for training text, what it costs at corpus scale, where it leaks, and how to evaluate a hybrid pipeline.
- LLM Re-identification Risk in De-identified TextHow LLMs infer identities and attributes from residual context in de-identified text, and how to run an LLM-assisted intruder test before licensing.
- Membership Inference and Canary Audits for Fine-Tuned ModelsDesign a pre-release leakage audit for a model fine-tuned on licensed sensitive data: canaries, membership inference, extraction tests and pass/fail gates.
- Models Trained on Unlawful Personal Data: Buyer ExposureHow EDPB Opinion 28/2024 treats models trained on unlawfully processed personal data, what deployers must assess, and how buyers limit downstream exposure.
- PII Detection Datasets for Training and Evaluating RedactionHow to choose PII detection datasets for NER training and redaction evaluation: public benchmarks, synthetic PII limits, contamination and real text.
- PII Masking vs Surrogate Replacement in Training DataPlaceholder tags or realistic fake values? How PII redaction style shapes fine-tuned model behavior, coreference and residual-leak risk, with a spec.
- Pseudonymisation Techniques for AI Training DataHow to pseudonymise licensed training records with keyed hashing, vault tokens and consistent surrogates, and why hashed identifiers are not anonymous.
- Redacting PII in Scanned Documents for Document-AI TrainingHow to verify PII is removed from scanned document images, OCR text layers and PDF metadata without destroying the layout signal document-AI models need.
- Redacting Spoken PII from Call Recordings for AI TrainingA buyer's acceptance spec for call-recording redaction: transcript detection, word-level alignment, audio masking, card data and residual voice identity.
- Releasing Open Weights Trained on Personal DataWhen can you publish model weights trained on licensed personal data? White-box extraction risk, GDPR anonymity tests and a pre-release privacy gate.
- Safe Harbor Dates, Ages and ZIP Codes for ML ModelsWhich dates, ages and ZIP codes survive HIPAA Safe Harbor, what date shifting and relative time do to temporal models, and how to scope a data request.
- Speaker Anonymization for Speech Data: Reading EER and WERHow to evaluate voice-anonymized speech datasets: VoicePrivacy EER under informed attackers, WER and emotion UAR, pseudo-speakers and residual leakage.
- Special Category Data in AI Training Datasets: ScreeningHow to find and handle GDPR Article 9 and US sensitive data hiding in support tickets, emails and call audio before a licensed corpus is used for training.
- Video Anonymization for AI Training Data: Buyer SpecWhat operational video must have anonymized beyond faces: bystanders, screens, badges, plates, audio and metadata, and how to verify it across frames.
Sources
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (current text). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)". https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- Bird & Bird, "EU: The SRB decision - a new era for personal data and data processing agreements" (2025). https://www.twobirds.com/en/insights/2025/eu-the-srb-decision-a-new-era-for-personal-data-and-data-processing-agreements
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- Acompli, "Digital Omnibus GDPR and Cookie Reforms Stall Without a Council Mandate" (2026). https://acompli.ie/news/digital-omnibus-gdpr-cookies-status-september-2026/
- Information Commissioner's Office, "Anonymisation guidance: About this guidance" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/about-this-guidance/
- Information Commissioner's Office, "How do we ensure anonymisation is effective?" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
- Information Commissioner's Office, "Pseudonymisation" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
- U.S. Government Publishing Office / U.S. Department of Education, "34 CFR 99.31 - Under what conditions is prior consent not required to disclose information?" (CFR 2018 edition). https://www.govinfo.gov/content/pkg/CFR-2018-title34-vol1/pdf/CFR-2018-title34-vol1-sec99-31.pdf
- Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
- U.S. Department of Health and Human Services, Federal Register, "Confidentiality of Substance Use Disorder (SUD) Patient Records (Final Rule)" (2024). https://www.govinfo.gov/content/pkg/FR-2024-02-16/html/2024-02544.htm
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
- Washington State Legislature, "Chapter 19.373 RCW - Washington My Health My Data Act". https://app.leg.wa.gov/RCW/default.aspx?cite=19.373&full=true
- Latanya Sweeney, "k-Anonymity: A Model for Protecting Privacy" (2002). https://dataprivacylab.org/people/sweeney/kanonymity.html
- Rocher, Hendrickx and de Montjoye, "Estimating the success of re-identifications in incomplete datasets using generative models" (2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
- National Institute of Standards and Technology (Garfinkel), "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- Microsoft presidio project, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- arXiv:2109.00281, "Benchmarking and challenges in security and privacy for voice biometrics" (2021). https://arxiv.org/pdf/2109.00281
- arXiv:2512.16086, "Privacy Blur: Quantifying Privacy and Utility for Image Data Release" (2025). https://arxiv.org/pdf/2512.16086
- NEMA / DICOM Standards Committee, "DICOM PS3.15 Security and System Management Profiles, Annex E: Attribute Confidentiality Profiles" (current edition). https://dicom.nema.org/medical/dicom/current/output/chtml/part15/chapter_E.html
- Carlini et al., "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
- Nasr et al., "Scalable Extraction of Training Data from Aligned, Production Language Models" (2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/cce0e917b050208170151f77b497fc71-Abstract-Conference.html
- CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
- Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
- Hukkelås and Lindseth, "Does Image Anonymization Impact Computer Vision Training?" (2023). https://openaccess.thecvf.com/content/CVPR2023W/WAD/papers/Hukkelas_Does_Image_Anonymization_Impact_Computer_Vision_Training_CVPRW_2023_paper.pdf
- Stanford University Open Journal Systems (GRACE), "Article on de-identified datasets in foundation-model training (title to confirm)". https://ojs.stanford.edu/ojs/index.php/grace/article/download/3837/1799/11712
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.