Skip to content

Privacy, de-identification and sensitive data

Special category and sensitive data hiding in operational training data

Quick answer

Operational corpora such as support tickets, sales emails and call recordings almost always contain special category data that nobody collected on purpose: a customer explaining a cancellation by naming a diagnosis, an agent noting a religious holiday, a union grievance in an HR thread. Before training, buyers should screen against GDPR Article 9 and US "sensitive data" lists together, then decide per category whether to exclude records, redact spans, or document a lawful condition.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why incidental sensitive data is the default in business records

Sensitive data appears in operational records because customers and employees explain themselves in free text, and nothing in a CRM or ticketing schema stops them. A Zendesk "cancellation reason" field, a Salesforce activity note or a Gong call transcript is designed to capture context, and context often includes health, family, faith or finances. The field schema looks clean; the payload does not.

The legal significance is that GDPR Article 9(1) prohibits processing data "revealing" a special category unless an Article 9(2) condition applies, and that prohibition does not depend on whether the controller intended to collect it [1]. US state laws also regulate "sensitive data" or "sensitive personal information," though the trigger differs: California's right to limit, for example, does not reach sensitive personal information collected or processed without the purpose of inferring characteristics about a consumer, so the intended use of the corpus matters [4]. For a buyer, that means a corpus licensed as "support conversations" may also be a health, biometric or children's data corpus in part.

For the general identifier pipeline, see PII redaction for LLM training data and the PII detection glossary entry. This page focuses on the narrower categories that carry heightened rules.

One screening taxonomy for GDPR Article 9 and US sensitive data

A single screening taxonomy works better than separate EU and US passes, because the category lists overlap heavily and the detectors are the same. GDPR Article 9(1) covers racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, genetic data, biometric data processed to uniquely identify a person, health data, and data concerning sex life or sexual orientation [1]. Criminal conviction data sits separately under Article 10 GDPR but should be screened alongside.

California's definition of sensitive personal information adds categories the EU list lacks, including government identifiers, account log-ins and financial account numbers with credentials, precise geolocation, citizenship or immigration status (added by AB 947 [5]), and the contents of a consumer's mail, email and text messages unless the business is the intended recipient [4]. Virginia's Consumer Data Protection Act (Va. Code 59.1-578, 59.1-580) [12] and many later state comprehensive laws go further than California's opt-out model: they define "sensitive data" (health diagnosis, religious beliefs, sexual orientation, biometric and genetic data, precise geolocation, known-child data) and require opt-in consent and a data protection assessment before processing it. Washington's My Health My Data Act defines "consumer health data" broadly enough to reach inferences, such as a purchase or a question that suggests a health condition [6].

Illustrative example: invented to show structure; it does not describe an available dataset.

Screening tagGDPR Art. 9 / 10US hooks to checkTypical hiding place in operational dataDefault handling
HEALTHHealth dataHIPAA (if PHI), WA MHMDA, VA/CA sensitive dataCancellation reasons, refund requests, accommodation emails, insurer callsRedact span; exclude if the record is about the condition
BIOMETRICBiometric for unique identificationIllinois BIPA, COPPA (children), state sensitive dataRaw call audio, video, voice-verification logs, signature imagesKeep audio only with documented basis; never extract speaker templates
RELIGION_BELIEFReligious or philosophical beliefsCA SPI, VA sensitive dataScheduling and leave requests, dietary notes, holiday referencesRedact span
UNIONTrade union membershipCA SPIHR tickets, grievance threads, payroll deductionsExclude HR queues by default
SEX_ORIENTATIONSex life or sexual orientationCA SPI, VA sensitive dataBenefits enrollment, partner names in context, dating or wellness productsRedact span or exclude
ETHNIC_POLITICALRacial or ethnic origin; political opinionsCA SPI, VA sensitive dataLanguage-preference notes, donation records, complaint narrativesRedact span
GENETICGenetic dataCA SPI, WA MHMDALab and testing product supportExclude record
CHILDNot a special category; heightened dutiesCOPPA, VA known-child dataParent-account tickets, education products, background voicesExclude record
CRIMINALArticle 10 dataState law variesFraud and collections notes, background checksExclude or redact
MESSAGE_CONTENTOrdinary personal dataCA SPI (contents of communications)Email and SMS bodiesConfirm basis for using message bodies at all

Where special category data hides in tickets, email and call audio

The highest-yield places to look are free-text fields, attachments and audio, not structured columns. Structured fields like plan_tier or region rarely carry Article 9 data; description, internal_note, close_reason and transcript turns do.

Common hiding places a screening plan should name explicitly:

  • Free-text reason codes. "Other: please specify" fields capture diagnoses, bereavement, pregnancy and disability accommodations.
  • Agent internal notes. Agents write "customer mentioned chemo, waived fee," which is health data about a named account even after the customer's name is masked.
  • Attachments. Doctor's notes, insurance letters and ID scans attached to tickets or emails. OCR them or exclude them; do not ship them unscanned.
  • Call audio. Recorded voice of an identifiable speaker is personal data under GDPR and most US state privacy laws. Under Illinois BIPA a voiceprint is a biometric identifier [7], so a pipeline that computes speaker embeddings for diarization or verification can create biometric data that the source recordings did not. See redacting spoken PII from call recordings.
  • Email signatures and footers. Union local numbers, pronouns, faith-based organization names and medical practice signatures.
  • Background speech and children. Children audible on calls or named in parent-account tickets. The amended COPPA Rule took effect June 23, 2025, with a compliance date of April 22, 2026 [8], and added biometric identifiers to its definition of personal information [9].

Health data in customer support transcripts deserves its own pass. If the supplier is a HIPAA covered entity or business associate, the transcripts may be PHI and need Safe Harbor or Expert Determination before they leave the supplier [10]; the trade-offs are covered in HIPAA Safe Harbor vs Expert Determination. If the supplier is not covered by HIPAA, Washington's MHMDA and state sensitive data rules may still apply [6].

Detection methods and their known failure modes

Detection for special categories needs a classifier layer on top of entity recognition, because the sensitive signal is usually a statement, not an entity. Presidio-style recognizers and NER models find names, phone numbers and card numbers well, but "I can't come in, it's Ramadan" contains no entity a standard recognizer tags.

A workable stack has three layers:

  1. Lexicon and pattern pass. Curated term lists per tag (condition names, medications, denominations, union names), with ICD-10 and RxNorm vocabularies for health. High recall on explicit mentions; misses paraphrase and misspelling.
  2. Sentence-level classifier. A fine-tuned encoder or LLM prompt that labels each turn or sentence with the screening tags above. Catches "my treatment" and "since my diagnosis," but drifts on domain jargon (a "plan" in insurance support is not a health signal by itself).
  3. Human review of a stratified sample. Sample by tag, queue and channel, and measure missed-detection rate per tag. Report recall per category, not one blended PII score.

Failure modes worth asking a supplier about directly:

  • Masking the name but leaving the condition. Pseudonymized records with intact health narratives are still personal data under GDPR, since only data rendered anonymous falls outside it [1].
  • Inference through combination. A pharmacy SKU plus a refund note can reveal a condition no single field states, which is the kind of data Washington's definition reaches [6].
  • Transcript-only redaction. Text is redacted but the aligned audio still speaks the diagnosis.
  • Near-duplicates. Canned replies quoting the customer's message replicate the sensitive span across threads; deduplicate before review.

Residual sensitive data also matters downstream because models can memorize rare strings; see training-data extraction and memorization risk.

Three handling options: exclude, redact, or document a lawful condition

Each screening tag needs a written default of exclude, redact or keep-with-basis, decided before the corpus is assembled. The EDPB's Opinion 28/2024 discusses data selection, preparation and minimisation at the design stage as measures relevant to assessing AI model development [3], and it addresses when a model can be considered anonymous and the consequences of unlawful processing during development [2].

Exclude the record when the record exists because of the sensitive fact: an accommodation request, a medical refund, an HR grievance. Redacting the condition leaves an incoherent example that harms training more than it helps.

Redact the span when the sensitive fact is incidental to an otherwise useful interaction. Use typed placeholders ([HEALTH_CONDITION], [RELIGIOUS_REFERENCE]) rather than deletion, so the model learns the dialog structure without the fact, and record the method and the sample-check results.

Keep with a documented basis only when the use case requires it. Under GDPR that means an Article 9(2) condition in addition to an Article 6 basis, such as explicit consent or the scientific research condition with Article 89 safeguards [1]. Under Virginia-style state laws it means opt-in consumer consent and a data protection assessment, which incidental mentions in a ticket archive almost never carry. For high-risk AI systems, Article 10(5) of the AI Act permits special category data for bias detection and correction under strict conditions [11]; as of October 2026, Regulation (EU) 2026/1744 has deferred high-risk obligations to 2 December 2027 (Annex III) and 2 August 2028 (Annex I), and deleted Article 10(5). See EU AI Act Article 10.

What to request from a supplier before licensing

Ask for the screening evidence as a document, not a verbal assurance. The minimum package for an operational text or audio corpus:

Illustrative example: invented to show structure; it does not describe an available dataset.

sensitive_data_screening:
  taxonomy_version: "screen-tags v3 (GDPR Art. 9/10 + CA SPI + VA sensitive data + MHMDA)"
  sources_in_scope: [zendesk_tickets, outbound_email, call_audio_wav_8khz]
  excluded_queues: [hr_cases, accessibility_requests, medical_billing]
  detectors:
    - layer: lexicon
      vocabularies: [ICD-10-CM terms, RxNorm ingredients, custom denominations list]
    - layer: classifier
      granularity: sentence
  handling_defaults:
    HEALTH: redact_span
    BIOMETRIC: no_speaker_embeddings_generated
    CHILD: exclude_record
    UNION: exclude_record
  placeholders: typed
  audio_alignment: redacted spans muted in audio via word timestamps
  sample_review:
    stratified_by: [tag, channel]
    reviewed_records: <number>
    missed_detections_by_tag: {HEALTH: <n>, RELIGION_BELIEF: <n>}
  residual_risk_statement: "no method is perfect; known gaps listed"

Pair this with the broader de-identification evidence package checklist, and check whether the supplier's own customer contracts and DPAs permit training use at all through customer contracts and DPAs. For jurisdiction overviews, see SourceX's privacy laws index.

How SourceX approaches sensitive data in sourced corpora

SourceX sources operational datasets, including support and sales histories and documents, from US companies on request, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Health records require HIPAA de-identification by Safe Harbor or Expert Determination. Buyers can describe the corpus and the categories they need excluded, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset.

Example dataset pages for context: customer support ticket datasets and contact center call recordings. More guides are in the privacy and de-identification hub and the AI data hub.

Sourcing support or call data with sensitive categories screened

Describe the records you need and the categories your counsel wants excluded or redacted; SourceX looks for US businesses that hold that data and assesses the data and licensing permissions before anything is agreed. Requests do not guarantee a match, and nothing is contracted until a supplier agrees. Start a buyer request at sourcex.si/buyers.

Sources

  1. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
  2. CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
  3. Securiti, "Summary of EDPB Opinion 28/2024 concerning AI models' processing of personal data" (2025). https://securiti.ai/summary-of-edpb-opinion-282024-concerning-ai-models-processing-of-personal-data
  4. California Legislative Information, "Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  5. California Legislative Information, "AB-947 California Consumer Privacy Act of 2018: sensitive personal information" (2023). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB947
  6. Washington State Legislature, "Chapter 19.373 RCW - Washington My Health My Data Act". https://app.leg.wa.gov/RCW/default.aspx?cite=19.373&full=true
  7. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  8. Federal Trade Commission, Federal Register, "Children's Online Privacy Protection Rule (Final Rule amendments), 90 FR 16918" (2025). https://www.federalregister.gov/documents/2025/04/22/2025-05904/childrens-online-privacy-protection-rule
  9. Finnegan, "COPPA's Amended Rule Is Now in Full Effect: What Operators Need to Know" (2026). https://www.finnegan.com/en/insights/articles/coppas-amended-rule-is-now-in-full-effect-what-operators-need-to-know.html
  10. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  11. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  12. Virginia Law Library, "Code of Virginia Chapter 53: Consumer Data Protection Act" (2026). https://law.lis.virginia.gov/vacode/title59.1/chapter53/
  13. California Legislative Information, "SB-1223 Consumer privacy: sensitive personal information: neural data" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240SB1223

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data