Privacy, de-identification and sensitive data
Data minimization for training datasets: deciding which fields never leave the supplier
Quick answer
Data minimization for AI training means your data specification names every field, time window and level of detail the model actually needs, and treats everything else as staying with the supplier. Write the spec from the training objective backward: list the target label, the input features, the join keys and the history depth, then decide per field whether to keep, transform, generalize or drop it at source. A tighter spec reduces de-identification work, shortens legal review and strengthens any later anonymity claim.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why the buyer's specification, not the supplier's redaction, sets the privacy baseline
The buyer's specification sets the privacy baseline because a supplier can only minimize against a stated purpose, and the purpose lives in your spec. GDPR Article 5(1)(c) requires personal data to be adequate, relevant and limited to what is necessary for the purposes of processing [1]. HIPAA's minimum necessary standard, with implementation specifications in 45 CFR 164.514(d), makes the same move for protected health information [5]. If your request says "all ticket data from 2019 onward," the supplier has nothing to minimize against.
Minimization also does work downstream. The EDPB's Opinion 28/2024 says whether a trained model is anonymous must be assessed case by case [2], and summaries of the opinion list data preparation and minimization at the selection and training stage among the measures a controller can point to [3]. A narrow input set is evidence you can show later; a broad one is a liability you have to explain. For the wider de-identification picture, start at the privacy and de-identification hub.
Start from the training objective and work backward to fields
Every requested field should trace to one of four roles: label, feature, join key or evaluation slice. If a column fits none of them, it does not belong in the spec. This forces the conversation that usually gets skipped: what is the model predicting or generating, and what evidence does it need to do so?
Take a support-resolution model trained on a helpdesk export from a system such as Zendesk or Salesforce Service Cloud. It needs ticket text, the resolution category, timestamps to compute time-to-resolution, and perhaps product or plan tier. It does not need requester name, email, phone, billing account number, IP address or the agent's employee ID. Health-data practice guidance makes the same recommendation in general terms: restrict the fields, the time ranges and who can access them [4].
Two traps recur. First, "might be useful later" fields such as free-text notes, attachments and custom fields are where most unplanned personal data hides; see special category data hiding in operational records. Second, fields that look harmless can proxy for protected attributes, so check them against proxy variables in structured business data before keeping them as features.
Five field dispositions to assign in the spec
Each field in a minimized spec gets exactly one disposition, and the disposition determines who does the work and when.
- Keep as is. The value is needed at native precision and is not identifying in context, for example a resolution code or product SKU.
- Transform. The signal is needed but the raw value is not: replace a customer ID with a keyed, salted pseudonym that preserves joins, or convert absolute timestamps to offsets from ticket open.
- Generalize. Reduce precision: birth date to year, five-digit ZIP to three digits or state, exact revenue to a band. HIPAA Safe Harbor itself works this way for dates and geography [5].
- Drop at source. The field never leaves the supplier's environment. Direct identifiers and anything outside the four roles belong here.
- Scrub inside free text. Structured fields can be dropped cleanly; names and account numbers inside ticket bodies, email threads and call transcripts need entity detection and replacement. Measure what that step misses using the methods in PII redaction for LLM training data.
NIST SP 800-188 describes suppression, generalization and pseudonymization as standard de-identification techniques and pairs them with governance over who decides [6]. Writing the disposition into the spec makes that decision explicit and reviewable.
A need-to-train specification template
The template below is what a minimized request should look like before it reaches a supplier: one row per field, a role, a disposition and a reason.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Source field (helpdesk export) | Role | Disposition | Delivered form | Reason |
|---|---|---|---|---|
ticket_id | Join key | Transform | HMAC-SHA256 pseudonym, supplier-held key | Joins ticket to comments without exposing source IDs |
requester_name, requester_email, phone | None | Drop at source | Not delivered | No training role |
account_number, billing_id | None | Drop at source | Not delivered | No training role; financial identifier |
org_name | Evaluation slice | Generalize | Industry code and employee-count band | Slice by segment without naming customers |
created_at, solved_at | Feature, label input | Transform | Month of creation; duration in minutes | Keeps seasonality and time-to-resolve, drops exact time |
description, comments[].body | Feature | Scrub inside free text | Typed placeholders such as [PERSON], [EMAIL] | Model needs the language, not the identities |
assignee_id | None | Drop at source | Not delivered | Employee data with no training role |
resolution_category | Label | Keep | Native value | Target variable |
custom_fields.* | Unknown | Drop unless justified | Listed individually if requested | Custom fields hide unplanned personal data |
attachments | None | Drop at source | Not delivered | Screenshots and PDFs carry unreviewed content |
| History window | Scope | Restrict | 24 months, closed tickets only | Enough cycles for seasonality; no open cases |
Add three lines under the table: the stated purpose (for example, "train and evaluate a ticket-routing classifier"), the population filter (for example, "exclude tickets tagged legal, HR or security incident"), and the sample size you need to validate the dispositions before full delivery.
Time windows, granularity and row-level scope are minimization too
Minimization applies to rows and history as much as columns, and these levers are often more powerful than dropping a field. A 24-month window of closed tickets exposes fewer people than ten years of history, and if the model's accuracy plateaus at 18 months the extra years add risk without signal. Ask your team for a learning-curve estimate before setting the window.
Granularity matters because re-identification usually comes from combinations of quasi-identifiers, not from any single field. NISTIR 8053 surveys cases where de-identified records were re-identified by linking dates, locations and demographics to outside data [7]. A spec that keeps exact date, five-digit ZIP and job title together has rebuilt an identifier even after every name is gone.
Row-level filters belong in the spec too. Exclude record types that carry high-risk content and add little to the task: legal holds, HR complaints, fraud investigations, minors' accounts and anything flagged under a regulated regime. If a regime does apply, the spec has to match it: HIPAA Safe Harbor lists 18 identifier categories to remove, and the limited data set route keeps some dates and geography but requires a data use agreement [5].
How minimization changes the deal and the delivery
A minimized spec usually makes a supplier's approval easier, because the supplier's own privacy and security reviewers are checking the same question you are: what leaves the building and why. Fewer fields means fewer identifiers to remove or replace, a smaller residual-risk analysis, and a shorter list of confidential business fields to argue over. For how suppliers keep trade secrets and customer terms out, see keeping confidential information out of a data license.
Delivery format can enforce the spec. Columnar formats such as Apache Parquet store each column in separate chunks [9], so a supplier can write only the approved columns rather than exporting a full table and deleting later. Share protocols such as Delta Sharing let a provider share specific tables with a recipient over REST APIs [10], so the supplier can share only a minimized table rather than the source system. When even a minimized copy is too sensitive to move, consider training inside a data clean room instead of receiving records.
If you are building a high-risk system in scope of the EU AI Act, Article 10 requires training, validation and testing data to be subject to data governance and management practices [8]. As of October 2026, those high-risk obligations reportedly start on 2 December 2027 for Annex III systems after Regulation (EU) 2026/1744. A field-level disposition table is a direct, auditable record of that governance.
Failure modes to check before you sign off on the spec
Most minimization failures come from fields that look structural but carry identity, so review the spec against this list before sending it.
- Stable pseudonyms that link across deliveries. An unsalted hash of email is reversible by dictionary attack; ask for a keyed hash with the key held by the supplier, and decide whether refresh deliveries share the key.
- Free text overriding column drops. Dropping
requester_emaildoes nothing if signatures incomments[].bodyrepeat it. - Metadata leakage. File names, attachment URLs, Parquet key-value metadata and JSON field names can carry customer or employee names.
- Rare categories. A single customer in an industry code and size band is identifiable; set a minimum cell size or merge small groups.
- Memorization of what survives. Whatever you keep can be reproduced by a large model; weigh that using training-data extraction and memorization risk.
- Evidence gap. Ask the supplier to document the method applied to each field; the de-identification evidence package checklist lists what to request.
Validate dispositions on a sample before full delivery. If a supplier cannot release raw rows for that check, the options in sample access for sensitive data show how to review without receiving them.
Where SourceX fits in a minimized request
SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents and finance and legal workflows, and manages the licensing process. Buyers describe the data they need rather than naming businesses, and every release is approved by the supplying company. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. A field-level spec like the one above is the most useful way to describe your data need to SourceX. Browse other buyer guides from the AI data hub.
Request a minimized operational dataset
Bring your field-disposition table, history window and stated purpose. SourceX assesses data and licensing permissions with suppliers, and nothing is contracted until a supplier agrees, with the dataset delivered under a license that defines records, uses, term and delivery. Describe the minimized dataset you need.
Sources
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- European Data Protection Board, "Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models" (2024). https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- Securiti, "Summary of EDPB Opinion 28/2024 Concerning AI Models' Processing of Personal Data". https://securiti.ai/summary-of-edpb-opinion-282024-concerning-ai-models-processing-of-personal-data
- AccountableHQ, "HIPAA and Machine Learning: What You Need to Know to Build Compliant Healthcare AI". https://www.accountablehq.com/post/hipaa-and-machine-learning-what-you-need-to-know-to-build-compliant-healthcare-ai
- Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- The Apache Software Foundation (Apache Parquet), "File Format". https://parquet.apache.org/docs/file-format/
- Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.