Privacy, de-identification and sensitive data
Legitimate interest for training on licensed personal data: the assessment buyers must document
Quick answer
Yes, an AI developer can rely on GDPR Article 6(1)(f) legitimate interests to train on licensed records that still contain personal data, but only after a documented three-step assessment: a specific, lawful interest; processing that is necessary for it; and a balancing test the data subjects do not win. The EDPB accepted in Opinion 28/2024 that this route can be available, assessed case by case [1]. For licensed data, the assessment must cover how the supplier collected the records, what the license allows, and which mitigations you control.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
When the licensed dataset still needs a lawful basis
A lawful basis is required whenever the records you receive are still personal data in your hands, which covers most pseudonymised transcripts, tickets and email threads. Recital 26 of the GDPR takes only anonymous information out of scope, judged against all means reasonably likely to be used to identify someone [4]. Pseudonymised data, where names are replaced with tokens but the mapping exists somewhere, is treated as personal data in the EDPB's Guidelines 01/2025 [5].
The practical test for a buyer is your own position, not the supplier's label. If you can link a customer_ref token across tickets, recover a signature block from an email body, or match a rare job title plus city to a LinkedIn profile, you hold personal data. Our pages on classifying licensed data as anonymised or pseudonymised and the recipient perspective after EDPS v SRB cover that classification; this page assumes the answer is "still personal data."
Do not assume the AI Act changes this. Article 2(7) keeps Union data protection law applicable to personal data processed in connection with the Act [11]. As of October 2026, the Digital Omnibus proposals that would narrow the personal-data definition remain proposals, not law [9], so build the assessment on the current GDPR text.
Why legitimate interest is usually the only workable basis for training
Legitimate interest is usually the realistic basis because the alternatives rarely fit records collected by a third party. Consent under Article 6(1)(a) was not gathered by the supplier for your model, and retrofitting it across thousands of support customers is not feasible. Contract under Article 6(1)(b) covers the supplier's service to its customer, not your training run [4].
The EDPB's answer in Opinion 28/2024, requested by the Irish supervisory authority, is that legitimate interest can be an appropriate basis for both development and deployment of AI models, assessed case by case [1][8]. Law-firm summaries read it the same way: the route is open, but the three-step test does real work [7]. Practitioner summaries describe the same three-step structure and stress that naming an interest is only the first hurdle [3]. For the concept itself, see our glossary entries on lawful basis and legitimate interest.
Step 1: write the interest as a specific, lawful purpose
The interest must be real, present and clearly articulated, not "developing AI." The EDPB gives examples such as building a conversational agent to assist users or improving threat detection in an information system [1][2]. A defensible statement for licensed operational data names the model, the task and the reason this data is needed.
Weak: "to train our models on business data." Stronger: "to fine-tune a support-resolution model that drafts first replies for B2B software tickets, using historical ticket threads to learn product-specific troubleshooting steps." The stronger version also scopes what the data may not be used for, which you will need in Step 3.
Step 2: prove necessity against less intrusive options
Necessity means the processing must be needed for the stated purpose and no less intrusive means will achieve it [3]. The EDPB ties this to data minimisation: the volume of personal data and whether the purpose could be met with less, or with none [1]. A DPO should be able to show reviewers the alternatives that were tested and rejected.
Questions that make the necessity record concrete:
- Could the model learn the task from fully de-identified text? If yes, the residual identifiers are not necessary; remove them. See re-identification risk assessment methods.
- Which fields does training actually read? Drop
email_from,phone,account_numberand free-text signature blocks if the label is onlyresolution_category. - Is the time window larger than needed? Five years of tickets when eighteen months cover the current product version is hard to justify.
- Could synthetic text do the job? If you rely on it, read whether synthetic data derived from records is still personal data.
Step 3: run the balancing test on the people in operational records
The balancing test asks whether the data subjects' interests, rights and freedoms override yours, and reasonable expectations carry much of the weight [1][3]. The EDPB lists factors such as the relationship with the controller, the context and source of collection, the nature of the service, whether the data was public, and whether people could know their data might be used this way [1]. In licensed operational data, the people are customers, employees and counterparties of a company you never dealt with.
Expectations differ by record type:
- Support tickets and chat logs. Customers expected their messages to resolve an issue, perhaps to train the vendor's own support staff, not a third party's model. Weighs against you unless identifiers are stripped.
- Employee email and meeting transcripts. Employment is an imbalanced relationship and messages contain opinions, health absences and performance remarks. See employee communications in training data.
- Finance and legal workflow records. Account numbers and dispute narratives raise impact; US GLBA obligations may run in parallel. See financial records and GLBA.
Any special-category content, such as health, union membership or religion visible in free text, needs an Article 9(2) condition in addition to Article 6 [4]; legitimate interest alone cannot carry it. The usual fix is detection and removal before delivery, and you should test that removal on a sample yourself.
Mitigations a licensed-data buyer controls
Mitigating measures can tip a close balancing test, and the EDPB treats them as part of the assessment, not an afterthought [1][2]. Pseudonymisation is the most cited: commentary on the EDPB's 2025 pseudonymisation guidance notes it can reduce risk and make legitimate interests easier to rely on, while the data stays personal [5][6]. As a buyer, you control several measures directly:
- Pre-delivery minimisation. Specify in the request which fields are excluded and which identifiers are replaced, and ask for the method record.
- Contractual re-identification ban. Prohibit linking tokens to external sources, and restrict access to the raw training corpus to named roles.
- Training-time controls. Deduplicate repeated strings, which reduces verbatim memorization of rare identifiers, and filter residual PII with a named detector before tokenization.
- Output-side controls. Run extraction tests for canary strings and known identifiers, and add output filters for phone, email and account-number patterns.
- An objection path. Article 21 gives people a right to object to legitimate-interest processing [4]; the EDPB also points to opt-out mechanisms going beyond the legal minimum as a mitigation [1]. Agree with the supplier how a suppression request reaches your corpus and your next training run.
- Transparency. Article 14 applies when data is not obtained from the data subject [4]. Decide with counsel whether the supplier's notices, your own public notice, or an exemption covers it, and record the reasoning.
Illustrative legitimate interests assessment record
The record below is the artifact reviewers ask for: one row per decision, each with evidence. Keep it versioned alongside the dataset manifest and the license.
Illustrative example: invented to show structure; it does not describe an available dataset.
| LIA field | Example entry | Evidence to attach |
|---|---|---|
| Dataset | 410,000 pseudonymised B2B support tickets, 2023-2025, English | Supplier manifest, field list |
| Interest (Step 1) | Fine-tune a reply-drafting model for software support | Product spec, model card draft |
| Necessity (Step 2) | Ticket bodies needed; requester name, email, phone, account ID removed; 24-month window | Field-drop log, ablation on de-identified sample |
| Data subjects | Customer end users; supplier support agents | Supplier description of collection context |
| Reasonable expectations | Supplier privacy notice mentions service improvement, not third-party AI | Archived notice version and date |
| Special categories | Health terms detected in 0.3% of tickets; removed pre-delivery | Detector name, sample QA report |
| Mitigations | Re-id ban in license; restricted bucket; dedup; output PII filter; suppression list sync | License clause reference, access policy, eval results |
| Objection handling | Supplier forwards suppression IDs; excluded at next refresh | Process note, contact owner |
| Transparency | Public training-data notice published; Art. 14 reasoning recorded | Notice URL, counsel memo |
| Outcome and review date | Balance favors processing with mitigations; review at next dataset refresh | DPO sign-off |
What happens if the supplier's collection was unlawful
Opinion 28/2024 also addresses the consequences of unlawful processing during development for later use of the model, and supervisory authorities assess this case by case [1][8]. A buyer cannot fully outsource that risk to a license. Ask for the supplier's collection context, the privacy notice in force when records were created, and how consent and notice records were kept.
If problems surface after delivery, follow a structured response such as our playbook for personal data found in a licensed dataset. Document what you knew at contract time; that record is part of your accountability under the GDPR.
Cross-border and UK variations
A buyer established in the EU applies the GDPR to its processing even when the records came from US companies and describe US residents, and EU or UK data subjects inside a US dataset raise transfer questions covered in buying US-sourced data with EU or UK personal data. In the UK, the Data (Use and Access) Act 2025 adds a statutory definition of scientific research that can include commercial research, and a list of recognised legitimate interests [10]. Neither removes the need for an assessment for ordinary model training, so treat UK reliance as a separate counsel question. Our privacy cluster hub and the question can I license data if I am under GDPR cover the wider frame.
How SourceX supports the documentation
SourceX sources operational datasets from US companies on request and manages licensing and ongoing purchases; nothing is held in stock and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Diligence materials on source, rights, preparation and allowed use are prepared per dataset, which feeds directly into the evidence column above. You can describe the data your assessment depends on before you finalize the LIA.
Request licensed training data you can document
If your legitimate interests assessment needs licensed operational records with a recorded de-identification method and per-dataset diligence materials, describe the data and the allowed uses you need. SourceX looks for US businesses that hold it, and every release is approved by the supplying company. Start a buyer request at sourcex.si/buyers.
Sources
- European Data Protection Board, "Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models" (2024). https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
- European Data Protection Board, "EDPB opinion on AI models: GDPR principles support responsible AI" (2024). https://www.edpb.europa.eu/news/edpb-opinion-on-ai-models-gdpr-principles-support-responsible-ai_en
- Baker McKenzie (Connect On Tech), "EDPB opinion on the processing of personal data in the context of AI models" (2024). https://connectontech.bakermckenzie.com/edpb-opinion-on-the-processing-of-personal-data-in-the-context-of-ai-models/
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- European Data Protection Board, "Guidelines 01/2025 on Pseudonymisation" (2025). https://www.edpb.europa.eu/our-work-tools/documents/public-consultations/2025/guidelines-012025-pseudonymisation_en?page=4
- VPH Institute, "The European Data Protection Board (EDPB) adopts pseudonymisation guidelines" (2025). https://www.vph-institute.org/news/the-european-data-protection-board-edpb-adopts-pseudonymisation-guidelines.html
- Osborne Clarke, "EDPB delivers view on using personal data to train and deploy AI models" (2024). https://www.osborneclarke.com/insights/edpb-delivers-view-using-personal-data-train-and-deploy-ai-models
- CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2025). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
- Acompli, "Digital Omnibus GDPR and Cookie Reforms Stall Without a Council Mandate" (2026). https://acompli.ie/news/digital-omnibus-gdpr-cookies-status-september-2026/
- UK Government (GOV.UK), "Data (Use and Access) Act 2025: data protection and privacy changes" (2025). https://www.gov.uk/guidance/data-use-and-access-act-2025-data-protection-and-privacy-changes
- European Parliament and Council of the European Union (EUR-Lex), "Regulation (EU) 2024/1689 (AI Act), Article 2: Scope" (2024). https://eur-lex.europa.eu/eli/reg/2024/1689/art_2/oj
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.