Privacy, de-identification and sensitive data
Receiving pseudonymised data after EDPS v SRB: when it is not personal data for the buyer
Quick answer
Since the CJEU's judgment in C-413/23 P (4 September 2025), pseudonymised records can be personal data for the company that holds the key yet not personal data for a recipient that has no reasonable means to re-identify anyone [1]. For an AI data buyer that conclusion is earned, not assumed. You need evidence that you cannot reverse the pseudonymisation or identify people by other means, and the content itself must not identify or relate to individuals, which free text often does [1][3].
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What the Court actually decided in EDPS v SRB
The Court held that whether pseudonymised data is personal data can be assessed from the perspective of the recipient, not only the original controller [1]. The facts are useful for buyers because they look like a data license. The Single Resolution Board collected comments from shareholders and creditors, replaced names with alphanumeric codes, and passed the coded comments to Deloitte, which had no access to the key [2]. The EDPS treated the transfer as a disclosure of personal data. The General Court annulled that decision for not examining Deloitte's position; on appeal the Court of Justice accepted the recipient-relative approach in principle but set the General Court's judgment aside on the points below [2][6].
Three holdings matter for intake teams:
- Recipient-relative identifiability. Pseudonymised data is not automatically personal data for every party. It falls outside the definition for a recipient only if the pseudonymisation measures actually prevent that recipient from attributing the data to a person, considering means reasonably likely to be used [1][7].
- Opinions relate to their authors. A comment, review or opinion is information relating to the person who wrote it, so the content test is not avoided by coding the author's name [3].
- The disclosing controller keeps its duties. The SRB's obligation to inform data subjects about recipients was assessed at the time of collection, from the controller's perspective, regardless of what Deloitte could do [2].
The case arose under Regulation (EU) 2018/1725, which governs EU institutions, but commentators read it as applying equally to the GDPR because the definitions mirror each other [1][5]. Commentators report that the case was withdrawn in December 2025, after the judgment, and treat the judgment as the operative standard as of October 2026 [4].
Why pseudonymised still usually means personal data
The default remains that pseudonymised data is personal data; SRB created a fact-dependent route out for a specific recipient, not a new category [2][8]. GDPR Article 4(5) defines pseudonymisation as processing so data can no longer be attributed to a person without additional information kept separately under technical and organisational measures, and Recital 26 says such data should be treated as information on an identifiable person [7]. The ICO guidance takes the same line and warns that weak or reversible pseudonymisation leaves the data fully in scope [8].
The relative test does not lower the anonymisation bar; it changes whose means are counted. Recital 26 already asked about means reasonably likely to be used, "either by the controller or by another person" [7]. SRB confirms that for the recipient's own classification, the relevant means are those realistically available to the recipient, including auxiliary data it holds or can obtain [1][5]. For background on the general classification, see the de-identified data for AI training hub, anonymised or pseudonymised training data under GDPR and the Pseudonymization glossary entry.
The recipient test applied to AI training intake
A buyer can treat a delivery as non-personal only if it can show that both the key and the content are out of reach. Ask two separate questions, because failure on either keeps the data personal for you [1][3].
Question 1: can you reverse the measures? Consider who holds the mapping table, salt or HMAC secret; whether tokens are deterministic across deliveries (so a later delivery with partial identity could link back); whether the supplier would re-identify on request under a support or audit clause; and whether your group companies or vendors hold the same customers' data under their real identifiers.
Question 2: can you identify people by other means? This is the motivated intruder question the ICO describes: what could a determined person achieve with public sources, your existing data and reasonable effort [9]? In business text the risk usually sits in the content, not the ID column. Job titles in small firms, named locations, rare incidents, order values and timestamps can single out a person even after every name is replaced. The indirect identifiers in business text guide covers these patterns.
Training adds a third consideration: what the model will do. If a model can memorise and emit rare strings, outputs may expose content that was identifying in combination. EDPB Opinion 28/2024 assesses model anonymity case by case, separately from the status of the input data [11]. See when a model trained on personal data is anonymous.
Free text, transcripts and reviews: where the argument fails
Unstructured records are the weakest candidates for a recipient-side assessment because the content relates to its author or subject even after coding [3]. A support transcript in which an agent writes "as discussed with you last Tuesday about the Lyon warehouse lease" relates to a person through its content, and an employee's performance comment relates to both writer and subject. Coding the speaker field does nothing to that.
Practical consequences for intake:
- Structured event logs with tokenised account IDs and coarsened timestamps are more defensible than free text.
- Free-text fields need entity redaction or surrogate replacement on top of ID tokenisation, followed by a residual-risk review of a sample.
- Call recordings and voice data are rarely defensible as non-personal; voice itself is an identifier.
- Opinions, ratings and complaints are personal for their authors by content, per the Court's reasoning on comments [3].
If you cannot get free text below the identifiability threshold, plan for the data to be personal data in your hands and build the lawful basis instead; the legitimate interest assessment for licensed training data sets out what to document.
Evidence file: what a buyer should hold before relying on SRB
A recipient-side conclusion should rest on a written file you can show a supervisory authority, because GDPR accountability means the party claiming the data is outside scope should be able to show its reasoning. The judgment does not prescribe a list; the items below are a practitioner synthesis of what the Court's reasoning makes relevant [1][2][9].
Illustrative example: invented to show structure; it does not describe an available dataset.
| Evidence item | What it shows | Where it lives | Failure mode |
|---|---|---|---|
| Key custody statement | Mapping table, salt or HMAC secret stays with the supplier; no copy, escrow or API access for the buyer | Supplier preparation memo; data license schedule | Supplier offers "re-identification on request" for audit or deduplication |
| Tokenisation method | Keyed hash or random surrogate, per-delivery or per-buyer salt, no plain hashes of emails | Preparation record | Unsalted SHA-256 of emails is reversible by lookup |
| Contractual re-identification ban | Buyer may not attempt to identify individuals or link with other data; flow-down to vendors | License; vendor terms | Ban absent from subprocessors and annotation vendors |
| Technical separation | Data lands in an isolated project with no join paths to CRM, ad or customer tables | Access-control config; data catalog lineage | Training bucket shared with analytics workspace |
| Auxiliary data inventory | Buyer holds no data on the same population, or holds it under controls preventing linkage | Internal data map | Lab also licenses a CRM export from the same supplier |
| Content residual-risk review | Sample checked for names, roles, places and rare events in free text | QA report with sample size and findings | Redaction run only on structured columns |
| Re-assessment trigger | Who re-runs the test when new data, partners or model features arrive | Governance procedure | Classification never revisited after a second delivery |
An intake record can capture the conclusion in a form auditors can read:
Illustrative example: invented to show structure; it does not describe an available dataset.
intake_id: DEL-0007
source_description: "Support tickets, EU-connected B2B software vendor"
pseudonymisation:
method: "HMAC-SHA256 on account_id and agent_id; per-buyer key"
key_holder: "supplier only"
free_text_treatment: "NER redaction + surrogate names; 500-record manual review"
recipient_assessment:
can_reverse_measures: false
auxiliary_data_on_population: "none held"
contractual_reidentification_ban: true
technical_separation: "isolated project, no joins to customer systems"
residual_content_risk: "low after review; 3 rare-event tickets removed"
classification_for_recipient: "not personal data (SRB recipient test)"
reassess_on: ["new delivery", "new data source on same population", "model release"]
What SRB does not settle for buyers
The judgment narrows one question and leaves several open, so plan for the cases where your conclusion is challenged [1][4]. Bird & Bird notes that the reasoning raises questions about when processor agreements under Article 28 are needed if the recipient sees only non-personal data [1].
- The supplier's side. The disclosing controller still needs a lawful basis and must have told data subjects about recipients at collection [2]. If that notice was missing, the supplier's disclosure may be unlawful even if the data is non-personal for you, which is a contractual and reputational risk worth diligence.
- Transfers. Whether a disclosure to a non-EU recipient is a restricted transfer when the data is personal only for the exporter is not answered by the judgment; counsel should decide whether to put transfer terms in place anyway. The transfer mechanisms guide covers the options.
- Regulator guidance. EDPB Guidelines 01/2025 on pseudonymisation went through public consultation in 2025 and predate the judgment [10]; watch for updated guidance rather than relying on the consultation text.
- Legislation. The Digital Omnibus proposal would narrow the definition of personal data along similar lines, but as of October 2026 the GDPR amendments are proposals, not law [12].
- UK recipients. The ICO's anonymisation guidance applies its own effectiveness tests and is under review after the Data (Use and Access) Act 2025 [8][9].
How SourceX prepares pseudonymised records for buyers
SourceX sources operational datasets from US companies on request, and every dataset is rights-reviewed for ownership and consents before delivery under a license that defines records, uses, term and delivery. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded, and a sample is checked, though no method is perfect. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which gives your counsel the inputs for a recipient-side assessment rather than the conclusion itself. You can describe the records you need on the SourceX buyer page; for the supply side of the same question, see GDPR and selling data to AI companies.
Licensing pseudonymised operational records for training and evaluation
SourceX finds US businesses that hold the data you describe, and each release is approved by the supplying company, with nothing contracted until a supplier agrees. Delivery runs through private, access-controlled workflows after an executed agreement. Describe your dataset and intended use at sourcex.si/buyers.
Frequently asked questions
Does EDPS v SRB mean GDPR never applies to data I cannot re-identify?
No. It means identifiability is assessed for your position, using the means reasonably likely available to you [1][7]. The content can still identify or relate to people, as opinions and comments do, and the disclosing controller keeps its own obligations [2][3].
Is hashing email addresses enough to rely on the recipient test?
Usually not. An unsalted hash of an email can be reversed by hashing candidate addresses, so a recipient with access to email lists can re-identify; ICO guidance treats weak pseudonymisation as personal data [8]. Keyed hashes with supplier-held secrets are a stronger starting point.
Does the withdrawal of the case in December 2025 weaken the judgment?
Commentators report the case was withdrawn in December 2025 but treat the Court's judgment as the operative standard on pseudonymised data [4]. Check later case law and EDPB guidance before relying on it in a new matter.
Sources
- Bird & Bird, "EU: The SRB decision: a new era for personal data and data processing agreements" (2025). https://www.twobirds.com/en/insights/2025/eu-the-srb-decision-a-new-era-for-personal-data-and-data-processing-agreements
- Jones Day, "CJEU Clarifies Scope of Personal Data in EDPS v SRB Decision" (2025). https://www.jonesday.com/en/insights/2025/09/cjeu-clarifies-scope-of-personal-data-in-edps-v-srb-decision
- Burges Salmon, "Pseudonymised Data: CJEU provides clarification on the concept of Personal Data" (2025). https://www.burges-salmon.com/articles/102l5dc/pseudonymised-data-cjeu-provides-clarification-on-the-concept-of-personal-data/
- Lewis Silkin, "It's nothing personal: Reassessing pseudonymised data and AI after EDPS v SRB" (2025). https://www.lewissilkin.com/insights/2025/12/18/its-nothing-personal-reassessing-pseudonymised-data-and-ai-after-edps-v-srb-102ly5u
- Future of Privacy Forum, "Rethinking Personal Data: The CJEU's Contextual Turn in EDPS vs. SRB" (2025). https://fpf.org/blog/rethinking-personal-data-the-cjeus-contextual-turn-in-edps-vs-srb/
- Clifford Chance, "Pseudonymized data after EDPS v SRB" (2025). https://www.cliffordchance.com/insights/resources/blogs/talking-tech/en/articles/2025/09/pseudonymized-data-after-edps-v-srb.html
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- Information Commissioner's Office, "Pseudonymisation" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
- Information Commissioner's Office, "How do we ensure anonymisation is effective?" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/
- European Data Protection Board, "Guidelines 01/2025 on Pseudonymisation" (2025). https://www.edpb.europa.eu/our-work-tools/documents/public-consultations/2025/guidelines-012025-pseudonymisation_en?page=4
- CMS, "EDPB Opinion 28/2024: key takeaways on processing personal data in the context of AI models" (2024). https://cms.law/en/int/legal-updates/edpb-opinion-28-2024-key-takeaways-on-processing-personal-data-in-the-context-of-ai-models
- Acompli, "Digital Omnibus GDPR and Cookie Reforms Stall Without a Council Mandate" (2026). https://acompli.ie/news/digital-omnibus-gdpr-cookies-status-september-2026/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.