Skip to content

Privacy, de-identification and sensitive data

Can you release model weights trained on licensed personal data?

Quick answer

Sometimes, but only after you can show that personal data cannot reasonably be pulled out of the weights. In the EU, a model trained on personal data is treated as anonymous only if extracting that data is insignificant using all means reasonably likely to be used [1]. Publishing weights makes white-box attacks possible, so that bar is harder to meet than for an API, and your data license must also allow distribution of the trained model. Treat release as a gated decision backed by evidence, not as a default.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why open weights change the privacy analysis

Releasing weights removes every control you had at deployment time, so the privacy case must rest on the model itself. Behind an API you can rate-limit queries, filter outputs for identifiers, log abuse and patch the model. Once a checkpoint is on a public hub, anyone can run unlimited queries offline, read exact token probabilities, strip your safety tuning and fine-tune the model on their own data.

Each of those capabilities strengthens a known attack. Carlini et al. extracted verbatim training sequences, including names, phone numbers and email addresses, by generating many samples and ranking them by likelihood [4]. Nasr et al. later reported extracting gigabytes of training data from open models such as Pythia and GPT-Neo as well as from semi-open and closed ones [5]. Research on fine-tuning shows that a small amount of additional training can amplify privacy leakage and bring back personal information a model had seen [6].

Size does not protect you either. A 2025 study found that a small open model emitted personally identifiable information when prompted [3]. The working assumption for an open-weights lab should be that any record the model memorized will eventually be found, because the attacker has unlimited time and full access. The companion page on training-data extraction and memorization risk covers the attack mechanics in more depth, and what model memorization is gives a short definition.

Are model weights personal data under the GDPR?

Weights can be personal data whenever personal data from training is extractable from them or can be obtained through queries. EDPB Opinion 28/2024 says anonymity must be decided case by case. A model counts as anonymous only if two likelihoods are insignificant: directly extracting training subjects' personal data, and obtaining that data through queries [1]. Commentary on the Opinion notes that a model able to output its training subjects' personal data cannot be anonymous, and that a non-anonymous model stays within the GDPR [2].

For a release, this has three practical effects. First, distributing the weights is itself processing, and you need a lawful basis for it [1], separate from the one you relied on for training. Second, data subject rights do not stop at the hub: you cannot erase a record from thousands of downloaded copies, so a release you cannot defend as anonymous creates obligations you have no way to meet. Third, the EDPB lists evidence a supervisory authority may review, including documented resistance to attribute and membership inference, extraction testing and regurgitation testing [1].

As of October 2026, the Commission's Digital Omnibus proposal to narrow the definition of personal data is not law, so Opinion 28/2024 remains the working reference. Our page on EDPB Opinion 28/2024 and model anonymity works through the full test. The page on unlawfully processed training data explains what happens downstream if the training itself lacked a lawful basis.

De-identified training data lowers the risk but does not settle the question

De-identifying data before training is the strongest lever you have, but the model can still memorize what de-identification missed. Automated PII detection is probabilistic. Microsoft's Presidio project says plainly that because it relies on trained models, there is no guarantee it will find all sensitive information [8]. Free-text fields such as support-ticket bodies, call transcripts and engineering comments are where detectors fail: a customer's name inside an email signature, an account number written with spaces, or a rare job title combined with a city.

Traditional de-identification also protects against linkage in a static table, which is a different threat from extraction out of a generative model. NIST SP 800-188 contrasts traditional techniques with formal privacy methods such as differential privacy and warns about the limits of the traditional approaches [7]. A HIPAA Safe Harbor or Expert Determination dataset [10] is a sound input. A model trained on it still needs its own extraction testing before release.

If release is the goal from the start, consider differentially private training. DP-SGD bounds how much any single record can influence the weights, which is the property an open release actually needs; see differential privacy for LLM fine-tuning. The cost is utility and engineering effort, and the guarantee depends on how you define a privacy unit. A guarantee per training example is weaker than one per person when a single customer appears in hundreds of tickets.

Check the data license before you check the model

A privacy pass does not authorize a release if the data license forbids distributing derived models. Licenses for operational data commonly separate internal training, deployment behind an API and distribution of weights or derivatives, and some restrict the last. Read for terms that cover "models," "derivatives," "outputs" and "sublicensing," for any obligation to delete or retrain when the license ends, and for record-level takedown duties you could not honor after distribution.

Those clauses decide the question before the privacy review even starts. A license that requires you to retrain without withdrawn records is incompatible with a public checkpoint. The pages on releasing open-weight models trained on licensed data and what happens to trained models when a data license ends cover the rights side.

In the EU, Article 53 of the AI Act requires general-purpose AI model providers to keep a copyright policy and publish a summary of training content. Models released under a free and open-source license are exempt from some documentation duties but not from these two, and models with systemic risk get no exemption [9]. As of October 2026, AI Office enforcement of these duties applies to new models from 2 August 2026. Your training-content summary will tell the public which kinds of licensed data went into the model, so agree that disclosure with the supplier in advance.

A pre-release privacy gate for models trained on licensed personal data

A defensible release decision rests on four kinds of evidence: what went into training, what the license allows, how the model behaves under attack and who signed off. The gate below is a template. Adapt the thresholds to your risk appetite and record the actual results, not just pass or fail.

Illustrative example: invented to show structure; it does not describe an available dataset.

GateEvidence to fileFails when
1. Data inventoryList of every licensed dataset, its fields, its de-identification method and the sample-check resultsAny training source lacks a recorded de-identification method
2. License scopeClause references allowing distribution of weights and derivatives, plus the takedown and termination termsLicense limits use to internal training or API deployment, or requires retraining on withdrawal
3. Canary and exposure testPlanted canary strings in training data, with exposure measured at the release checkpointCanaries are recoverable by sampling or rank highly by likelihood
4. Targeted extractionPrompts using prefixes of real records (ticket headers, email openings), with outputs scanned for names, emails, phones and account numbersAny verbatim personal identifier from training appears in output
5. Membership inferenceLoss-based and reference-model attacks on held-in and held-out recordsAttack advantage is clearly above your agreed threshold
6. Fine-tuning red teamA short fine-tune on look-alike data followed by repeating tests 4 and 5Leakage rises noticeably after fine-tuning
7. Regulatory fileAnonymity assessment against Opinion 28/2024, lawful basis for distribution, AI Act Article 53 summary draftAnonymity cannot be argued and no basis covers distribution
8. Sign-offPrivacy, legal and data-supplier approval with date and checkpoint hashAny approver missing or the hash differs from the released file

Run tests 3 to 6 on the exact checkpoint you will publish, including any quantized variants, because quantization and merging change model behavior. Keep the record in your training data use register so later audits can tie the release back to specific licenses.

Release options when the gate fails

A failed gate does not have to mean the model never ships; it means full open weights are not the right release mode for this checkpoint. Weigh these options in order of residual risk.

  • Retrain without the problem data. Remove the free-text fields or records that drove leakage, tighten de-identification and retest. This is usually cheaper than defending a weak anonymity case.
  • Train with differential privacy. Retrain or fine-tune with DP-SGD at a per-person privacy unit, then repeat the gate.
  • Split the release. Publish a base model trained only on data cleared for distribution, and keep the adapter trained on sensitive licensed data behind an API.
  • Gated access. Release weights only to vetted researchers under a use agreement. This reduces exposure but is still distribution, so the GDPR analysis and the license still apply.
  • API only. Keep the weights private and rely on output filtering, rate limits and monitoring.

Each option changes what your data licenses have to permit. When you plan a release before you buy data, ask for evidence that supports a release case: the de-identification evidence package, consent and notice records, and license terms that address derived models directly.

Sourcing data with the release decision in mind

The cheapest point to fix a release problem is when you specify the data. SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records and documents, and manages the licensing process. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery.

Before delivery, personal details such as names, emails, phones and account numbers are removed or replaced; the method is recorded and a sample is checked, though no method is perfect, which is exactly why the extraction tests above still matter. Health records require HIPAA de-identification by Safe Harbor or Expert Determination. If you plan to release weights, say so in your request, because allowed uses are agreed in the license and nothing is contracted until a supplier agrees. You can describe the data you need on the buyers page, and the privacy and de-identification hub collects related guides.

Request licensed data for an open-weights release

SourceX finds US businesses that hold the operational data you describe, prepares diligence materials on source, rights, preparation and allowed use for each dataset, and agrees pricing and allowed uses in a license, with nothing contracted until a supplier agrees. Data is sourced on request, so a request does not guarantee a match. Describe your dataset and intended release at the SourceX buyers page.

Sources

  1. European Data Protection Board, "Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models" (2024). https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf
  2. Herbert Smith Freehills Kramer, "EDPB issues Opinion on personal data in AI models" (2025). https://www.hsfkramer.com/notes/data/2025-posts/EDPB-issues-Opinion-on-personal-data-in-AI-models
  3. arXiv, "Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models" (2025). https://arxiv.org/pdf/2507.04478
  4. arXiv (Carlini et al., USENIX Security 2021), "Extracting Training Data from Large Language Models" (2020). https://arxiv.org/pdf/2012.07805
  5. arXiv (Nasr et al.), "Scalable Extraction of Training Data from Aligned, Production Language Models" (2025). https://arxiv.org/abs/2311.17035v1
  6. arXiv, "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
  7. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  8. Microsoft (microsoft/presidio project), indexed on pkg.go.dev, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  9. European Commission AI Act Service Desk, "Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  10. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data