Privacy, de-identification and sensitive data
Differentially private synthetic text from licensed sensitive data
Quick answer
Differentially private (DP) synthetic text is generated so that the output distribution changes only within a bounded amount (epsilon, delta) whether or not any single record is in the source corpus. Teams build it two ways: fine-tune a generator with DP-SGD and sample from it [5], or query a foundation model through an API and select candidates with a DP mechanism (Private Evolution) [1]. It lowers memorization risk for SFT, eval and augmentation, but utility drops, and the guarantee only covers what the privacy unit defines.
By SourceX Editorial · Updated
What the DP guarantee actually covers in synthetic text
The guarantee covers the privacy unit you declare, at the epsilon and delta you account for, and nothing outside the mechanism. If the unit is "one ticket" but one customer appears in 400 tickets, that customer's influence is bounded only per ticket, so group privacy degrades the effective bound roughly in proportion to the number of contributions. For licensed operational corpora (support threads, claim notes, legal matter files), the right unit is usually the user, account or matter, not the row.
Three things sit outside the guarantee. Hyperparameter search on private data consumes budget unless you account for it. Public seed prompts, few-shot exemplars or a pretrained model that already saw the same records are not protected by the mechanism at all. And any non-DP step, such as a quick manual "fix" of synthetic outputs while looking at real records, breaks the accounting chain.
For the broader question of whether synthetic output can still be personal data, see is synthetic data derived from licensed records still personal data. This page stays on methods; the wider privacy and de-identification guide for licensed AI data covers the rest of the cluster.
The two families: DP fine-tuned generators and API-based Private Evolution
You choose between training a generator under DP-SGD or never training on private data and instead using DP selection over API samples. They differ in compute, required access and failure modes.
DP fine-tuned generator. Start from pretrained weights, fine-tune on the private corpus with DP-SGD (per-example gradient clipping plus calibrated Gaussian noise, with a privacy accountant composing loss across steps) [6], then sample as many synthetic records as you like; sampling is post-processing and costs no extra budget [5]. The cost is GPU memory for per-example gradients, large batch sizes to keep noise manageable, and careful conditioning (labels or control codes) so the synthetic set covers the classes you need. The detailed training mechanics are in differential privacy for LLM fine-tuning with DP-SGD.
API-based Private Evolution (Aug-PE style). A foundation model generates candidate texts from public prompts; each private record "votes" for its nearest candidates in an embedding space; noise is added to the vote histogram; the top candidates are varied and resampled over several rounds [1]. No private gradient ever touches model weights, so it works with closed models you can only call by API. Its weak points are domains far from the base model's distribution, where candidates never get close to the private data, and the fact that the embedding model and the vote counts are where privacy is spent.
DP distillation. A third pattern uses a DP-fine-tuned teacher to generate synthetic text that trains a smaller student, which moves the privacy cost into one teacher run and lets you train many students without additional budget [4].
Illustrative example: invented to show structure; it does not describe an available dataset.
| Decision factor | DP fine-tuned generator | API-based Private Evolution |
|---|---|---|
| Access needed | Open weights, GPUs for per-example gradients | Inference API plus an embedding model |
| Where budget is spent | Every DP-SGD step | Each noisy voting round |
| Best fit | Domain-specific jargon, structured fields, long documents | Text close to general web language, short records |
| Main failure mode | Low utility at small epsilon; mode collapse to frequent templates | Candidates never reach rare domain content |
| Private data leaves your enclave | No | Only if the API runs outside it; keep voting local |
| Typical downstream use | SFT, classification data, eval set seeding | Augmentation, classifier training, prompt corpora |
Utility under realistic privacy budgets
Utility drops as epsilon shrinks, and conservative noise calibration can cost more utility than the privacy requirement demands [2]. In practice, small budgets keep frequent patterns (common intents, standard phrasing) and lose rare classes, long-tail entities and exact numeric relations, which are often what a domain model needs.
Plan for these failure modes before you commit compute:
- Tail erasure. Rare categories (escalations, fraud flags, uncommon diagnoses) shrink or vanish because their gradient or vote signal is swamped by noise.
- Template collapse. Generators converge to a few high-likelihood shapes; distinct-n and self-BLEU drop while perplexity looks fine.
- Field incoherence. Structured elements (amounts, dates, SKUs, account states) become internally inconsistent across a record.
- Label drift. Conditioned generation produces text that a classifier trained on real data assigns to a different label.
Measure utility with train-on-synthetic, test-on-real (TSTR) against a licensed real holdout that never entered the DP pipeline. The holdout design is covered in using a licensed real-data holdout to validate synthetic training data, and broader metrics in assessing synthetic data quality.
Why published DP synthetic-text benchmarks need a skeptical read
Treat headline numbers from earlier DP synthetic-text papers as provisional, because at least one later study found that some reported results depended on violated data assumptions [3]. The common problems are contamination and leaky setups.
Check four things before relying on a paper's epsilon-versus-accuracy curve. Was the "private" benchmark (often a public review or news dataset) already in the base model's pretraining data, so the generator recalls rather than learns it? Was the privacy unit a sentence or document when users contribute many? Was hyperparameter tuning done on private data without accounting? Were the evaluation tasks easy enough that templated text scores well?
The safest evidence comes from your own corpus. A sensitive licensed corpus of internal business records is, by construction, unlikely to be in public pretraining data, which is exactly what makes it a fair test of whether a DP method works for you. Use novelty testing against public web corpora to confirm that before reading results.
Preparing a licensed sensitive corpus for DP synthesis
Prepare the corpus so the privacy unit is explicit and the guarantee is not doing work that de-identification should do first. DP bounds per-unit influence; it does not stop a generator from reproducing a phone number that appears thousands of times across units.
- Assign a stable unit key. Map every record to a pseudonymous user, account or matter ID so you can cap contributions per unit (for example, sample at most k records per account).
- Scrub direct identifiers first. Run PII detection and surrogate replacement before DP. Tools such as Presidio help but state that they cannot guarantee finding all sensitive information [8]; see PII redaction for LLM training data.
- Separate public from private inputs. Seed prompts, schemas and label descriptions must come from public or non-sensitive sources, and you should document which inputs count as public.
- Fix the holdout before any DP run. Hold out real records for TSTR and canary tests; they never enter training or voting.
- Insert canaries. Plant synthetic secrets at known frequencies and test whether the synthetic output or a model trained on it exposes them, as described in training-data extraction and memorization risk.
Privacy budget record to keep with every synthetic release
Keep a written budget record for each synthetic release so reviewers can audit the claim, not just the epsilon value. A bare "epsilon = 4" is not reviewable without the unit, delta, accountant and pipeline scope.
Illustrative example: invented to show structure; it does not describe an available dataset.
dp_synthetic_release:
release_id: synth-support-2026-10-a
source_corpus: licensed_support_threads_v3 # license ref attached
privacy_unit: customer_account_id
contribution_cap: 20 threads per account
mechanism: dp_sgd_finetune # or private_evolution
base_model: open-weights 8B, pretrained (public)
accountant: RDP / PRV (state which)
epsilon: 4.0
delta: 1e-6 # below 1 / number of units
budget_spent_on_tuning: included
public_inputs: [label taxonomy, seed prompts v2]
pre_dp_scrubbing: surrogate replacement, method doc v1.3
holdout: 5% of accounts, excluded before training
canary_test: 50 canaries at 1x/10x/100x, exposure reported
utility: TSTR macro-F1 vs real-trained baseline, per class
synthetic_records: 200000
retention_and_use: per license section on derived data
License terms a DP synthesis plan depends on
A DP pipeline is only usable if the source license lets you create derived synthetic data and keep it after the term. The guarantee says nothing about contractual rights, and license metadata on datasets is frequently missing or wrong: one audit of more than 1,800 text datasets found license omission above 70% and error rates above 50% on popular hosting sites [7].
Before you start, confirm in writing:
- Derived and synthetic outputs are permitted uses, including the specific mechanism (fine-tuning a generator is itself model training).
- Whether synthetic data and the DP generator may be retained after the license ends, or must be deleted with the source.
- Whether synthetic data may be shared with vendors, used for commercial models, or released externally.
- Who owns the generator checkpoint and the canary and holdout results.
The rights question is worked through in generating synthetic data from licensed data: rights to derived datasets, and the blend of real and synthetic data in combining licensed and synthetic data. For terminology, see the synthetic data glossary entry.
Where SourceX fits for DP synthesis projects
SourceX sources operational datasets from US companies on request and manages the licensing process, so a team planning DP synthesis can start from a real, rights-reviewed corpus rather than scraped text. Every dataset is delivered under a license defining records, uses, term and delivery, so derived-data and retention questions can be raised during the Agree step. Personal details are removed or replaced before delivery with the method recorded and a sample checked, though no method is perfect, which is why DP still matters downstream. You can describe the corpus you need on the buyers page; a request does not guarantee a match, and SourceX does not train models.
Sourcing a sensitive corpus for differentially private synthetic text
If your DP synthesis plan needs real operational text such as support histories, engineering records or finance and legal workflows, SourceX looks for US businesses that hold the data you describe, with each release approved by the supplying company. Data is rights-reviewed and delivered through private, access-controlled workflows after an executed agreement. Tell us what you need at sourcex.si/buyers.
Frequently asked questions
What epsilon is acceptable for synthetic text?
There is no universal threshold. Report epsilon together with delta, the privacy unit and the accountant, and decide with reviewers using your own TSTR and canary results rather than a value borrowed from a paper [2][3].
Does DP synthetic text remove the need for de-identification?
No. DP bounds each unit's influence, but identifiers repeated across many units can still surface, and pre-DP scrubbing tools are imperfect [8]. Do both, and record the method for each.
Can I use a closed model API for DP synthesis?
Yes, Private Evolution-style methods only need inference access and a DP voting step [1]. Keep the private records and the voting computation inside your controlled environment and check that the API terms allow your use.
Sources
- NSF Public Access Repository (ICML 2024 paper record), "Differentially Private Synthetic Data via Foundation Model APIs 2: Text" (2024). https://par.nsf.gov/biblio/10575568-differentially-private-synthetic-data-via-foundation-model-apis-text
- Association for Computational Linguistics, "DPGA-TextSyn: Differentially Private Genetic Algorithm for Synthetic Text Generation" (2025). https://aclanthology.org/2025.findings-acl.831.pdf
- arXiv, "Private Synthetic Text Generation with Diffusion Models (arXiv:2410.22971v1)" (2024). https://arxiv.org/html/2410.22971v1
- arXiv, "Differentially Private Knowledge Distillation via Synthetic Text Generation" (2024). https://arxiv.org/html/2403.00932v2
- arXiv, "Synthetic Text Generation with Differential Privacy: A Simple and Practical Recipe (arXiv:2210.14348)" (2022). https://arxiv.org/pdf/2210.14348v1
- Tonic.ai, "How to preserve the privacy of text data with differential privacy and large language models". https://www.tonic.ai/blog/how-to-preserve-the-privacy-of-text-data-with-differential-privacy-and-large-language-models
- Longpre et al., arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Microsoft (presidio project), via pkg.go.dev, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.