Skip to content

Privacy, de-identification and sensitive data

Data clean rooms for AI training: when you can train without receiving raw records

Quick answer

A data clean room can replace delivery of sensitive records when your workload is small enough to run where the data lives and only approved outputs need to leave: aggregate analytics, cohort joins, feature engineering, evaluation runs and modest fine-tuning. It is a weak substitute for large-scale pretraining or iterative LLM fine-tuning, where you need GPUs, many epochs and the trained weights. The deciding question is not privacy alone but whether the license lets model weights exit the environment.

By SourceX Editorial · Updated

Clean rooms exist because sensitive training data is hard to obtain: even de-identified datasets are often difficult to access, which pushes data holders toward controlled access models instead of copies [1]. For definitions, see the data clean room glossary entry and what is a data clean room?. This page covers the buyer decision for model training and evaluation, within the wider privacy and de-identification hub.

What a clean room actually changes in a data deal

A clean room changes who holds the copy, not whether a license is needed. In a delivered-copy deal the records land in your storage and the license governs what you do with them. In a clean room the supplier, or a neutral operator, keeps the records in an environment it controls, and you submit code or queries whose outputs pass through a release policy.

The mechanisms vary. Warehouse sharing such as BigQuery sharing gives the subscriber a linked dataset that can be queried but not modified, without copying the underlying tables [2], and VPC Service Controls rules govern how that shared data can be reached across project perimeters [3]. Open protocols such as Delta Sharing expose Delta Lake and Parquet tables over REST from cloud object storage [4]. Trusted execution environments go further: AWS KMS can release a decryption key only to a Nitro Enclave whose attestation measurements match the key policy, so the data is decrypted only inside approved code [5].

Federated approaches keep data with each holder and move training updates instead; domain-adaptive pretraining of language models has been studied in this setting [6]. Private set intersection lets two parties find overlapping records without exposing non-matches, a common first step for joins [8]. None of these is a clean room by itself, but commercial clean rooms often combine them with query restrictions and output review. The access models guide compares clean rooms with delivered copies, term copies and remote access.

Workloads clean rooms handle well and poorly

Clean rooms handle workloads well when compute is light and outputs are small and reviewable; they struggle when you need heavy GPU training and the full weights. Use this table before you spend weeks negotiating an environment that cannot run your job.

Illustrative example: invented to show structure; it does not describe an available dataset.

WorkloadClean-room fitWhyWhat leaves the room
Aggregate profiling (counts, distributions, label balance)StrongSQL with minimum cell sizesAggregates above a threshold
Joining your records to supplier records for coverage checksStrongPSI or hashed-key joins [8]Match rates, overlap counts
Model evaluation on held-out sensitive dataStrongInference only, fixed test setMetrics, error categories
Classical ML (gradient-boosted trees, logistic regression)ModerateCPU-friendly, small model artifactsModel file, if the license allows
LoRA or adapter fine-tuning of a small LLMModerate to weakNeeds GPUs in the operator's tenancy; adapters can memorizeAdapter weights, if allowed
Full fine-tuning or continued pretrainingWeakLarge GPU clusters, many epochs, multi-terabyte checkpointsFull weights, rarely allowed
RAG index constructionWeakEmbeddings and chunks are derived copies of the textUsually nothing

Evaluation is the cleanest fit, because nothing trained on the data needs to leave. The enclave and supplier-hosted evaluation guide covers that case in depth. For training, the friction usually comes from three places: the operator's environment lacks the accelerators your stack expects, your training code and dependencies must be audited before they run, and every iteration of hyperparameters passes through a review queue.

Why model export is the real negotiation

Model export is the central term because trained weights can carry the data out of the room. Fine-tuned language models can memorize and regurgitate training records, so a clean room that blocks raw-row export but lets unrestricted weights leave has moved the risk rather than removed it. The memorization and extraction risk guide explains how extraction attacks work and how to test for them.

Suppliers typically choose among four export postures. They can allow only metrics and aggregates out. They can allow weights out after a memorization test, such as canary insertion or extraction probes on known sensitive strings. They can require differentially private training, accepting the privacy-utility cost that DP imposes on language models [7]; the DP-SGD fine-tuning guide covers epsilon choices. Or they can keep the model hosted in the room and serve it to you by API only.

Ask which posture applies before you design the experiment. A team that plans LoRA fine-tuning and learns at the end that adapters cannot leave has wasted the compute budget.

Clean room vs data license: terms to negotiate

A clean-room arrangement still needs a written license, and its terms differ from a delivered-copy license in predictable ways. Treat the operator's platform terms and the supplier's data license as separate documents, and make sure they agree.

Illustrative example: invented to show structure; it does not describe an available dataset.

Clean-room term sheet checklist

  • Permitted computations: named job types (SQL aggregates, evaluation, training of specified model classes), not "analytics."
  • Output release policy: minimum aggregation threshold (for example k of at least 50 per cell), who reviews releases, and the review turnaround you can plan around.
  • Model export: whether weights, adapters, embeddings or distilled models may leave; required tests before release; DP parameters if required.
  • Derived artifacts: status of feature tables, synthetic data and logs created inside the room. See the synthetic data privacy guide.
  • Code and dependency approval: who audits your container images and whether proprietary model code is visible to the supplier.
  • Compute: GPU types and quotas available, who pays, and whether you can bring your own accelerators via an attested enclave [5].
  • Logging and audit: query logs retained by whom, and whether the supplier can see your prompts, queries or evaluation sets.
  • Term and teardown: what happens to your code, outputs and intermediate state when access ends.
  • Jurisdiction and access location: where the environment runs and which personnel can reach it.

The last point matters for regulated data. The DOJ rule at 28 CFR Part 202 restricts and prohibits certain transactions that give countries of concern or covered persons access to bulk U.S. sensitive personal data, and access, not only transfer, is in scope [9]. A clean room reduces exposure but does not by itself settle that analysis; see the DOJ bulk sensitive data rule guide.

Privacy controls a clean room does not replace

A clean room limits where data goes, but the records inside still need de-identification appropriate to the use. Operators and suppliers' staff can see the data, your code can log it, and outputs can leak it through small cells or memorized text. Redaction tools reduce exposure but are not complete; Microsoft's Presidio project itself warns that its ML-based detection offers no guarantee of finding all sensitive information [10].

So ask for the same evidence you would for a delivered dataset: the de-identification method, the residual-risk assessment and sample checks. The de-identification evidence package checklist lists the documents. If you plan to join clean-room data with your own, review linkage and mosaic risk first, because joins are exactly what clean rooms make easy.

A decision path for procurement and platform teams

Choose a clean room when the workload is evaluation or analytics, the supplier will not release a copy, and your team can work inside someone else's environment. Choose a delivered, de-identified copy under license when you need GPU-scale training, many iterations and weights you own.

  1. Write down the workload: model class, parameter count, expected epochs and GPU hours.
  2. Write down what must leave: metrics, model weights, adapters, embeddings or nothing.
  3. Ask the supplier which export posture it accepts and which compute the environment provides.
  4. Run a pilot job in the room to measure turnaround per iteration before committing budget.
  5. Compare the total cost and calendar time against a de-identified delivered copy with a license defining records, uses, term and delivery.

If step 3 rules out weights leaving, a clean room is an evaluation or analytics arrangement, not a training one. That is still valuable, for example to qualify a dataset before you license a copy. When you need the copy itself, you can also describe the data to SourceX as an alternative route.

Sourcing sensitive training data when a clean room will not fit

SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, documents, and finance and legal workflows; it does not hold stock, and a request does not guarantee a match. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced before delivery and the method recorded. Describe the sensitive training data you need.

Frequently asked questions

Can I train a large language model inside a data clean room?

Only if the operator provides the accelerators and lets the result leave. Many clean rooms are built around SQL and lightweight compute, so continued pretraining or full fine-tuning is rare; evaluation and adapter-scale training are more realistic.

Is federated learning the same as a clean room?

No. Federated learning keeps data with each holder and sends model updates to an aggregator [6], while a clean room keeps data in one controlled environment that you query or run code against. Both reduce raw-data movement, and both still need output controls because updates and weights can leak training data.

Does a clean room remove the need for de-identification?

No. Operator staff and your code can still see records inside the room, and outputs can still reveal individuals, so de-identification and output thresholds are complementary controls, not alternatives.

Sources

  1. Stanford University (GRACE journal), "Training on Sensitive Data: Governance and De-identification for Machine Learning". https://ojs.stanford.edu/ojs/index.php/grace/article/view/3837
  2. Google Cloud, "Introduction to BigQuery sharing". https://docs.cloud.google.com/bigquery/docs/analytics-hub-introduction
  3. Google Cloud, "Sharing VPC Service Controls rules". https://docs.cloud.google.com/bigquery/docs/analytics-hub-vpc-sc-rules
  4. Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
  5. Amazon Web Services, "Using cryptographic attestation with AWS KMS - AWS Nitro Enclaves". https://docs.aws.amazon.com/enclaves/latest/user/kms.html
  6. arXiv, "FDAPT: Federated Domain-Adaptive Pre-Training for Language Models" (2023). https://arxiv.org/pdf/2307.06933
  7. arXiv, "Differentially Private Knowledge Distillation via Synthetic Text Generation" (2024). https://arxiv.org/html/2403.00932v2
  8. arXiv, "Asymmetric Private Set Intersection with Applications to Contact Tracing and Private Vertical Federated Machine Learning" (2020). https://arxiv.org/pdf/2011.09350
  9. U.S. Department of Justice, National Security Division, "Preventing Access to U.S. Sensitive Personal Data and Government-Related Data by Countries of Concern or Covered Persons (Final Rule, 28 CFR Part 202)" (2024). https://www.justice.gov/d9/2024-12/NSD%20104%20-%20Data%20Security%20-%201124-AA01%20-%20Final%20Rule_0.pdf
  10. Microsoft (microsoft/presidio project), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data