Privacy and preparation
Compute-to-data: can a buyer train on your data without receiving it?
By SourceX Editorial · Updated
Short answer
A buyer can sometimes train on your data without receiving a copy, using compute-to-data setups such as seller-hosted training, confidential computing enclaves or federated learning. These cost more effort than a standard delivery, and the trained model still leaves your environment, so the deciding question is whether the reduced exposure justifies the extra engineering on both sides.
Key takeaways
- Compute-to-data moves the training or evaluation job to the records instead of moving the records to the buyer.
- Model weights trained on your records leave your control even when the raw records stay put.
- Seller-hosted training is the easiest variant to explain and usually the heaviest for your own IT team.
- Remote evaluation is often the most practical starting point when source code or trade secrets are involved.
- Rights review and privacy preparation are still required, whichever delivery method is chosen.
Can a buyer train on data it never receives?#
A buyer can train on data it never receives by sending the computation to where the data lives, an approach often called compute-to-data. The buyer's training code runs inside an environment the supplier controls, or inside a hardware-isolated enclave, and only agreed outputs such as model updates or evaluation scores come back.
The idea appeals to CTOs because a raw export feels like the least reversible step in a license. In practice compute-to-data is less common than standard delivery, because it shifts infrastructure, security and support work onto both parties. It is a delivery design, not a privacy treatment, and the records still need rights review and de-identification.
Privacy-enhancing technologies are now part of the standard vocabulary of dataset documentation. The Data & Trust Alliance's Data Provenance Standards, for example, list federated learning, homomorphic encryption, secure multi-party computation and differential privacy among the privacy-enhancing tools a dataset record can declare.
Which compute-to-data options exist?#
Compute-to-data options range from simple remote evaluation to cryptographic methods that remain impractical for most large-scale training. Five patterns cover most conversations a supplier is likely to have with a model developer.
- Seller-hosted training or remote access: the buyer's code and base model run on infrastructure you own or rent, sometimes through a supervised remote workspace its engineers log in to, and the trained weights are released after review.
- Confidential computing enclave: records are decrypted only inside attested hardware in a cloud region, a design intended to keep both the cloud operator and the buyer from reading them in the clear.
- Federated learning: model updates are computed where each data source sits, and only those updates are combined centrally.
- Remote evaluation: the buyer's model is tested against your records inside your environment, and only scores and error summaries leave.
- Cryptographic computation: homomorphic encryption or secure multi-party computation lets parties compute on encrypted data, at a heavy performance cost.
Feasibility and cost compared#
Feasibility and cost differ sharply across these options. The table uses relative terms, because real costs depend on dataset size, model size and how much relevant engineering each side already has in place.
| Option | What leaves your control | Effort for your team | Compute cost | Practical fit |
|---|---|---|---|---|
| Standard encrypted delivery | A de-identified copy of the records | Low once preparation is done | Carried on the buyer's own infrastructure | Most licensed datasets |
| Seller-hosted training | Trained weights and job logs | High: hardware, access control, support | Usually high; who pays is negotiated | Very sensitive records and a committed buyer |
| Confidential computing enclave | Weights and attested outputs | Medium: cloud setup and key management | Cloud costs plus enclave overhead | Buyers already using enclave tooling |
| Federated learning | Model updates from each site | Medium to high: software at every site | Spread across participating sites | Records split across many locations or affiliates |
| Remote evaluation | Scores and error summaries | Medium: host a test harness | Low to moderate | Benchmarks and evaluation sets |
| Homomorphic or multi-party computation | Encrypted intermediate values | Very high, specialist skills | Very high for large models | Narrow analytics, rarely full training |
What still leaves your environment?#
Model weights are the part of a compute-to-data arrangement that always leaves, and they carry information about the training records. Large models can memorize rare strings such as names, account numbers or distinctive phrasing, and a released model can sometimes reproduce them when prompted.
That is why compute-to-data does not replace de-identification. Removing personal and confidential details before training protects you whether the records travel or the model does. The license should also define which outputs may leave, how logs are handled and whether the buyer may keep intermediate checkpoints.
Federated learning has its own leak path: the updates sent from each site can reveal properties of the underlying records. Differential privacy reduces that risk but also reduces the usefulness of the updates, so the trade-off has to be tested on real workloads rather than assumed.
When does compute-to-data make sense?#
Compute-to-data makes sense when records are too sensitive or too large to move comfortably and the buyer is committed enough to share the engineering. For most mid-sized companies licensing support, CRM or operations records, a well-prepared copy delivered securely is simpler to run and easier to audit.
| If this is true | Lean toward |
|---|---|
| Records are de-identified and the buyer needs repeated experiments | Standard encrypted delivery |
| The archive is multi-TB and costly to transfer | Encrypted drives, or training close to your storage |
| Records reference source code or trade secrets you will not release | Remote evaluation or seller-hosted training |
| The buyer only needs to measure model quality on your workflows | Remote evaluation |
| Affiliated companies hold similar records they cannot pool | Federated learning, if the buyer supports it |
| Your IT team has no capacity to run GPU infrastructure | Avoid seller-hosted training |
Questions to ask a buyer who proposes compute-to-data#
Questions to a buyer should establish who builds, runs and pays for the environment before anyone designs it. Clear answers at this stage prevent a technical pilot from turning into an open-ended infrastructure project.
- Which option do you propose, and have you run it with other suppliers at this data size?
- What hardware, cloud services and software must we provide, and who pays for them?
- Which outputs leave our environment: weights, checkpoints, logs, metrics or samples?
- How do you prevent the model from memorizing and reproducing specific records?
- Who on your team gets access to our environment, under which security controls and for how long?
- What happens to weights trained in our environment if the license ends?
Illustrative: a software company splits its delivery#
Illustrative: a fictional vertical software company holds years of Jira issues, GitHub pull requests and Zendesk tickets linked to releases. A model developer wants both the support history and the code review history. The CTO is comfortable licensing de-identified tickets but not the proprietary source code that the pull requests reference.
The company splits the request. Tickets and issue discussions go through standard preparation and encrypted delivery. For the code review set, the buyer agrees to remote evaluation: its model runs against review tasks inside the company's own cloud account, and only scores and anonymized error categories leave. The repository never moves, and the license lists exactly which outputs may be exported.
How SourceX handles delivery choices#
SourceX treats the delivery method as part of the Delivery step of the SourceX five-step transaction, decided after Rights and Preparation have settled what may be shared and in what form. Large datasets stay in the seller's own storage or ship on encrypted drives; SourceX does not host multi-TB datasets.
Whatever the method, the SourceX Evidence Packet records permitted use and release authorization, so the supplier and the buyer work from one record of what may leave the supplier's environment and under whose approval.
Frequently asked questions
Is compute-to-data the same as a data clean room?
They overlap. A data clean room usually lets parties run approved queries or joins on combined data and see only aggregated results. Compute-to-data for AI focuses on running training or evaluation jobs where the records sit. Some clean room products support model training, so check what a specific setup actually allows.
Does compute-to-data remove the need for a license?
No. The buyer still uses your records to build or test a model, so permitted use, outputs, term and restrictions belong in a written license. The agreement should also cover access to your environment, security obligations and what happens to weights trained there.
Can compute-to-data help with records a contract says we cannot share?
Only sometimes. If a customer contract limits use of its data to providing your service, letting a buyer's code train on those records may still fall outside what the contract permits, even though nothing is transferred. The question is what the contract allows, not where the computation happens, so the rights review comes before any delivery design.
Who pays for the compute in a seller-hosted setup?
That is negotiated. Training compute can be significant, and buyers who ask for seller-hosted training often expect to cover it, directly or by supplying hardware or cloud credits. Settle it in writing before any infrastructure is ordered or reserved.
Can we move from remote evaluation to standard delivery later?
Yes, if both sides agree and the records are prepared for release. Some suppliers start with remote evaluation so a buyer can measure fit, then license a de-identified copy once the value is clear. Each stage needs its own scope and approval.
Sources
- The Data Provenance Standards' "Privacy Enhancing Tools" code list includes data anonymization, encryption, masking, minimization, redaction, differential privacy, federated learning, homomorphic encryption, k-anonymity, l-diversity, pseudonymization, secure multi-party computation, t-closeness and tokenization. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.