Privacy and preparation
Should de-identification happen before data leaves your systems?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Yes: de-identification should be finished before records leave your control for a licensee, either inside your own systems or in a controlled processing environment bound by contract. Your own systems give the most control; a controlled environment adds capacity and tooling. Letting the buyer de-identify means identifiable data leaves first, which defeats the purpose.
Key takeaways
- Identifiable records should never reach the licensee; de-identification is complete before delivery.
- Processing inside your own systems keeps raw records under your existing security controls and customer contracts.
- A controlled processing environment can work when it is contractually bound, isolated, logged and deletes raw copies afterward.
- Calling a cloud detection API, copying exports to a shared drive or emailing a sample all count as data leaving.
- Record where and how processing happened, because a buyer's counsel will ask.
Where can de-identification happen?#
De-identification can happen in three places: inside your own environment, in a controlled processing environment run under contract, or at the buyer after delivery. Only the first two are sound choices for a data license.
The location decides who sees raw records. That one fact drives most of the legal analysis, because privacy laws and customer contracts care about who receives identifiable data, not just what the final dataset looks like.
Public-sector guidance frames the same choice. NIST SP 800-188, published in September 2023, tells agencies to assess goals and risks first and then pick a data-sharing model: publishing de-identified data, publishing synthetic data, offering a query interface that applies de-identification, or using a protected enclave. A data license usually corresponds to the first model, which is why preparation finishes before anything is released.
| Option | Who sees raw records | Your control | Effort on your side | Fits when |
|---|---|---|---|---|
| Your own environment | Your staff only | Highest | Highest | Sensitive records, strict customer contracts, very large archives |
| Controlled processing environment | Named people at a contracted processor | High, through contract and audit | Moderate | You lack tooling or people, and contracts allow a processor |
| At the buyer | The licensee | Low | Lowest | Not recommended for licensed records |
De-identifying inside your own systems#
De-identifying inside your own systems means exports, detection, replacement and QA all run on infrastructure you control, under the access rules and contracts you already have. Raw records never move, so no new party enters the picture.
The cost is effort. Someone has to stand up tooling, write custom rules for internal identifiers, run QA and keep logs. Open-source detectors such as Presidio run on your own servers, which makes this route practical for teams with some engineering capacity.
This route also suits very large archives. Multi-terabyte exports are slow and risky to move, so preparing them where they already sit and shipping only the de-identified output, online or on encrypted drives, avoids moving raw data at all.
What makes a processing environment controlled?#
A processing environment is controlled when a contract, technical isolation and logging together ensure that raw records are used only for preparation and then deleted. A vendor's general-purpose cloud service does not meet that bar by default.
Before choosing this route, check your customer agreements. Many enterprise contracts list approved subprocessors or require notice before adding one, and a preparation vendor can count as a subprocessor for customer content.
- A written processing agreement that limits use to preparation for your package and forbids any other use.
- An isolated workspace, ideally inside a cloud account you own, with access granted to named individuals.
- Access and activity logs you can review.
- No copies on personal devices, shared drives or email.
- Deletion of raw and intermediate files at the end, with written confirmation.
- Subprocessor approvals or notices required by your own customer contracts.
Why the buyer should not do the de-identification#
The buyer should not do the de-identification because the transfer of identifiable records is itself the event that privacy laws, privacy notices and customer contracts govern. Once raw data reaches the licensee, the disclosure has happened, however carefully the buyer cleans it afterward.
Buyer-side cleaning also changes the legal shape of the deal. A license of de-identified records may fall outside personal data rules, while a transfer of raw records for the buyer to clean can look like a disclosure, or a sale, of personal data. It also puts your customers' details in a system you do not control and cannot audit.
Offers to strip identifiers on the receiving side are sometimes made in good faith to save the seller effort. The better answer is to accept help with tooling and formats while keeping the raw records at home.
How to choose between your own systems and a processor#
The choice between your own systems and a contracted processor usually comes down to contracts, capacity and the nature of the records. Work through the questions below with IT and counsel before the first export, and write the answers into the project file.
When answers point in different directions, contracts win. A processor with better tooling does not help if your customer agreements forbid adding one without notice, and fixing that after records have moved is far harder than asking first.
| Question | If yes | If no |
|---|---|---|
| Do customer contracts restrict new subprocessors? | Use your own systems, or obtain approvals first | Either route can work |
| Can your team run tooling, custom rules and QA? | Your own systems | A controlled processing environment |
| Is the archive too large to move safely? | Prepare it where it sits | Either route can work |
| Do records include health, financial or other sensitive details? | Your own systems, with a tighter review | Either route can work |
| Does a client own part of the content? | Exclude that content before any processing | Proceed with the remaining scope |
The quiet ways data leaves before anyone decides#
Data often leaves your systems before anyone formally decides to share it. Sending text to a cloud detection API moves raw records to that provider, which your agreements and customer contracts need to cover. So does uploading a sample to a vendor's demo portal, attaching an export to an email, or copying files to a shared drive that outside parties can reach.
Staging is the usual weak point. Exports land in a downloads folder, a temporary bucket or a laptop, and stay there long after the project ends. Pick one staging location inside your control before the first export, and delete intermediate files on a schedule.
Samples deserve the same care as full exports. If a buyer wants to see what the data looks like, show a de-identified sample or a synthetic mock-up, never a raw extract.
Illustrative: a consulting firm keeps preparation in-house#
Illustrative: a fictional operations consulting firm wants to license internal playbooks and project review notes stored in SharePoint and Teams. The documents name client companies, client staff and the firm's own consultants, and some include candid performance comments.
The firm's general counsel rules out sending documents to an outside service, since its client engagement letters restrict disclosure of client information. Instead, the firm runs detection and review on a virtual machine inside its own cloud subscription. A preparation specialist helps through the firm's own accounts, with access logged and removed at the end. Client deliverables and performance comments are excluded entirely.
Only de-identified playbooks and review notes leave the firm. The general counsel keeps a short memo describing where processing ran, who had access and when raw copies were deleted.
How SourceX approaches processing location#
SourceX is built around keeping raw records with the supplier. Nothing is shared during the initial assessment, which collects metadata only. Large datasets stay in the supplier's own storage or ship on encrypted drives after preparation, and SourceX never hosts multi-terabyte datasets.
In the SourceX five-step transaction, Preparation finishes before Approval and Delivery, so the supplier approves a de-identified package, not a promise to clean one later. Where processing happened and who had access is recorded in the privacy record of the SourceX Evidence Packet.
Frequently asked questions
Is encryption in transit enough if the buyer will clean the data?
No. Encryption protects records on the way, but the buyer still receives and can read identifiable data on arrival. The question privacy laws and contracts ask is who receives personal data, and encryption does not change the answer.
Can the buyer's engineers help configure our tools?
Yes, as long as they never see raw records. They can advise on formats, tag conventions and file structure using synthetic examples or already de-identified samples. Keep the configuration work and the raw data on separate tracks.
Our records already live in a SaaS helpdesk. Have they already left our systems?
They sit with a processor working for you under your agreement, which is different from disclosure to a licensee. Exporting them into your own environment for preparation is normal. Check whether the vendor's terms limit bulk exports or certain uses before you start.
Do we need a separate contract with a preparation vendor?
Usually yes. A processing agreement should limit use to your project, set security and access rules, require deletion of raw data at the end, and address any subprocessors the vendor itself uses. Check your own customer contracts for subprocessor notice duties too.
Does local processing slow the project down?
It can add setup work, but it often saves time later because there is no vendor onboarding or contract negotiation. Starting with one record family, proving the pipeline and then widening the scope keeps the effort predictable.
Sources
- NIST SP 800-188 (September 2023) tells agencies to assess goals and risks and then choose a data-sharing model: publishing de-identified data, publishing synthetic data, offering a query interface that applies de-identification, or using a protected enclave. Source
- Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization in text and images. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.