Getting started
Delete or keep? Weighing data minimization against data value
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Data minimization vs data retention comes down to four factors weighed per record family: legal duty to keep or delete, breach and discovery risk, storage and management cost, and value to the business or outside buyers. Personal data needs a current purpose to stay. Operational content with personal details removed can often stay under a documented business reason.
Key takeaways
- Minimization principles target personal data kept beyond its purpose, not every old business record.
- Legal duty is the first factor: a hold or retention rule overrides cost and value arguments.
- Breach and discovery risk rise with the personal and sensitive content inside a record family.
- De-identification can separate the risky part of a record from its useful part.
- Every keep or delete decision should carry a written reason, an approver and a review date.
Why delete-or-keep is a real dilemma#
Delete-or-keep is a real dilemma because privacy law, security teams and finance push toward deleting old data, while operations and AI projects reward long, deep history. Both sides have a point, and a general counsel has to reconcile them record family by record family.
Minimization principles appear in GDPR and in several US state privacy laws, and some newer state laws set tighter limits on collecting and keeping personal data than earlier ones did. These rules focus on personal data held without a purpose. Operational content with personal details properly removed is generally governed more by contracts and internal policy.
This is general information, not legal advice. Which laws may apply depends on the people, places and data involved, so assess each decision with counsel.
The four-factor decision matrix#
The four-factor decision matrix scores each record family on legal duty, breach risk, storage cost and value, then reads the combination. Score each factor low, medium or high with the record owner, the security lead and counsel in the same room, so no single function decides alone.
| Factor | Questions to ask | Points toward delete | Points toward keep |
|---|---|---|---|
| Legal duty | Does a retention rule, hold, contract or privacy law apply? | A deletion duty, or personal data whose purpose has ended | A retention rule, open hold or contract requirement |
| Breach and discovery risk | How much personal, financial or sensitive content is inside? | Dense personal or sensitive details | Little personal content, or content that can be removed |
| Storage and management cost | What does it take to secure, back up and administer? | A legacy system kept running only for this data | Low-cost, controlled storage with an index |
| Value to the business or buyers | Is it used, or could it be used once prepared? | Duplicates, logs or trivial content | Expert work with decisions and outcomes |
How to read the matrix#
Reading the matrix starts with legal duty, because a duty to keep or to delete settles the question before cost or value come in. Only when no duty applies do the other three factors decide.
The high-risk, high-value pattern is where most debates happen. De-identification lets a company keep the reasoning in a support ticket or job record while removing the names, contact details and account numbers that create most of the risk.
| Pattern | Usual decision |
|---|---|
| Legal duty to keep | Keep until the duty ends, whatever the other factors say |
| Legal duty to delete | Delete, whatever the value, unless counsel identifies an exception |
| High risk, low value | Delete on schedule |
| High risk, high value | De-identify: keep the operational content, delete the personal details |
| Low risk, high value | Keep with a named owner and a review date |
| Low risk, low value, high cost | Delete, or move to low-cost storage if a reason remains |
What does keeping data cost beyond storage?#
Keeping data costs more than storage because every record family needs security, backups, access reviews and a way to answer access requests and discovery. A retired system kept alive for occasional lookups still needs patches, accounts and someone who knows how to query it.
Discovery is the hidden cost. Records that exist can be requested in litigation, and reviewing them takes lawyer time. Data with no duty or value behind it adds that cost without any offsetting benefit, which is the strongest practical argument for minimization.
Is future value a valid reason to keep data?#
Future value can be a valid reason to keep operational content, but it is a weak reason to keep personal data. Privacy principles generally expect personal data to be tied to a defined purpose, and a vague future use may not satisfy that expectation.
The cleaner approach separates the two. Remove or transform personal details on a defined schedule, and keep the operational content under a documented reason, such as internal training, quality analysis or a licensing review. Each decision should leave a short written record.
- The record family and every system that holds it.
- The factor scores and who assigned them.
- The decision: delete, keep, de-identify or archive.
- The method used to remove personal details, if any.
- The approver and the next review date.
Mistakes on both sides of the decision#
Mistakes on the delete side and the keep side tend to mirror each other. Both come from deciding by system or by habit rather than by record family and a documented reason, and both are hard to explain after the fact.
- Deleting a whole system to meet a minimization goal, including records a retention rule still requires.
- Purging old data while a hold is pending or litigation is reasonably anticipated.
- Keeping everything because storage looks cheap, without counting discovery and breach exposure.
- Treating removal of names as de-identification when unique details remain in the text.
- Holding personal data for an undefined future AI use.
- Making decisions without a written record that a regulator or court could review.
Illustrative: a staffing firm decides what to keep#
Illustrative: a fictional staffing firm runs Bullhorn as its ATS and holds years of candidate profiles, job orders, client notes and placement records. A new state privacy law prompts the general counsel to review what the firm keeps and why.
Stale candidate profiles score high on risk and have no current purpose, so the firm deletes them on its schedule after checking for holds. Job orders and client requirement notes score low on personal content and high on value, so they stay with an owner. Placement records with interview feedback are mixed: the firm keeps them for the period counsel sets, then de-identifies the feedback and deletes the candidate identity.
Every decision is logged with a review date. Candidate personal data is excluded from any outside use from the start, and the firm's privacy notice is updated to match the new practice.
How SourceX handles minimization in a licensing review#
SourceX works on the keep side of the matrix only where content is owned by the company and can be prepared. In the Preparation step of the SourceX five-step transaction, personal and confidential details are removed, and the SourceX Evidence Packet records the privacy record alongside provenance, licensing rights, permitted use and release authorization. Record families dominated by personal data are usually left out of scope.
Frequently asked questions
Does de-identified data still count as personal data?
It depends on the law and the method. Some laws treat properly de-identified data as outside their scope, often with conditions such as commitments not to re-identify and contractual controls on recipients. Weak techniques, such as removing names but leaving unique details, may not qualify. Under GDPR Recital 26, identifiability is judged by all the means reasonably likely to be used, such as singling out, weighing cost and time against available technology. Have counsel review the method against the applicable definitions.
Who should make delete-or-keep decisions?
The record owner proposes, security assesses risk, finance weighs cost, and counsel confirms legal duties and signs off. For large or sensitive record families, the decision may go to the executive team or board. Whoever decides, write the reasoning down so it can be explained later.
Can we delete data under a legal hold if it is costly to keep?
Not without counsel lifting the hold or approving a defensible alternative. A hold overrides retention schedules, cost arguments and minimization goals for the records it covers. If a legacy system is expensive to run, counsel may approve preserving the relevant data in another form that keeps its integrity.
How often should we revisit keep decisions?
Give every keep decision a review date, and revisit early when a law changes, a system is retired, the company is acquired or a hold ends. Tying the review cycle to the retention schedule review keeps both documents consistent and avoids decisions that quietly expire.
Does moving data to the cloud change the analysis?
Cloud storage can lower storage cost and improve controls, but it does not change the legal duty, the risk from personal content or the value question. It can add vendor terms on where data is processed and how deletion works, and those terms belong in the same review.
Sources
- GDPR Recital 26 says that to decide whether a person is identifiable, account should be taken of all the means reasonably likely to be used, such as singling out, weighing objective factors such as cost and time against available technology. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.