Leadership and readiness
Data minimization vs AI value: do you have to delete old records?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
Data minimization does not require deleting every old record. Minimization and storage limitation rules target personal data kept longer than its purpose needs. Operational content stripped of personal details is a separate question, governed mainly by contracts and retention policy. Decision rule: delete personal data you no longer need, de-identify content with continuing value, and document both.
Key takeaways
- Minimization rules are aimed at personal data; operational content without personal details is governed mainly by contracts, retention schedules and legal holds.
- Keeping personal data in case it becomes useful for AI later is one of the weakest retention reasons a company can give.
- De-identifying a record before its deletion date can keep the operational content while removing what minimization rules target.
- Legal holds, tax retention and customer contracts can require keeping records, and return-or-destroy clauses can require deleting them.
- Every keep, delete or de-identify decision should be written down with its reason, legal basis and owner.
What does data minimization actually require?#
Data minimization is a privacy principle that limits the personal data a company collects, uses and keeps to what is needed for a stated purpose. GDPR names data minimization and storage limitation among its core principles, and several US state privacy laws now tie collection and retention to what is reasonably necessary for the purposes a company disclosed.
The principle applies to personal data: names, email addresses, phone numbers, street addresses, account numbers and any detail that can identify a person. A maintenance log describing how a compressor failed, or a code review discussing a race condition, is not personal data in itself once the identifying details are gone.
Which laws may apply depends on where the people in the records live, which entity collected the data and what the privacy notice said at the time. That analysis is assessed deal by deal with counsel, not settled by a general rule.
Personal data and operational content are two separate questions#
Most old business records mix two layers: identifiers about people and operational content about the work. Minimization pressure falls heavily on the first layer and lightly, if at all, on the second.
Deleting whole tickets to remove a requester's email address destroys the troubleshooting history as collateral damage. Redacting the identifiers can meet the same privacy goal for many record families, although whether redaction is enough depends on the law, the notice and how easily someone could re-identify the person.
| Record family | Personal data inside | Operational content | Minimization pressure |
|---|---|---|---|
| Support tickets in Zendesk or Intercom | Requester names, emails, phone numbers, order numbers | The problem, troubleshooting steps and resolution | High on identifiers, low on resolution text |
| CRM opportunity history in Salesforce or HubSpot | Contact names, titles, emails, call notes about people | Deal stages, objections, loss reasons | High on contacts and personal notes |
| Job records in ServiceTitan | Homeowner names, addresses, phone numbers, photos of homes | Equipment, diagnosis, parts used, callbacks | High on addresses and photos |
| Issues and code reviews in Jira and GitHub | Developer usernames, occasional customer details | Defects, review comments, fixes | Low to moderate |
| Quality records such as NCRs and CAPAs | Inspector names, sometimes customer contacts | Defect, root cause, corrective action | Low |
| Candidate records in Bullhorn | Resumes, contact details, interview notes | Placement process | High throughout; often unsuitable |
A decision rule for old records#
The decision rule for old records is simple to state: personal data follows its retention schedule, and operational content can be kept when it is de-identified, not restricted by contract and documented. The steps below apply that rule one record family at a time.
- Identify whether the record family contains personal data and which kinds: customers, employees, candidates or sensitive categories.
- Check for obligations to keep: litigation holds, tax and accounting retention, industry retention rules and warranty commitments.
- Check customer and vendor contracts for return-or-destroy clauses that require deletion whatever privacy law says.
- Where personal data has no current purpose, delete or de-identify it on schedule rather than extending retention for a hypothetical future use.
- Where the operational content has value, de-identify before the deletion date and record the method used.
- Log the decision, its legal basis and its owner in the retention register.
Is keeping data for future AI use a valid retention reason?#
Keeping personal data for an undefined future AI purpose is generally a weak retention justification. Minimization rules ask for a specific purpose, and a vague expectation that records might be valuable someday is close to the opposite.
A second issue is the notice. If the privacy notice in effect at collection did not describe licensing or model training, using personal data for that purpose may require a new notice or consent under some laws. Retaining de-identified operational content is a different decision, and counsel evaluates whether the de-identification meets the applicable standard.
A common mistake is quietly pausing deletion jobs in the help desk or CRM because someone heard old records could be licensed. That leaves personal data past its schedule and creates an inconsistent record of why. A better pattern keeps the schedule running and adds a de-identification step ahead of it for record families that pass a rights review.
Timing matters because some deletion tools are irreversible. Zendesk, for example, lets admins create ticket deletion schedules that keep deleting archived tickets matching their criteria, and deleted tickets cannot be restored. Any de-identified copy of those tickets has to be made before the schedule reaches them, and the export should be logged against the schedule it ran ahead of.
When you must keep records anyway#
Minimization is never the only rule in play. The obligations below can stop deletion or force it, even where privacy law is silent, and a licensing review that ignores them produces a package the company cannot stand behind.
| Obligation | What it can require | Who to ask |
|---|---|---|
| Litigation hold | Preserve relevant records and suspend deletion | Litigation counsel |
| Tax and accounting retention | Keep invoices, ledgers and payroll records; the IRS says to keep employment tax records for at least 4 years after the tax is due or paid, whichever is later | CFO or tax advisor |
| Industry or regulatory rules | Keep specific records such as safety or quality logs | Compliance lead |
| Customer contracts | Keep, return or destroy customer data at term end | Contract owner and general counsel |
| Warranty and service commitments | Keep service history to handle claims | Operations lead |
Illustrative: a scheduling software vendor with an overdue help desk#
Illustrative: a fictional B2B scheduling software vendor has run Zendesk for support and Jira for engineering since its early years. During a retention review, its general counsel finds that closed tickets kept customer contact details far longer than the written retention schedule allows.
The company runs the scheduled purge of contact fields and attachments. Before it does, engineering exports the ticket bodies that link to Jira issues, redacts names, emails, phone numbers and account identifiers, and stores a de-identified copy with a note describing the method. Counsel reviews the customer agreements and excludes tickets that quote customer files, because those fall under confidentiality terms.
The outcome is a closed retention gap and a smaller, documented set of troubleshooting records the company can later assess for licensing without having kept personal data past its schedule.
How SourceX handles minimization questions#
Minimization questions sit in the Rights and Preparation steps of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. The rights review identifies which record families carry personal data and which notices and contracts restrict reuse; preparation removes personal and confidential details before anything leaves the company.
The method is written into the privacy record of the SourceX Evidence Packet. SourceX does not ask companies to extend retention of personal data for licensing, and the initial fit check collects metadata only, not files.
Frequently asked questions
Does de-identified data still count as personal data?
It depends on the law and on how thorough the de-identification is. Some laws treat data as out of scope only when it cannot reasonably be linked back to a person and the company commits not to re-identify it. Pseudonymized data, where a key still exists somewhere, is often still treated as personal data. Counsel should assess the standard for each dataset.
Can we pause deletion schedules while we evaluate licensing?
A pause for personal data is hard to justify on commercial grounds alone. A legal hold can justify suspending deletion; an internal evaluation usually cannot. The safer pattern is to keep schedules running and de-identify the operational content you want to keep before the deletion date, with counsel weighing in on any exception.
Do minimization rules cover employee messages in Slack and email?
Employee details are personal data under many privacy laws, although US state laws differ on how they treat workforce data. The Colorado Attorney General, for example, states that the Colorado Privacy Act does not apply to data maintained for employment records purposes, while other states may treat workforce data differently. Internal messages also mix work content with personal remarks, so they need careful redaction, a review of what employees were told, and often an updated employee notice before any licensing use.
Does minimization apply to business contact details?
Often, yes. GDPR covers business contact details, and US state laws differ on whether people acting in a business role are covered. CRM contact fields are therefore usually treated as personal data for planning purposes, while deal stages and loss reasons can often be kept once the names are removed.
What should a retention register record for these decisions?
For each record family: the system, the personal data categories, the retention period, the legal basis for keeping or deleting, any holds, the de-identification method if used, the decision owner and the date. That record answers a regulator, an auditor or a buyer asking why the records still exist.
Sources
- Zendesk admins can create ticket deletion schedules that delete archived tickets after a set period. Deleted tickets cannot be restored, and the schedules keep deleting any tickets that match their criteria. Source
- The IRS says to keep employment tax records for at least 4 years after the date that the tax becomes due or is paid, whichever is later. Source
- The Colorado Attorney General states that the Colorado Privacy Act does not cover personal data of individuals acting in a commercial or employment context and does not apply to data maintained for employment records purposes. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.