Leadership and readiness
What is dark data, and what should you do with it?
By SourceX Editorial · Updated
Short answer
Dark data is information a company collects and stores during normal operations but never uses again: closed tickets, finished project folders, retired system archives, shared inboxes and call recordings. Dark data carries cost and risk because nobody owns or reviews it. Every dark data store should end in one of four decisions: delete, de-identify, archive or assess for licensing.
Key takeaways
- Dark data is defined by neglect, not by format: any record family nobody owns or reviews can become dark.
- The main risks are personal data nobody knows about, retention obligations nobody tracks and storage nobody can explain.
- Discovery comes first, because no decision can be made about a store nobody has found.
- Different stores in the same company usually need different answers among delete, de-identify, archive and assess for licensing.
- Operational records that link requests to decisions and outcomes are the dark data most worth assessing before deletion.
What counts as dark data?#
Dark data is any information a company created or collected while doing its work and then stopped using, reviewing or managing. It is usually still sitting in a live system, a backup, a file server or a cloud account that keeps renewing.
Dark data is not the same as unstructured data. A legacy ERP table of closed service orders is perfectly structured and can still be dark if nobody has queried it in years. Equally, an active knowledge base made of free-text articles is not dark because people use it every day.
The useful test is ownership. If nobody can say who is responsible for a store, what it covers and how long it should be kept, it is dark, whatever its format.
Dark data examples in operating companies#
Dark data looks different in each kind of business, but the pattern repeats: the work finished, the records stayed and nobody went back. The table below lists common examples by company type.
What these examples share is that people created them while doing skilled work. An exception email shows how a dispatcher recovered a missed pickup; an NCR shows how an inspector traced a defect to its cause. That is why some dark data is worth a second look before anyone deletes it.
| Company type | Dark data example | Where it tends to hide |
|---|---|---|
| B2B software | Closed support tickets and abandoned engineering projects | Zendesk, Jira, archived Confluence spaces |
| Engineering and architecture | RFI logs, submittal reviews and markups from finished projects | Project servers, Procore, Bluebeam sessions |
| Consulting | Proposals, project reviews and draft playbooks | SharePoint sites and departed partners' drives |
| Home services and trades | Technician notes, job photos and callback history | ServiceTitan, Housecall Pro, old dispatch systems |
| Logistics and distribution | Exception emails, claims files and EDI error logs | Shared inboxes and TMS archives |
| Manufacturing | NCRs, CAPAs and maintenance work orders | QMS and CMMS tools, spreadsheets on a plant server |
What are the risks of dark data?#
The first risk of dark data is personal information nobody tracks. Old support tickets, HR files saved to a shared drive by mistake and call recordings can contain customer and employee details well past any retention schedule.
The second is security. Retired systems that still run tend to miss patches, keep stale admin accounts and sit outside current monitoring. The third is obligation: records a company keeps may have to be produced in a dispute, and contracts may have required them to be returned or destroyed long ago.
Cost is real but usually the smallest of the four risks. Storage is cheap; not knowing what you store is not.
How to find dark data#
Finding dark data is an inventory exercise, and it works best when several sources are combined. No single list, whether from IT, finance or the people who have been around longest, catches everything.
Expect the first pass to miss things. Treat the result as a living register that is reviewed whenever a system is added, migrated or retired, and whenever a team is reorganized, because those are the moments when new dark data is created.
- Pull every software subscription from accounts payable and company card statements.
- Review admin consoles for inactive workspaces, archived projects and unused shared drives.
- Ask long-tenured staff which old systems and folders they still remember.
- List systems retired in past migrations and confirm whether anything still runs or renews.
- Check shared inboxes, team mailboxes and departed users' drives.
- Note on-premises servers, backup drives and storage closets with physical media.
- Record for each store: system, owner, date range, record family, personal data present and known obligations.
Four options for each dark data store#
Each dark data store should end with one of four decisions, made by its owner with counsel checking holds and obligations. Applying one answer to everything, such as deleting it all or keeping it all, usually gets at least one store wrong.
A practical decision rule: delete what has no legal, contractual or business reason to exist; archive what must be kept; de-identify what is useful but full of personal details; and assess for licensing only records the company owns that show how real work was done.
Deferring is a legitimate choice as long as it is recorded. An archive with an owner and a review date is a decision; a folder left alone because nobody wanted to choose is simply dark data with a longer life.
| Option | When it fits | What to watch |
|---|---|---|
| Delete | No legal, contractual or business need, and mostly personal or duplicate content | Check legal holds first and record what was deleted and why |
| De-identify | The content is useful but identifiers are not needed | Free text needs human review, and the method should be documented |
| Archive | Must be kept for tax, warranty or legal reasons, or the decision is deferred | Name an owner, restrict access and set a review date |
| Assess for licensing | Company-owned operational records linked to outcomes across several years | Rights review, privacy burden and preparation cost |
Illustrative: an engineering firm's project server#
Illustrative: a fictional structural engineering firm is replacing an on-premises project server that holds folders for every project since the firm began, alongside Deltek Vantagepoint for project accounting. The COO uses the replacement as the moment to sort the server's contents.
The inventory finds four kinds of material. HR files saved to a shared folder by mistake move into the HR system with restricted access. Client drawings and deliverables are archived under the terms of each client agreement. Duplicate scans and superseded drafts are deleted after a check for holds.
The firm's own RFI responses and internal review comments are de-identified, with client names and project addresses removed, and set aside for a licensing assessment. The new server starts with an owner for every folder.
Where SourceX fits#
SourceX becomes relevant only for stores that land in the assess-for-licensing column. The fit check uses metadata, such as systems, years of history and record families, and nothing is shared during the initial assessment.
In the SourceX Enterprise Data Value Framework, dark operational records often rate well on human-generated signal and domain expertise, because they capture people solving real problems. They often need work on data cleanliness and carry privacy burden and preparation cost, which reduce net value. Records that are easy to reproduce from public sources rate lower, because reproducibility reduces value. The framework is a SourceX-developed methodology with qualitative ratings, not a price list.
Frequently asked questions
Is dark data the same as ROT data?
They overlap but are not the same. ROT, short for redundant, obsolete or trivial, is a judgment that information has little value. Dark data simply means information nobody uses or manages. Some dark data is ROT and should be deleted; some is valuable history that has been ignored.
Who should own decisions about dark data?
The COO or a similar operating leader usually coordinates, the business owner of each record family makes the call, counsel checks holds and contract obligations, and IT carries out the deletion, archiving or export. Writing down who decided what is as important as the decision itself.
Does cloud software make dark data worse?
It can. Many SaaS tools keep records until someone deletes them, and subscriptions renew without anyone reviewing what accumulates. Cloud storage also makes it easy for teams to create new workspaces that outlive the project that needed them. A yearly review of subscriptions against the inventory catches most of these.
Can dark data be valuable to AI developers?
Some can. Records that show how a company handled real requests, decisions and outcomes over years, such as linked support and engineering history or quality investigations, may interest model developers. Value depends on uniqueness, rights, cleanliness and preparation cost, and many stores will not qualify.
Should we delete dark data before selling the company?
Not on impulse. Deletion during a sale process can raise questions in diligence and may conflict with legal holds or contract obligations. A documented inventory with clear owners and decisions is usually more useful to an acquirer than a recently emptied server.
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.