Skip to content

AI data market

What is dark data, and is it worth anything?

By SourceX Editorial · Updated

Short answer

Dark data is information a company keeps but no longer uses, such as closed tickets, old project folders and retired system databases. Much of it is worth little, and some is a liability. Dark data has licensing value only when it records real work with outcomes, the company holds the rights and preparation costs stay reasonable.

Key takeaways

  • Dark data is stored but unused information, and few companies have a complete map of what they hold.
  • Worth is net: uniqueness, human-generated signal and AI utility raise it, while preparation cost and privacy burden reduce it.
  • Machine logs, duplicates and expired drafts rarely have licensing value; linked work records with outcomes sometimes do.
  • A dark data store with unknown rights is a cost until someone reviews those rights.
  • Assess before deleting or migrating, because retired systems are where valuable history is most often lost.

What is dark data?#

Dark data is information a business collects, processes and stores in the course of its work but does not use again. Nobody queries it, reports on it or owns it. It sits in retired databases, archived inboxes, shared drives and the export a vendor sent when a subscription ended.

The term covers structured data, such as old ERP tables, and unstructured data, such as email threads, meeting notes and scanned forms. What makes data dark is not its age or format but the fact that nobody in the company is looking at it. Typical examples in an operating company include:

  • Closed support tickets and chat transcripts from a help desk the company replaced.
  • Proposals, staffing plans and project reviews on an old file server.
  • Mailboxes and Slack history of employees who have left.
  • Job, dispatch and warranty records stranded in a field service system before a migration.
  • Quality records, NCRs and maintenance logs kept for compliance and never analyzed.
  • Application and server logs retained by default.

Is dark data worth anything?#

Dark data is worth something only in specific cases: most stores are worth less than owners hope, and some cost more to keep than they could ever return. Worth takes three forms: value for internal use, value for licensing to AI developers, and negative value, meaning the cost and risk of keeping it.

Internal value comes from analysis your own team could run, such as studying why warranty claims cluster around one product line. Licensing value appears when an AI developer would pay to use the records for training or evaluation under a defined license, while the company keeps ownership. Negative value comes from storage fees, breach exposure, discovery burden and records kept past your retention schedule.

Many dark data stores carry only negative value. The work is finding the few that do not, and four quick questions sort most stores before anyone opens a file:

  • Does the store show people doing real work, with decisions and outcomes, rather than system output?
  • Does the company clearly hold the rights to it, free of client, vendor or employee restrictions it cannot clear?
  • Can it be exported with links intact and prepared without heavy redaction, OCR or format recovery?
  • Is anything in it under a legal hold or past its retention period?

Which dark data stores tend to have value?#

The dark data stores that tend to have value are the ones that show people doing real work and what happened as a result. Being unstructured is not a disqualifier; the most useful stores are often messy text written by experienced staff.

Which dark data stores tend to have value?
Dark data storeLicensing potentialWhy
Help desk history with internal notesOften meaningfulShows problems, reasoning and resolutions written by people
Retired ERP or WMS databaseSometimesUseful if orders, exceptions and outcomes still link; weak if only balances remain
Project files and proposalsSometimesReal expertise, but client-owned deliverables are usually carved out
Departed employees' mailboxesMixedBusiness threads can be useful; personal content raises privacy burden
Call recordingsMixedRich signal, but consent questions and redaction work are heavy
Machine and server logsRarelyHigh volume, little human judgment, easy to reproduce
Duplicates, drafts and personal filesNoneNo outcome, unclear ownership, mostly cost

How to judge worth with the drivers that matter#

Dark data worth can be judged with the drivers in the SourceX Enterprise Data Value Framework, a SourceX-developed methodology that gives qualitative ratings rather than prices or index values. Most drivers raise value; reproducibility reduces it, and preparation cost and privacy burden reduce net value.

Exclusivity sits slightly apart: offering it can raise the price a buyer pays, but it does not change what the records contain. Value itself is known only once a buyer engages with a specific package.

How to judge worth with the drivers that matter
DriverRaises worth whenLowers worth when
UniquenessThe records exist nowhere else in your industrySimilar records are easy to find or buy
Human-generated signal and domain expertiseSpecialists wrote, decided and corrected in their own wordsSystems produced entries with no explanation
Recency and scaleCoverage spans recent years and many casesA thin slice from long ago
Data cleanliness and AI utilityRecords link request, action and outcomeFields are blank, codes changed meaning, links are broken
RightsCompany-owned internal records with known termsClient, vendor or employee restrictions are unknown
ReproducibilityHard to recreate with a simulator or scriptA generator could produce something similar
Preparation cost and privacy burdenFew personal details and clean exportsHeavy redaction, OCR or format recovery needed

Illustrative: a consulting firm's old engagement server#

Illustrative: a fictional operations consulting firm is retiring an on-premises file server that holds years of proposals, staffing plans, project reviews and internal playbooks. IT plans to move only active client folders to SharePoint and wipe the rest.

The managing partner asks for a short review first. Client deliverables and anything under client confidentiality terms are set aside, because clients own or restrict them. What remains is firm-owned: proposal drafts with win or loss notes, staffing decisions, post-project reviews and the playbooks partners wrote.

The firm exports that firm-owned set with folder structure and dates intact, records which client agreements it checked, and flags client names for removal in any later preparation. The export goes to restricted storage, and the firm runs a metadata-only fit check before deciding whether to license any of it. The rest of the server is deleted on schedule.

Mistakes that destroy dark data value#

Most dark data value is lost through routine IT decisions made before anyone asks what the data shows. The same mistakes come up across industries.

  • Migrating only open records and letting the old system lapse with the closed history inside.
  • Exporting tables without the links between them, such as tickets without comments or orders without status changes.
  • Deleting departed employees' accounts before business threads are preserved.
  • Keeping everything forever instead of following a retention schedule, which raises risk without raising value.
  • Sending an archive to a vendor before rights and privacy have been reviewed.
  • Assuming age equals value, when old records without outcomes are still low value.

How SourceX looks at dark data#

In SourceX terms, a dark data store is possible supply, the first stage of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. The first pass collects metadata such as system names, record families, date ranges and known restrictions, and no files are shared.

Stores that pass move to rights review and preparation, where personal and confidential details are removed before anything leaves the company. Anything licensed is documented in a SourceX Evidence Packet, and the company keeps ownership of its records.

Frequently asked questions

Is dark data the same as big data?

No. Big data describes volume and processing scale. Dark data describes use: information that is stored but not used. A small company can hold dark data, and a large warehouse that analysts query every day is not dark. Some dark data is large, such as log archives, but size says little about its worth.

Should we delete dark data to reduce risk?

Delete what your retention schedule says to delete and what has no business, legal or licensing purpose. Before deleting a large store, check for legal holds and take a quick look at what it contains. Deletion is permanent, while a short review costs little, especially just before a system is shut off.

How do we find dark data in the first place?

Start with systems you have retired or plan to retire, shared drives nobody owns, archive mailboxes, vendor exports from cancelled subscriptions and old backups. Ask IT and finance which subscriptions were cancelled in recent years, because the export taken at cancellation is often the only copy left. Record each find in a simple inventory.

Can dark data that contains personal information still be licensed?

Sometimes, after preparation. Personal and confidential details are removed or replaced before delivery, and some stores stay out because the privacy burden is too high. Which privacy laws may apply depends on whose information is involved and where they live, and it is assessed deal by deal with counsel.

Does dark data lose value as it ages?

Some does. Recency is a value driver, and records about obsolete products or processes teach less. Yet records created before generative AI tools became common can carry clearer human authorship, which some buyers care about. The bigger risk is usually not slow decay but sudden loss when a system is retired.

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify