AI uses for records
Dark data: what it is and why AI developers want it
By SourceX Editorial · Updated
Short answer
Dark data is information a business collects and keeps but no longer uses beyond its original purpose: closed support tickets, a retired ERP database, departed staff mailboxes, old job notes and quality logs. AI developers want the parts that record real work from request to result, because that material never reached the public web and is hard to recreate.
Key takeaways
- Dark data is defined by neglect, not by format: clean database tables and free-text notes can both be dark.
- AI developers look for dark data that keeps a request, the work done and the result connected in one trail.
- Machine logs without human decisions, client-owned material and records centered on individuals are usually passed over.
- Rights, preparation cost and privacy burden decide whether a dark archive is worth assessing, not its size alone.
- A first look needs only metadata: the system, the date range, the record families and known restrictions.
What is dark data?#
Dark data is information a business collects, processes and stores in the course of its work but does not use beyond the purpose it was created for. It is called dark because nobody looks at it: it sits in retired systems, old exports and archives, often without an owner, an index or anyone who remembers the schema.
IT industry analysts popularized the term, and in IT and records management it usually comes with a warning about storage cost and security exposure. For business owners it now has a second meaning, because some neglected records describe real work in a way AI developers find hard to obtain anywhere else.
Dark data is not a file type. A retired ERP database with clean tables, years of shared-drive folders and a box of scanned work orders can all be dark. What they share is that the business kept them after it stopped using them.
Dark data examples, and what each shows an AI developer#
Dark data examples differ by industry, but the useful ones have something in common: each records people doing skilled work and shows how that work turned out. The table pairs typical dark records with what they reveal to a team building AI systems.
| Dark data | Typical source | What it shows an AI developer |
|---|---|---|
| Closed support tickets linked to engineering issues | A retired help desk at a B2B software company | How problems are diagnosed, escalated and fixed |
| RFI logs and submittal reviews | Closed project folders at an engineering firm | How technical questions get answered within contract limits |
| Past proposals and project reviews | Departed partners' drives at a consulting firm | How scope and staffing are reasoned through and later judged |
| Job notes and callback histories | A previous field service platform at a trades contractor | How technicians diagnose equipment, and when a fix failed |
| Order exceptions and EDI error logs | A legacy WMS, TMS or ERP at a distributor or 3PL | How operations recover when a transaction breaks |
| Nonconformance reports and CAPAs | An old QMS or spreadsheets on a manufacturer's file server | How quality decisions are made and verified |
| Former employees' mailboxes | Archived email accounts in any company | How negotiations and coordination unfold across many messages |
How dark data differs from ROT, cold and legacy data#
Dark data differs from ROT, cold and legacy data in what defines it: dark data is defined by neglect, ROT by a lack of value, cold data by how rarely it is accessed and legacy data by the age of the system holding it. The distinctions matter when a company decides what to delete, what to keep and what to assess.
The practical split is between dark data that is also ROT and dark data that is merely neglected. The first is a cleanup task; the second deserves a look before anyone deletes it.
| Term | Meaning | Overlap with dark data |
|---|---|---|
| ROT data | Redundant, obsolete or trivial content with no ongoing value | Some dark data is ROT and should go; much of it is not |
| Cold data | Data accessed rarely and kept on cheaper storage | Cold data may be managed and indexed; dark data usually is not |
| Unstructured data | Text, images, audio and documents without a fixed schema | Much dark data is unstructured, but whole databases can be dark too |
| Legacy data | Data held in outdated or retired systems | One of the most common sources of dark data |
| Real-world data | Records produced by actual operations rather than generated for a test | Dark data is often real-world data nobody is using |
Why AI developers want certain dark data#
AI developers want certain dark data because it shows how professionals handled real work inside companies, which public web text rarely captures. A model that answers trivia has no use for a contractor's callback history; an agent that triages service calls does.
Public text is also finite. Epoch AI researchers estimate the stock of human-generated public text at around 300 trillion tokens and project that, if trends continue, language models will fully use it sometime between 2026 and 2032. Records kept inside companies sit outside that public stock by definition.
The strongest material keeps the whole sequence in one connected trail: the original ask, the steps someone took, the choice they made and what happened afterward. A ticket tied to a code fix, an RFI tied to a revised drawing and a carrier exception tied to a credit memo are all examples. Trails like these can help train agents and grade them, and because they were never published, they can form private evaluation sets that are unlikely to have appeared in any model's training data.
Isolated files carry far less. A folder of invoices without the orders, notes and disputes around them shows results without the reasoning behind them.
What makes some dark data worth more than the rest#
Some dark data is worth more because it scores better on the drivers in the SourceX Enterprise Data Value Framework, a methodology SourceX developed rather than an industry standard. The drivers that increase value are uniqueness, domain expertise, human-generated signal, scale, recency, data cleanliness, rights and AI utility, while exclusivity increases price.
Other drivers pull the opposite way. Reproducibility reduces value, because data anyone could generate is easy to replace. Preparation cost and privacy burden reduce net value, because records full of personal details or locked in an unreadable format cost more to make usable. A quick screen asks a few plain questions.
- Was it written by people doing the work, or produced automatically by a system?
- Does it connect what was asked, what was done and how it turned out?
- Can it still be exported from the system where it lives?
- Does the company have the right to license it under customer, vendor and employee terms?
- How much personal or confidential content would need to be removed?
Which dark data AI developers usually pass over#
AI developers usually pass over dark data that shows no human judgment, belongs to someone else or is mostly about individuals. Knowing the common rejects saves a company from preparing archives nobody will license.
Machine output is the biggest category. Server logs, sensor readings and system-generated notifications can be large but rarely show a person deciding anything, unless they sit alongside the tickets or work orders that explain what people did in response.
Ownership and sensitivity rule out the rest. Client deliverables, customer source code and material received under a client's confidentiality terms usually belong to someone else. HR files, candidate records and anything centered on health or personal finances carry a privacy burden that typically outweighs their usefulness, and content the company bought or copied from third parties brings rights it cannot pass on.
Illustrative: a 3PL finds its old WMS database#
Illustrative: a fictional third-party logistics company migrated to a new warehouse management system and kept the old one running read-only on a single server, in case a customer disputed an old shipment. Nobody has opened it in years, and IT wants to switch it off.
Before decommissioning, the COO asks for a quick inventory. The old database holds years of receiving discrepancies, pick errors, damage claims and customer service notes tied to orders, plus EDI error logs. Some client contracts restrict sharing of shipment details, so those accounts are marked as excluded.
The company exports the database to its own encrypted storage, logs it in a data inventory with date range and record types, and retires the server. The export is now an indexed archive with an owner instead of dark data, and the company can decide later whether to assess it for licensing.
A first look without moving any files#
A first look at dark data can happen entirely on metadata, without exporting or sharing a single record. A short note for each archive is enough to rank which ones deserve more attention.
SourceX works the same way. Its fit check collects metadata only, and nothing is shared during the initial assessment. If a record family looks promising, the SourceX five-step transaction moves through Supply, Rights, Preparation, Approval and Delivery, with the company approving each stage and a SourceX Evidence Packet documenting what was released and on what terms.
- The system or location, and whether it still runs or exists only as an export.
- The date range the records cover.
- The record families present, such as tickets, work orders, emails or quality reports.
- Whether records link to each other through IDs, job numbers or order numbers.
- Known restrictions: client contracts, vendor terms, holds and sensitive content.
Frequently asked questions
Is dark data the same as big data?
No. Big data describes large volumes processed with specialized tools, usually for active analysis. Dark data describes information that is not used at all, whatever its size. A small company's retired help desk can be dark data while being modest in volume.
Should we just delete our dark data?
Delete what is ROT, past its retention period and free of legal holds. Keep, index and assess what still describes real work, especially when it is linked and company-owned. Deletion is irreversible, so confirm holds and retention requirements before any purge rather than after.
Is dark data worth money?
Some of it may be, but there is no price list. Licensing value depends on the record type, rights, the preparation needed and whether a buyer wants that kind of data, and it becomes clear only once a buyer engages. Some archives are worth more as an organized internal resource than as a licensing package.
Does dark data contain personal information?
Often. Old tickets, mailboxes and job records include names, contact details and sometimes sensitive notes. Any reuse, internal or external, needs privacy review, and licensing requires removing personal and confidential details before anything is shared. Which privacy laws may apply is assessed with counsel.
Can dark data in a retired system still be used?
Usually, if the database or an export still exists and someone can read the format. The risk is in waiting: hardware fails, licenses lapse and the people who understood the schema leave. Exporting to an open format with a short data dictionary keeps the option open.
Sources
- Epoch AI researchers estimate the stock of human-generated public text at around 300 trillion tokens and project that, if trends continue, language models will fully utilize this stock between 2026 and 2032. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.