What is Dark Data?

Dark data is data an organization collects and stores but does not use for analysis, decisions, or AI. Typical examples include log files, archived emails, call recordings, scanned forms, and sensor readings that sit in storage after their original purpose ends. The term describes how the data is used, not where it lives: a dataset in a modern cloud warehouse can still be dark if no one queries it.

Data usually stays dark for practical reasons. It may lack documentation of what fields mean, sit in formats that tools cannot parse, contain sensitive details that block reuse, or be spread across systems with no shared identifiers. For example, years of maintenance notes may describe equipment failures in free text that no model has been trained to read. Bringing dark data into AI work starts with finding it, describing it, and deciding which task it could serve.

Frequently asked questions

Why is it called dark data?

The data is stored but not visible to the people and systems that make decisions, so its value stays unused, much like unlit space.

Is dark data the same as unstructured data?

No. Much dark data is unstructured, but structured tables can also be dark if nobody uses them, and unstructured data can be actively used.

What are the risks of dark data?

It adds storage cost and can contain sensitive information that nobody is tracking, while the organization misses value it already holds.