Article icon
Article

The Dark Data Tax: Why Organizations Lose Track of Their Own Data

Key Takeaways

  • Industry research finds that more than half of organizational data is dark, unknown, or untapped.
  • Storage investment alone does not turn data into an asset. Data can remain dark when metadata is incomplete, formats cannot be read by existing tools, information cannot be retrieved through queries, or datasets remain isolated in silos.
  • Enterprise data architecture provides a structured connection between business objectives, business meaning, data sources, formats, relationships, flows, logical structures, and the physical systems in which data is implemented.
  • To avoid foundational breakpoints, architecture and governance must be treated as sustained organizational practices, and policies, standards, and architectural decisions must carry sufficient authority, ownership, and escalation mechanisms to guide delivery.

Why Organizations Lose Track of Data They Already Have

This is the first of a two-part series on dark data: data an organization stores but cannot readily understand, locate, access, or use without consulting specific individuals. The data may be poorly stored, insufficiently documented, inconsistent with agreed formats, or invisible to existing data value chains. These value chains are the connected processes through which data is captured, stored, governed, transformed, shared, analysed, and used for decisions. Enterprise data architecture (EDA) helps maintain this connection from business meaning and logical structure to the physical systems that store, move, and expose the data. Part 1 examines two organizational breakpoints, funding and authority, that can weaken this chain before it is fully embedded. Part 2 explores the value the chain provides when intact and two additional breakpoints that may still limit it.

Dark Data Explained

Organizations are generating and retaining data at unprecedented scale. IDC estimates that global data creation will increase from 213,557 exabytes in 2025 to 527,469 exabytes by 2029. Yet, having more data does not necessarily mean knowing more about it.

Low-cost, scalable storage has made it easier to retain data for longer. What organizations do not always maintain with the same consistency is its context: what the data means, where it came from, how it relates to other data, whether it follows agreed formats, how it moves through systems, and which processes or decisions depend on it.

This is where the dark data problem begins. Estimates vary because studies use different definitions and rely heavily on respondents’ assessments of their own data environments. They nevertheless point to a substantial gap between the data organizations retain and the data they can readily classify, understand, and use.

What Is Being Measured

Figure

Source and Interpretation

Organizational data whose value is unknown and is therefore classified as dark

52%

Veritas, Global Databerg Report, 2016. Based on a survey of 2,550 senior IT decision-makers across 22 countries.

Organizational data identified as redundant, obsolete, or trivial

33%

Veritas, Global Databerg Report, 2016.

Organizational data classified as business-critical

15%

Veritas, Global Databerg Report, 2016. This means classified as business-critical, not necessarily the total share actively used.

Organizational data reported as dark, untapped, or often unknown

55%

Splunk, The State of Dark Data. Based on responses from more than 1,300 business and IT leaders across seven economies.

Business executives reporting that their organizations create measurable value from data

32%

Accenture and Qlik, The Human Impact of Data Literacy, 2020. This measures value realization rather than dark data.

Estimated average annual organizational cost of poor data quality

At least $12.9 million

Gartner, Data Quality: Why It Matters and How to Achieve It. This concerns poor data quality and should not be interpreted as a dark-data estimate.

Data Architecture Bootcamp

Learn how to design modern data architectures that unify operational, analytical, and AI data – September 2026.

The evidence does not demonstrate that most stored data has no value or contributes nothing to decisions. It shows that a large proportion is unclassified, insufficiently understood, untapped, redundant, obsolete, or difficult to connect to business use. Some dark data may be valuable, but that value cannot be assessed or realized until the data becomes visible and understandable.

These findings also do not prove that organizations lack investment or intent. Storage platforms, data teams, governance programmes, and architecture documentation may already exist. The breakpoint appears when these capabilities do not preserve a unified and accessible context around the data. Storage without sufficient metadata, relationships, lineage, ownership, standards, and connection to business processes behaves less like a usable asset and more like an ongoing cost, risk, and unrealized opportunity.

The Knowledge Problem

What does it mean to know your data?

It does not mean knowing every dataset the organization stores. It means knowing the data that supports its most critical operations, decisions, regulatory obligations, and services. The Veritas Global Databerg study estimated that only 15% of enterprise data was classified as business-critical, while 52% was dark because its value had not been determined. The real risk is that some critical business data may sit within that dark data.

Knowing critical data means being able to establish where it lives, what business concept it represents, how it relates to other data, and how it flows from the system where it was created to the report, decision, regulatory submission, or application that depends on it. It also means knowing how it was transformed, by which rules, what its quality is, who is accountable for it, and what would be affected if it changed.

Many organizations cannot answer these questions reliably for their most important data – not because the information was never documented, but because it was captured within a project while systems, processes, and requirements continued to change. The documentation did not always follow. Knowledge became incomplete, outdated, or dependent on specific individuals.

Closing this gap requires more than enterprise data architecture alone. Data governance provides the accountability, decision rights, ownership, and standards needed to manage critical data, while stewardship, metadata management, and other data management practices help keep that knowledge usable and current. Within this broader system, enterprise data architecture is one of the most important structural elements. It connects business meaning, logical structures, and rules to the systems, databases, and flows where data is implemented.

The answer is usually not technical. The chain may have no durable organizational home. Responsibility may be unclear, authority may be limited, and funding may end when the project closes. Two of these breakpoints are foundational enough to weaken the chain before it becomes embedded in systems.

Where the Chain Breaks

The chain breaks in four distinct ways, described across this two-part series. They are not independent: Each one degrades the conditions that the layers below it depend on, which is why organizations that fix individual symptoms keep encountering the same underlying problems. This article covers the two most foundational breaks. Part 2 covers the two that determine whether the chain reaches daily work and stays current once it does.

Break 1: The Discipline Is Funded as a Project, Not as a Practice

The knowledge chain from business meaning to physical systems requires continuous maintenance. Conceptual models, which define the organization’s core business entities and what they mean, must be updated when the business changes. Logical models, which translate those concepts into attributes, relationships, and rules, must be revised when requirements change. Data lineage, which maps how data moves and is transformed between systems, must be corrected when systems are added, modified, or retired. The business glossary, which provides shared definitions for important terms, must remain current as their use evolves. None of this work usually has its own project code or produces a visible deliverable for leadership. Yet it is what enables every other data investment to create value rather than accumulate cost.

The pattern in many organizations is to fund this work during a major initiative, such as a platform migration, regulatory response, or analytics programme, and treat the resulting artifacts as completed outputs. The models are produced. The glossary is populated. The data flows are documented. The project closes. Funding ends. The business and its systems continue to change.

The artifacts may be accurate when delivered, but within months the gap between what they describe and how the organization operates begins to grow. A new system may introduce business entities that have no governed relationship to the enterprise conceptual model. A changed process may use an existing definition differently. A transformation rule may be revised without the lineage being updated. Within a few years, a model intended to describe the business can become a historical record, accurate for an earlier state of the organization but misleading for the current one. Data governance initiatives frequently stall, not because governance is impossible, but because the strategic commitment that launched them was not converted into the sustained responsibility and investment needed to keep the knowledge chain current between initiatives.

Immediate: Reframe the next budget discussion around the cost of the last time an incomplete or outdated chain was discovered. This may be a migration overrun, an analytics initiative delayed because teams could not agree on what the data meant, or a regulatory submission requiring weeks of manual reconciliation. Attach a cost to the maintenance gap before requesting funding to close it.

Tactical: Conduct a line-item review of where data investment is concentrated: infrastructure, visualization, architecture and modelling, governance operations, stewardship, and metadata management. Metadata management keeps essential information about data current, including its meaning, origin, ownership, format, and use. The imbalance between spending on visible technology and spending on the practices that maintain context reveals where the knowledge chain is structurally underfunded.

Structural: Establish a recurring operating budget for maintaining the knowledge chain. Rather than relying on project allocations that expire at delivery, provide annual funding for the roles, tools, and ongoing work required to keep business definitions, models, metadata, lineage, and system mappings current as the organization and its data landscape change.

Break 2: The Chain Has No Authority to Bind

This version stays very close to the original length while adding the proliferation of small initiatives and briefly explaining each role and concept.

The abstraction chain, meaning the link from a business concept through its structured definition to the physical systems that store it, is useful only when it shapes delivery. The organization’s agreed definitions of critical data must be standards that new systems, integrations, dashboards, and analytical products follow, not optional references. This requires authority: the ability to require alignment, challenge contradictory definitions, and escalate unresolved deviations.

Many organizations create the appearance of this authority without giving it practical force. A chief data officer may be unable to require a business unit’s new system to align with the enterprise conceptual model before go-live. A data steward, who manages a domain’s definitions and quality expectations, may maintain a glossary that projects can disregard. An architecture team may produce models that are consulted selectively. Meanwhile, small initiatives proliferate across departments: Local applications, dashboards, automations, data marts, and pilots are developed without assessment against enterprise data architecture. The roles and artifacts exist, but alignment remains optional.

This matters because the abstraction chain crosses organizational boundaries. One locally convenient but conceptually incompatible definition creates a breakpoint. Repeated across small initiatives, these breakpoints become difficult to see and govern. Each solution may appear harmless in isolation, yet connected systems inherit its definitions, transformations, and quality assumptions. By the time the inconsistency is detected, it may be embedded across applications, reports, and years of downstream data.

Enterprise data architecture therefore needs accountable roles to operate as a governing capability rather than a reference library. Enterprise data architects maintain models across domains. Data stewards maintain definitions and quality standards. A data governance council, a standing cross-functional body, resolves disputes and makes binding decisions when local priorities conflict with enterprise standards. Delivery governance must also make smaller initiatives visible early enough for review.

Immediate: Identify who can require a new system or local initiative to align its definitions, structures, and flows with the enterprise model before production. If the answer differs by project size or funding source, document the gap and bring it to the governance forum.

Tactical: Publish a data accountability matrix linking significant data to a named owner, steward, and architect. Add an architecture checkpoint for local applications, dashboards, automations, data marts, and pilots so small initiatives do not remain invisible until integration problems appear. Where accountability is unclear, say so.

Structural: Give data stewards defined authority, not only responsibilities, and give the chief data officer or equivalent function a formal mandate to require architecture alignment and escalate non-compliance. Embed proportionate EDA review into project, procurement, and change governance. Evaluate stewardship on the quality, consistency, and currency of governed data, not on meeting attendance.

Where Part 2 Picks Up

Both breaks above are, at root, data governance failures. Data governance is the discipline of deciding who owns a data domain, who approves its definitions, and who has the authority to make that decision binding across the organization. Without it, enterprise data architecture has accurate models and no way to enforce them.

Funding and authority can both be fixed, and the chain can still fail. Part 2 picks up there.

Want to become a data governance specialist?

Gain a solid foundation in data governance and train for certification.