Ask ten executives what data governance is. You’ll get ten answers: a catalogue, a compliance policy, a team, a quality dashboard, a privacy project. That scatter is telling. It explains why so many organisations invest in governance without ever reaping the benefits: they treat a problem of architecture and accountability as a problem of tooling.

The symptom is universal. Two departments produce two “revenue” figures that don’t reconcile. The same customer exists under four identifiers across four systems. A regulatory report takes three weeks of manual reconciliation because no one knows which source is authoritative. So the organisation buys a data catalogue, appoints a lead, launches a programme — and eighteen months later, the two figures still don’t reconcile.

The reason is simple: governance doesn’t first answer “which tool?” but three structural questions only architecture can settle — who owns the data, what does it mean, and what is its lifecycle? This article lays out those three questions, the model that resolves them, and the failure modes to avoid.

Why this is becoming urgent

Data governance isn’t new. What has changed is the cost of not having it.

Data has sprawled. Cloud, managed services and the proliferation of SaaS have exploded the number of places where the same piece of information lives. The neatly bounded central warehouse has given way to a fragmented landscape where data duplicates faster than anyone can govern it.

Decisions have automated. As organisations industrialise analytics and decision support, data quality stops being a reporting nicety: it becomes the raw material of operational choices. Bad data no longer produces a wrong report — it produces a wrong decision, at scale and without human supervision.

Regulation has hardened. Personal-data protection, sector-specific requirements, processing traceability: regulators no longer only ask “do you protect the data?” but “do you know where it is, where it came from, and who accesses it?”. Those are questions of lineage and ownership — questions of governance.

What these three forces share: none is solved by buying software. All require that you have first answered who decides what about the data.

The three questions governance must settle

1. Who owns the data?

This is the founding question, and the most dodged. “Ownership” doesn’t mean legal ownership or technical possession of the server. It means accountability: who has the authority to decide a data element’s definition, its expected quality, its access rules and its lifecycle?

The classic mistake is to hand data to IT. IT hosts and operates the data; it has no legitimacy to decide what an “active customer” or a “terminated contract” is — those definitions belong to the business. Hence a distinction every mature organisation ends up formalising:

  • The data owner is a business leader. They are accountable for the definition, quality and usage rules of a data domain. It’s a role of authority, not execution.
  • The data steward operates that accountability day to day: maintaining definitions, monitoring quality, arbitrating anomalies. They are the owner’s operational arm.
  • The data custodian, on the IT side, ensures storage, security and technical availability.

Until these roles are named — by name, not “team X” — governance stays an intention. Data with no owner is data no one has a mandate to fix.

2. What does the data mean?

Two systems storing a “customer” field almost never store the same thing. One counts prospects, the other doesn’t; one includes closed accounts, the other excludes them. No platform resolves this divergence: it is semantic, not technical.

The answer is a business glossary — a shared reference of definitions, arbitrated by the owners, that is authoritative across the whole organisation. It’s not a technical dictionary of columns; it’s the contract of meaning that finance, sales and operations agree on. Without it, every data reconciliation replays the same dispute over what “revenue” means.

To that glossary, architecture must add two more references: reference data (the shared nomenclatures — countries, currencies, product codes) and master data (the pivot entities — customer, product, supplier — of which only one authoritative version should exist). Master Data Management isn’t a parallel project: it’s the bedrock that guarantees a “customer” is the same customer everywhere.

3. What is the data’s lifecycle?

Data is born, flows, transforms, expires and, one day, must disappear. To govern is to master that trajectory: where the data comes from (its lineage), which transformations it undergoes, how long it’s retained, and when it must be purged.

This dimension has become non-negotiable for two reasons. Compliance first: proving that a piece of personal data is deleted at the right time requires knowing every place it has spread to. Trust second: a decision-maker who can’t trace a figure back to its source won’t rely on it — and data no one trusts is worthless, whatever its actual quality.

The model: federate, don’t centralise

Faced with all this, the natural temptation is centralisation: one team, one platform, one point of control. It’s the approach that fails most reliably. A central team will never know each business domain finely enough to arbitrate its definitions; it becomes a bottleneck, then a service the rest of the organisation routes around.

The model that holds is federated. It rests on a simple principle: accountability for data belongs to those who produce and understand it — the business domains — while a light central body sets common rules and guarantees interoperability. This is the logic behind modern data mesh approaches: treat data as a product, owned by each domain, with shared standards that make those products interoperable.

Three pillars structure this federated model:

  • Accountable domains. Each business domain owns its data, exposes it cleanly to others, and answers for its quality. Ownership is distributed, not diluted.
  • Common standards. A minimal central governance defines what must be shared: the glossary, exchange formats, security and classification rules, lineage requirements. It arbitrates interoperability, not the content of each domain.
  • Data contracts. At the boundary between domains, a contract makes explicit the structure, semantics and quality guarantees of what is exchanged. It’s the data equivalent of the interface contract between services: it lets a domain evolve without breaking those that depend on it.

This model is not a tool choice. It’s a choice of organisational and technical architecture — the distribution of responsibilities and boundaries — that the enterprise architect is best placed to draw, because they see both the business domains and the flows that connect them.

Why a catalogue is never enough

One organisation in two confuses “buying a data catalogue” with “governing its data”. The catalogue is useful — it makes data findable and documents lineage. But it only reflects governance; it doesn’t create it.

A catalogue deployed without named owners fills up with orphan entries no one maintains. A catalogue without an arbitrated glossary documents the existing confusion instead of resolving it. A catalogue without an accountability model becomes a graveyard of metadata: exhaustive, current in month one, obsolete by month six.

The rule is constant: the tool comes after the model. You’ll only usefully catalogue what you’ve first decided the ownership, definition and lifecycle of. Reversing the order — tool first, governance later — is the shortcut that costs the most, because it gives the illusion of progress while leaving the underlying problem untouched.

The failure modes to know

Governance programmes fail in predictable ways. Naming them is already half of avoiding them.

Committee governance. A programme that produces policies, charters and meetings, but no change in the systems or in real accountabilities. Governance exists on paper; the data stays ungoverned.

The grand inventory. Trying to catalogue and document all data before governing anything. The effort exhausts itself before reaching the data that mattered, and the programme dies of its own ambition.

Phantom ownership. Owners named on an org chart but with no authority, no allocated time, no consequence for inaction. An owner who can decide nothing is not an owner.

Quality without purpose. Measuring data quality for its own sake, unconnected to a concrete decision or risk. You produce metrics no one uses, instead of hardening the specific data a specific decision depends on.

All-or-nothing. Waiting for the perfect platform and the complete model before starting. Useful governance is built domain by domain, beginning where the cost of uncertainty is highest.

A trajectory, not a programme

Governance isn’t deployed in one block; it’s built in stages, starting where the pain is real.

Start with a high-stakes domain. Choose a scope where uncertainty about the data visibly costs a lot — a regulatory report, a customer view, a contested steering figure. Name an owner there, arbitrate the definitions, trace the lineage. Prove it on that domain before extending.

Institutionalise the roles. Formalise the owner and steward functions, give them a mandate and time. Governance lives or dies by the reality of these roles, not by the sophistication of the platform.

Set the common minimum. Establish the glossary, the classification and security rules, the exchange standards. As little as possible, but respected everywhere.

Tool what’s already governed. Deploy the catalogue, quality controls, automated lineage — on domains that already have an owner and definitions. The tool then amplifies real governance, instead of simulating absent governance.

Extend through contracts. Domain after domain, connect the scopes with explicit data contracts. Governance grows as a mesh, not as a control tower.

What leadership must decide

Data governance is often framed as a technical topic. That’s a level error. The decisions that make it succeed or fail are leadership decisions: accepting that data belongs to the business and not to IT; giving owners real authority and real time; agreeing to start small rather than wait for the perfect programme; and treating data as an asset one answers for, not a by-product of applications.

None of these trade-offs is technical. All are architectural in the strong sense: they concern the distribution of accountability, the boundaries between domains and the contracts that connect them. That is precisely what enterprise architecture makes visible and decidable — before a single tool is chosen. The question is never “which catalogue should we buy?”. It is: who owns what, what does our data mean, and where does it go? Until those answers exist, no tool will invent them.