ISSUE 2026-09-27SERIES BREC 31TYPE BRIEFgarnetgrid.com
Governance is a control, or it is documentation
Insights
Most governance programmes are scoped as artefacts, and artefacts decay. The failures that matter are structural: rules with no enforcement point, owners who cannot change the system producing the data, quality tests that check shape rather than meaning, lineage bounded by what a parser can see, and a deletion path nobody has ever executed. This is what actually goes wrong, and what to build instead.
The question worth asking before any of this is funded is not whether you need data governance. It is which of the things you are about to build will still be true in eighteen months, when the person who ran the programme has moved on and nobody's working day depends on maintaining it.
Most programmes are scoped as artefacts: a catalogue, a classification taxonomy, a policy set, a council with a monthly meeting. Each can be completed, and none of them changes what happens when an engineer ships a breaking change on a Tuesday afternoon.
So ask of every line item: what fails when this rule is broken, and who finds out? If the answer is that the catalogue becomes inaccurate, you have documentation. Documentation decays at the rate the underlying systems change, which on a growing platform is faster than anyone's appetite for maintaining it, and a catalogue that is quietly wrong is worse than no catalogue because people trust it.
A control, by contrast, has a failure mode. A schema registry that rejects an incompatible change is a control. A build that fails when a new table arrives with no owner and no retention class is a control. A policy document stating that tables must have owners is a statement of intent, and intent is not in the code path.
This is a common way these programmes fail, and it is structural rather than cultural. People do not ignore governance because they are careless. They ignore it because nothing they do that day breaks if they do.
Ownership lands on people who cannot change the producer
The steward model puts a name against a dataset. In practice that name usually belongs to someone in a data or analytics team, while the defect is upstream of them: an application emitting an event whose field means two different things depending on the code path, a third-party API that made a field nullable, a CRM where a sales team has been using a free-text box as a status column.
The steward can document that. They cannot fix it, because fixing it means changing a product team's code, and that team has no local incentive to change it. The governance layer then becomes a well-maintained register of known-wrong data, consumers learn to route around it, and that is how an organisation ends up with several incompatible definitions of an active customer.
Ownership is only real when the owner controls the producer and feels it when the data is wrong. The blunt test is: who gets paged? If the answer for a dataset is nobody, or whoever last touched the pipeline, it is unowned whatever the catalogue says.
There is an unpopular implication. Governance that works pushes work onto the teams that emit data, and those teams did not ask for it. If you lack the organisational mandate to make that stick, say so at the start and build the other thing: a consumer-side quarantine, where the platform validates at its boundary and rejects rather than repairs. That is a legitimate design with honest limits. Calling it upstream governance is not.
Quality tests catch absence, not meaning
Most quality suites assert the things that are easy to assert: not-null, uniqueness, row counts, freshness, accepted ranges. They catch pipelines that broke loudly. They rarely catch the failures that damage decisions, because those failures keep the shape of the data and change its meaning.
Consider what survives a full green suite. A status enum gains a value upstream and the downstream case statement files it under other. An amount column moves from minor units to major units after a payments migration. One of several producers starts writing local time into a column everything else treats as UTC. A soft-delete flag is introduced upstream, so deleted rows persist in every mart that does not know about it. A backfill re-runs with a corrected algorithm and last month's published number no longer reproduces. Not one of those trips a null check. Several pass a range check. The enum case is the worst, because the default branch exists precisely to keep the pipeline green.
The mitigations are unglamorous. Fail closed on unrecognised categorical values and accept the pages, rather than bucketing the unknown into a catch-all. Reconcile against an independent system of record rather than against your own staging tables, because a check derived from the artefact it is judging cannot discriminate between right and wrong. Put units, timezone and grain in the contract, not in a column comment. Snapshot the aggregates you publish, so you can tell when history has moved underneath a report.
Lineage is a parser's opinion
Column-level lineage is sold as a solved problem and is mostly a SQL-parsing exercise. Inside a warehouse and a well-structured transformation project it works, and it is useful. Its boundary is whatever the parser can see, and the paths that matter most for governance tend to sit outside it: application code that joins two datasets in Python, dynamic or templated SQL with a runtime table name, reverse ETL pushing customer attributes into a third-party tool, a BI tool's own extracts, a scheduled export, a notebook, anything a person downloaded.
Two consequences follow. Do not treat lineage as an impact-analysis guarantee; it gives you a lower bound on what breaks, and the gap between that bound and reality is invisible by construction. And if you want a defensible answer to where a field has actually gone, you need egress evidence at the boundary, such as warehouse query logs and export events, rather than a graph inferred from transformation code alone.
Access control at the wrong grain
Classification and masking are usually applied to source tables, because that is where the obviously sensitive columns are, and the model feels complete once they are tagged.
Derived data is where it comes apart. A transformation concatenates names into a display field. A mart joins a pseudonymised identifier to a table that happens to carry the mapping. An aggregate with a small enough group becomes a lookup. None of that inherits the tag unless propagation is enforced when the model is built, and propagating classification through arbitrary SQL is a hard problem rather than a configuration setting.
The workable approach is to make classification part of the model's contract and fail the build when a model with classified inputs declares no classification on its outputs. It is coarse and it over-tags. Over-tagging is the cheaper error.
And again: what enforces it at query time? If masking is attached to a role that the ELT tool's own service account does not use, you have a control that exists and does not hold.
Deletion is the honest test
If you want to know whether a governance programme is real, ask for one individual to be deleted completely, and ask to see the evidence rather than the procedure.
This is where the copies surface. The warehouse and its time-travel window. Backups and their retention. A development snapshot someone took last quarter. Cached extracts in the BI layer. Feature stores. Model training sets. Logs. Event streams, which are append-only by design, where deletion means tombstones, compaction and waiting.
There are honest mechanisms. Encrypting per-subject fields under a per-subject key, then destroying the key, turns an intractable deletion problem into a key-management one. It is not free: you must design for it early, you lose cheap joins on those fields, and the key store becomes the most sensitive system you operate. Strict copy discipline also works, where every copy of production data is registered and expires. That requires the registered path to be the easy path, which requires platform work.
What does not work is a deletion runbook written against the primary store while nobody has enumerated the copies. Do the enumeration first. It usually changes the architecture, and it is better to learn that before you have promised anything to a regulator or a customer.
Coverage metrics get gamed; incident metrics do not
Programmes report coverage: the share of tables with an owner, the share classified, the number of assets catalogued. Every one of those can be satisfied in an afternoon by bulk-assigning a distribution list and a default class, and near a reporting date it often is.
Incident-shaped measures resist that, because something has to actually happen. How many data incidents were found by a consumer before a producer noticed? How long from a wrong number appearing in a report to knowing which upstream change caused it? How many breaking changes reached a consumer with no notice? Those numbers are uncomfortable at first, which is the point. A metric nobody is embarrassed by is not measuring anything.
If you are about to fund a programme, start narrow and make it bite. Take the two or three datasets where a wrong number is genuinely expensive, and give each of them a contract with its producer, a gate in the producer's own pipeline, an owner who is paged, reconciliation against something independent, and a deletion path you have executed at least once for real. Then widen it. A narrow governance layer that holds is worth considerably more than a complete one that is advisory, and building the narrow version tells you early whether you have the mandate to enforce anything, which is the actual precondition and not something a tool choice can supply. Two structural choices earn their keep alongside it. Centralise the mechanism and federate the content: one contract format, one gate, one registry, with definitions and owners belonging to the teams that produce the data. And give every governance artefact an expiry date, because anything nobody is obliged to renew drifts into being wrong while still looking authoritative.