What We Do Insights About Get in touch

Data quality is an operating problem, not a tooling problem

Buying a monitoring tool is the easy decision. Deciding who gets woken up when it fires, and what they are expected to do, is the one that determines whether quality improves.

Most organisations that have a data quality problem have already bought something to fix it. The tool is installed, tests exist, alerts fire. And the business still does not trust the numbers.

That is not a failure of the tool. It is what happens when detection is installed without a response attached to it.

An alert nobody owns is noise

The first thing that happens after a monitoring rollout is a wave of alerts. Some are real, most are not tuned yet, and all of them land in a channel that belongs to everybody, which is to say nobody. Within a few weeks the channel is muted. The tests are still running, faithfully, into a room with no one in it.

The fix is unglamorous and organisational: every check that matters gets a named owner, a severity, and an agreed expectation of what happens when it fires. That is more work than configuring the tool, and it is the work that actually changes outcomes.

If nobody would be woken up for it, it should not page anyone. If somebody should be, say who, in writing, before it fires.

Not everything deserves the same standard

Teams often try to apply uniform quality rules across the whole estate, which guarantees either that the important tables are under-protected or that the unimportant ones generate most of the noise.

A more workable approach is to tier explicitly. A small set of assets carry regulatory, financial or customer-facing consequences: those get tests, monitors, owners and a response expectation. A larger set support internal analysis: those get lightweight checks and no paging. The rest are exploratory and are labelled as such, so nobody builds a board report on them by accident.

Making the tiering visible does something else useful. It converts a vague complaint that "the data is bad" into a specific question about which tier an asset is in and whether that is the right tier.

Measure the response, not the rules

Counting tests is a poor proxy for quality. A team can double its test count and change nothing about how quickly a broken pipeline gets fixed.

More honest measures are operational: how long between a failure occurring and somebody noticing, how long between noticing and resolving, how often a consumer discovers a problem before the monitoring does. That last one is the most telling of all, because it measures the gap between what you are watching and what people actually depend on.

Start where the pain is

Broad quality programmes struggle to hold attention because their benefit is diffuse. Narrow ones succeed because they are legible: pick the report that causes the most arguments, make that one dependable end to end, and let the approach spread on its own evidence.

It is slower to announce and considerably faster to achieve. And it produces the only thing that really matters here, which is a business that stops checking the numbers by hand because it no longer needs to.

Written by the Semantic team. If it is relevant to something you are dealing with, tell us about it

More from Insights

Related reading

All insights

Have a data problem you're trying to solve?

Let's talk about it.