Data Quality
If a dashboard shows the wrong number, who catches it?
|
8 minutes

Software engineering has solved the regression problem with automated testing. Data engineering is still catching up. The risks are the same, and something that worked last week quietly breaks this week, but in data, there is no equivalent of a failing test suite to warn you. There is only a stakeholder who eventually notices that the numbers do not look right. Here is why data platforms need the same discipline that software teams have been applying for years, and what it looks like when you build it in from the start.
The problem software solved, but data has not
Ask any software development team how they know their application still works after a change. They will tell you about their test suite: unit tests, integration tests, end-to-end tests that run automatically on every commit and every morning. If something breaks, a pipeline goes red. The feedback loop is fast and specific.
Now ask a data team the same question. How do you know the revenue number in this morning's dashboard is correct? In most cases, the honest answer is some version of: we check it manually, we compare it to last week, we notice when something looks off. The process is human, intermittent, and dependent on someone having the time and context to spot the anomaly.
This is not a criticism. It reflects how the two disciplines evolved differently. Software QA built its automation culture over decades of painful experience with bugs in production. Data is going through the same maturation now, just faster and with higher stakes, because more decisions are being made on data than ever before.
The parallel is direct. In software, a regression is when a change breaks something that used to work. In data, the equivalent happens constantly: a source system changes its schema, a transformation breaks silently, a join condition that was correct six months ago is no longer correct for the business logic it represents. The output looks the same. The number is wrong. The difference is that in software, automated tests catch it before it ships. In data, it ships. And then someone eventually notices.
What "automated data quality" actually means
When software teams talk about test automation, they mean a suite of tests that runs without human involvement, reports failures specifically, and blocks bad code from reaching production. The goal is to make quality a property of the system, not a property of how carefully someone checked on a given day. Data quality automation follows the same logic, applied to pipelines and data models instead of application code.
In practice, it means checks that run every time a pipeline runs or on a schedule, if the pipeline is continuous. Schema tests verify that columns exist, types are correct, and key fields are not null. Freshness checks confirm that data is arriving on schedule and that tables are not silently stale. Referential integrity checks confirm that the relationships between tables still hold. Range checks flag values that are statistically implausible: the transaction amount of zero, an age of two hundred, a date in 1970. And end-to-end reconciliation checks compare the output of one calculation path against an independent path to confirm the final numbers are internally consistent.
None of this is new thinking. It is the same discipline software QA has been applying for years: schema tests are the data equivalent of type checks, freshness checks are the data equivalent of uptime monitoring, and end-to-end reconciliation is the data equivalent of integration testing. What changes is the medium.
When a check fails, the failure is specific. Not "the dashboard looks wrong," but "the orders table has not been updated in eighteen hours" or "fourteen percent of rows in the customer dimension are missing a region code." The same specificity that makes a failing test useful in software makes a failing data check useful in a pipeline. You know where to look. You fix it before it reaches anyone making a decision.
Checks handle the detection side. Data contracts handle the coordination side. A data contract is a formal agreement between the team producing a dataset and the teams consuming it, specifying the schema, quality rules, and freshness guarantees, enforced as code rather than documentation. When an upstream team changes a schema, the contract catches the incompatibility before it reaches production. The same principle that made API versioning standard in software engineering applies here: explicit, testable agreements between producers and consumers, rather than implicit assumptions that break silently.

Real world example: a silent error
We built a centralized platform for a global manufacturing company, producing periodic summary calculations that fed into regulatory submissions and board-level reports. Quality checks ran on every pipeline execution, including range checks on key input fields.
At some point, the unit in which one key metric was reported changed silently. No version flag, no schema change, nothing a conventional pipeline would treat as an error. The calculations ran. The outputs looked plausible. A single input field was off by a factor of a thousand, and no one on the team had noticed, because nothing about the output invited a second look.
That is the part worth sitting with. The report did not fail. It produced a result, and the result was wrong by an order of magnitude. On a platform without automated validation, that number would have moved downstream unchallenged, into a regulatory filing and a board pack, and the error would have surfaced weeks later, if at all.
Instead, our range check flagged the values on the next pipeline run: they sat three orders of magnitude outside the historical range for that field. We followed the lineage back to the ingesting source system and corrected the input before any output reached production. The fix took less than a day. The submissions were filed on time, with accurate data.
How a data quality issue travels
A data error rarely stays where it starts. A source system stops sending a field. A schema change goes unnoticed. An upstream table gets refreshed at the wrong time. Any of these can introduce a bad value at the point of ingestion.
From there, the value moves. It flows through the transformation layer, where business logic is applied on top of it. It enters the data model, where other calculations reference it. It reaches the reporting layer, where it appears in a dashboard as a number that looks like any other number.
By the time a stakeholder acts on it, the error may have touched a dozen downstream outputs and informed decisions across multiple teams. No single point in that chain announced the problem. Automated checks at each stage- ingestion, transformation, model, output- are what break the chain before the damage compounds.

The case for building it early
In software, the consensus on test automation is clear: the earlier you build it, the cheaper it is to maintain, and the more valuable it becomes over time. In practice, most codebases were not built that way. Retrofitting coverage takes more effort, but it is routine work that reliably lowers risk and cuts long-term maintenance costs.
The same is true in data. A data platform with quality checks built in from the first pipeline is fundamentally different from one where quality was left to manual validation and then retrofitted later. In the first case, every new pipeline gets added to an existing framework, and the cost of quality stays flat as the platform grows. In the second case, introducing checks later means auditing pipelines that were never designed to be testable, tracing lineage that was never documented, and trying to define what "correct" looks like for outputs that have been running unchecked long enough for no one to be certain.
We have seen enough data projects to know what happens when quality is left to manual checks on a platform that keeps growing; validation gets skipped when time is short, errors get normalized because "the numbers are always a little off", and trust in the data erodes gradually until the platform is producing outputs that nobody fully believes.
We build data quality checks at the start of projects alongside the first pipelines, before anyone has felt the pain of a bad number. Where a platform already exists, we introduce them where the risk is highest and expand from there. Either way, it is the same call software teams made about test automation a decade ago, anticipating a problem that is structurally inevitable rather than reacting once it becomes visible.
What the team does with the time saved
The parallel to software QA extends here too. When regression testing is automated, QA engineers stop spending time re-verifying that last sprint's features still work and start spending it on the things that require human judgment: exploratory testing, edge cases, flows that are technically correct but feel wrong.
Data quality automation creates the same shift. When pipelines are monitored and validated automatically, data engineers and analysts stop spending time manually checking whether yesterday's numbers match today's. That time goes toward the work that actually requires their expertise: modeling new domains, improving transformation logic, understanding edge cases in source data, and asking the harder questions: whether a metric is defined correctly, whether two dashboards that measure the same thing are doing so consistently, whether the data that is being collected can actually answer the business questions being asked of it. Automation does not replace that judgment. It creates the conditions for it.
What this looks like for the business
For the teams and organizations relying on a data platform, the most visible outcome is trust. Not the qualified kind, "we think the numbers are right," but the kind that comes from knowing that the data was validated this morning, automatically, against explicit criteria, and nothing failed. That trust changes how data gets used. When stakeholders are confident in the output, they use it to make decisions rather than spending meeting time questioning it. When engineers know that quality checks will catch regressions, they can move faster on new models without worrying that a change will silently break something upstream.
It changes what happens when something does go wrong, and something always eventually goes wrong. On a platform with quality checks in place, a pipeline failure surfaces immediately, with context, before it reaches anyone downstream. In a platform without them, the failure surfaces when someone notices an anomaly, which may be days later and may be too late.
The connection between QA and data quality is not accidental. Both are answers to the same question: how do you know the thing you built still works? In software, automated testing became the standard because manual checking did not scale. Data is at the same inflection point. The teams that treat data quality the way mature engineering teams treat QA systematically, continuously, and without waiting for something to go wrong are the ones whose stakeholders trust the numbers on the screen.
FAQ
Is this only relevant at scale?
The same argument applies to data as it does to software testing: the foundation is cheapest to build early. A data platform with five pipelines is the right time to introduce quality checks, not a platform with five hundred. By the time it feels necessary, the cost of doing it is much higher.
How much overhead does it add?
Writing quality checks alongside a pipeline adds time upfront. That investment is recovered quickly: checks run automatically, failures are caught before they propagate, and the team spends less time on manual validation. The overhead of maintaining checks as the model evolves is consistently lower than the overhead of validating everything by hand at scale.
How do we know if our current process is insufficient?
If stakeholders regularly question whether the numbers are right, if the team has discovered errors that were present in reports for days before anyone noticed, or if manual validation gets skipped when time is short, those are the same signals that software teams recognized before they invested in test automation. The response is the same, too.
Explore more stories



