All articles
Nava Labs Team 8 min read

Data Observability for Engineering Teams: What to Monitor and Why

Observability for your data layer follows the same principles as application observability. Here is a practical framework for what signals matter and how to surface them.

Data Observability for Engineering Teams

Application observability has a clear mental model: collect logs, metrics, and traces. Correlate them when something breaks. Ask questions about system behavior by querying the collected signals rather than reading code. Data observability borrows the same model but applies it to your data layer. The goal is the same: know when something is wrong before a stakeholder tells you.

The practical challenge is that data observability has more failure modes than application observability. An application service either runs or it does not. Data pipelines can produce technically successful outputs that contain wrong values, missing records, or stale data. Success at the infrastructure level does not imply correctness at the data level.

The five signal categories that cover most incidents

When we look at the incident patterns across the data engineering teams we work with, almost every data quality incident is traceable to a failure in one of five signal categories. The categories correspond closely to the five pillars described in the Monte Carlo model and similar frameworks, which in turn reflect what practitioners have found actually matters in production.

Freshness

Is data arriving when it is supposed to? Freshness monitoring answers questions like: has the daily order table been updated in the last 26 hours? Has the Kafka topic consumer advanced its offset in the last 10 minutes?

Freshness failures are the most common category we see. They happen because upstream batch jobs run late, source API rate limits slow down ingestion, or a silent infrastructure failure causes a connector to stop writing without raising an exception. The pipeline looks healthy; it just stopped delivering data.

The signal to monitor is last-updated timestamp, tracked per table and per partition where partitioning is in use. Set alert thresholds slightly wider than your normal delivery window to avoid alert fatigue from minor scheduling drift. A daily table that normally arrives at 03:00 UTC should alert if it has not arrived by 07:00 UTC, not at 03:01 UTC.

Volume

Are you receiving the expected number of records? Volume monitoring tracks row counts per ingestion run and raises an alert when the count falls outside a historical range.

Volume anomalies catch two distinct classes of failure. On the low side, they catch partial delivery: your connector ran but only received 40 percent of the expected records before timing out. On the high side, they catch duplication: a restart or a misconfigured offset committed the same batch twice.

Absolute thresholds are rarely the right approach. Row counts from real data sources vary with business activity. A volume monitor that alerts when fewer than 50,000 rows arrive will fire every Sunday, every holiday, and every time your users take a vacation. A monitor that alerts when the count is more than three standard deviations below the trailing 30-day average for the same day of the week is more useful.

Schema

Have the columns, types, or nullability constraints changed since the last run? We have written about this at length in the schema drift post in this series, but from an observability standpoint the key signal is: capture the schema on every run and diff it against the previous snapshot. Any change should produce a structured event that routes to your alerting system, not just a log line that gets lost in the noise.

The schema signal feeds directly into the lineage signal. When a column change happens, you want to know downstream: which transforms reference this column? Which dashboards are downstream of those transforms? The schema change event plus lineage information gives you impact radius before you start getting stakeholder messages.

Distribution

Are the values in your data within the expected range? Distribution monitoring is the most computationally expensive of the five categories and the one teams usually add last, but it catches a class of failure that the other categories miss entirely: data that is structurally correct but semantically wrong.

The canonical example is a currency conversion bug that produces revenue figures in the wrong currency. Row counts are correct. Schema is unchanged. Freshness is fine. But every revenue figure is off by a factor of roughly 1.1 to 1.5, depending on the exchange rate on the day the bug was introduced. A distribution monitor on revenue values would have flagged this within a few hours. Nobody notices the 3 percent shift in the average revenue per transaction until a quarterly business review.

Start with simple distribution checks: min, max, mean, null rate, and unique value count for key numeric and categorical columns. Run them as assertions after each transform step. A sudden change in null rate or a mean that moves outside the expected range should trigger an alert.

Lineage

What depends on what? Lineage is not a real-time signal in the same way as the other four categories. It is the context that makes the other signals useful. When freshness monitoring tells you a table is stale, lineage tells you which dashboards are about to show incorrect data. When schema monitoring tells you a column was dropped, lineage tells you which transforms will break.

The practical barrier to lineage is that capturing it requires instrumenting your transform code or your SQL. Column-level lineage in particular, where you can trace which input columns contributed to which output columns, requires either static analysis of your SQL or runtime instrumentation. Most teams start with table-level lineage, which is easier to capture and covers the majority of impact-radius use cases.

What to instrument first

If you are starting from zero, here is the order that delivers the most incident prevention per implementation hour.

First, freshness on your most critical tables. Pick the five tables that, if stale, would cause the most visible problems for the business. Add last-updated monitoring with a threshold that gives you advance warning but does not alert on normal variance. This catches the most common failure mode, silent staleness, in an afternoon of work.

Second, volume on your ingestion runs. Add row count assertions that compare each run against the trailing average. This catches partial delivery and duplication, the two most common connector failure modes.

Third, schema drift on your source connections. Every schema change that affects a column used in a downstream transform is a potential incident. Catching it at ingestion rather than at query time shortens the detection window from hours to minutes.

Lineage and distribution monitoring are valuable but take more time to implement correctly. Add them once the first three are stable and you have some sense of the distribution patterns in your data.

The instrumentation trap

One mistake we see repeatedly is over-monitoring. Teams instrument every column in every table and set tight thresholds, resulting in hundreds of alerts per day. The signal-to-noise ratio drops below the point where anyone trusts the alerts. The on-call engineer starts ignoring alert emails. The monitoring system that was supposed to catch incidents before stakeholders do ends up being turned off because it is too noisy.

The principle we use when setting up observability in Nava Labs is: monitor what you would actually investigate if it changed. Not every column matters equally. A created_at timestamp column in an orders table is critical. An internal reference ID column that no downstream query ever references is not. Start with the high-impact columns and add coverage incrementally as you learn which signals actually matter for your specific data stack.

See cross-source queries in action

Nava Labs connects your databases, warehouses, and event streams. Run SQL across all of them without moving data.

Get Early Access