Data pipelines are software, and software needs tests. The principle is not controversial. The challenge is that data pipeline testing has different failure modes than application code, and strategies borrowed directly from software development practice often miss the important failure classes.
A unit test suite that gives you 90 percent code coverage does not tell you whether your pipeline will produce the right results when the source schema changes, or whether the output data is semantically correct, or whether the pipeline produces duplicates after a retry. Catching those failures requires a testing strategy that is shaped around how data pipelines fail, not around how function calls fail.
What data pipelines fail at, specifically
Before choosing what tests to write, it is worth cataloging what actually breaks in production data pipelines. From the failure patterns we have observed building connectors across multiple source systems, the categories are:
Transform logic bugs: the SQL or code that transforms data does the wrong thing. A date calculation is off by one. A JOIN produces a fan-out that inflates revenue counts. A filter condition uses the wrong comparison and silently drops records. These are bugs in the transformation logic itself.
Upstream schema changes: a source system adds, removes, or renames a column. The pipeline ingests the new data but either errors on the missing column or silently returns null values where the old column used to be.
Data quality degradation: the schema did not change, but the values have. A field that was never null is now nullable. A field that contained valid email addresses now sometimes contains raw form-submitted garbage. A field that was a Unix timestamp has changed to an ISO 8601 string.
Infrastructure failures with incorrect recovery: the pipeline fails mid-run, retries, and produces duplicate records. Or it fails, fails to checkpoint its position correctly, skips records when it resumes, and produces a gap.
Each of these categories requires a different type of test to catch.
Unit testing transform logic with static fixtures
Transform logic bugs are the most straightforward to test and the most suitable for conventional unit testing approaches. The transform is a function: it takes data in, produces data out. Test the function with representative inputs and assert on the outputs.
The key is fixture quality. Tests that only exercise the "happy path" (clean, well-formed inputs in the expected schema) will not catch the bugs that occur at the boundary conditions. Write fixtures that include: null values in every nullable column, empty strings in text columns, the minimum and maximum valid values for numeric columns, timezone variations for timestamp columns, and the specific edge cases that your transform logic has conditional branches for.
For SQL transforms (dbt models, stored procedures, or inline SQL in your pipeline code), use a framework that lets you test SQL in isolation with controllable input data. dbt's built-in testing framework, SQLMesh's test harness, or a lightweight approach of creating test tables with known content and running your SQL against them all work. The key requirement is reproducibility: the test must produce the same result every run regardless of what is in the production tables.
Contract tests for source interfaces
Schema change bugs require contract tests: tests that verify the source system's API or schema matches what the pipeline expects. Unlike unit tests that verify your code, contract tests verify your assumptions about an external system.
A contract test for a database source might query the source schema metadata (from information_schema.columns or equivalent) and assert that the columns your pipeline depends on exist with the expected types. When this test runs in CI and the source schema changes, the test fails before the pipeline runs against production data, giving you a chance to update the pipeline code before users see broken reports.
For API sources, a contract test makes a live API call (to a test or staging environment if available, or to a read-only endpoint in production) and validates the response structure matches the schema your connector expects. Specifically: are all required fields present? Do the types match what the connector expects to receive? Are the optional fields that the connector uses actually present in the response, not just documented as optional fields that never appear?
Contract tests sit between unit tests and integration tests in your pipeline CI pipeline. They run more frequently than integration tests (because they do not require a full end-to-end run) and catch a class of bug that unit tests cannot (because unit tests exercise your code but not the contract between your code and the external system).
Data quality tests on pipeline outputs
Data quality tests verify that the data the pipeline produced is semantically correct. These tests run after the pipeline writes to the destination, checking the output data rather than the pipeline code.
The categories of data quality tests worth implementing:
- Freshness tests: is the most recent record in the destination table more recent than N hours ago? If not, the pipeline may have stalled or failed silently.
- Volume tests: did the pipeline write a reasonable number of records? A threshold like "between 50 percent and 200 percent of last run's record count" catches both under-ingestion (something is being filtered or dropped) and over-ingestion (duplicates from a bad retry).
- Null rate tests: for columns that should rarely or never be null, track the null rate. If the null rate for
customer_idjumps from 0.01 percent to 15 percent, something changed in the source or the join logic. - Uniqueness tests: for tables that should have unique records on a primary key, check for duplicates. This catches the idempotency failure where retries produce duplicate records.
- Referential integrity tests: for foreign key relationships, check that values in child tables exist in parent tables. A
product_idin the orders table that does not exist in the products table is a join that will silently drop rows in downstream queries.
dbt tests implement most of these out of the box. For pipelines not using dbt, the same checks can be written as SQL queries that return zero rows on success.
Integration tests: the part most teams skip
Integration tests run the full pipeline end-to-end against a real or realistic source system. They are expensive to run and hard to maintain, which is why most teams skip them. But they catch the failure class that no other test level reaches: the interaction between your pipeline code and the actual behavior of the source system.
A unit test verifies that your pagination code works with your mock response. An integration test verifies that it works with the actual API response, including edge cases in the API's behavior that the mock does not replicate.
The practical challenge is environmental control. Integration tests need a source system to run against, and running against production creates problems (rate limit consumption, sensitive data, irreproducibility). The best approaches are: a dedicated test account with synthetic data in the source system, a vendor-provided sandbox environment (many SaaS APIs offer one), or a locally-running mock server that replays recorded real API responses.
Integration tests do not need to run on every commit. Running them nightly, or triggered on changes to connector code, provides a reasonable coverage-to-cost tradeoff. The goal is to catch the integration-layer failures before they reach production, not to maximize test frequency.
Testing failure recovery: the part almost everyone skips
Recovery behavior after failures is the most under-tested part of most data pipelines. The common pattern is to test the happy path thoroughly and assume the retry logic works. The retry logic is often where the production failures live.
At minimum, test two scenarios: a mid-run failure followed by a successful retry (does the retry produce duplicates?), and a failure at the cursor checkpointing step (does the pipeline skip or double-ingest records on retry?). These tests are not complicated to write. They require being able to inject a failure at a specific point in the pipeline run, which usually means adding a testing hook or using a dependency injection approach that lets tests swap in a failing component.
The test that most teams have never written: deliberately run the pipeline against a source that returns a different schema than expected and verify that the failure is clean (error raised, partial write rolled back, cursor not advanced) rather than silent (pipeline completes successfully with null or missing values in the destination).
Pipeline health monitoring built in
Nava Labs tracks freshness, volume, and schema consistency across all connected sources, alerting you when output data patterns shift.
Get Early Access