Production pipelines break for reasons that feel unique each time but cluster into three patterns when you look across enough incidents. We have seen this across the teams we work with and hit all three ourselves while building Nava Labs. The good news is that each pattern has a corresponding mitigation you can build in from the start, rather than bolting on after the first 2 a.m. page.
Pattern one: the schema you trusted changed without telling you
The most common production failure mode is a source schema change that silently invalidates a downstream transform. A column gets renamed. A nullable field starts arriving as NULL where it previously always had a value. An upstream team adds a new required field, changes a type from INTEGER to BIGINT, or drops a column they assumed nobody was reading.
Your transform either fails with a column-not-found error (visible and recoverable) or silently produces wrong output (invisible and much worse). The silent case is the dangerous one. A column rename where the old name and the new name carry the same data produces no error at all. Your dashboard starts showing zeros or wrong aggregates, and you find out from a stakeholder three days later.
The mitigation is schema contract enforcement at ingestion. Before your transform runs, compare the incoming schema against the expected schema. This does not require a heavyweight data contract framework. At minimum, a checksum comparison of column names and types will surface the rename case. For the subtle cases, especially type promotions or nullability changes, you need field-level comparison.
Where teams go wrong is treating schema validation as a one-time setup task rather than a continuous check. You define the schema at setup time, it matches, and you never look again. Six months later the upstream team promotes a column from NUMERIC(10,2) to NUMERIC(18,2) because they started processing larger values. Your downstream aggregation still works, but you now have a schema mismatch that will mask more severe future changes.
The fix is storing the schema version at each run and diffing it against the previous run. When something changes, the pipeline should stop and alert, not continue and corrupt. Silent forward progress on a changed schema is worse than a hard failure.
Pattern two: the assumptions baked into your transforms
The second category is volume and cardinality assumptions embedded in the transform logic itself. These are bugs that exist from day one but only surface at scale or when the data distribution shifts.
A concrete example: a transform joins a transactions table to a customer table with the assumption that every transaction has exactly one matching customer record. On day one that assumption holds. A year later, after a data migration or a schema relaxation upstream, some transaction IDs match multiple customer records. Your join fans out. A table that was 10 million rows becomes 40 million rows after the join. The transform takes four times as long, or worse, produces duplicated aggregations that look plausible but are wrong.
The variant we see most often is date range assumptions. A pipeline processes "the last 30 days of data" using a hardcoded lookback window. This works until the source starts delivering data late, or until a backfill run pushes older records through the pipeline. The lookback window now misses legitimate records, and the aggregation is systematically understated.
The mitigation here is row-count and cardinality assertions after every major transform step. Your pipeline should be able to answer: did the join fan out unexpectedly? Did the output row count fall outside the expected range? These checks do not need to be precise. A check that says "the row count after this join should be within 20 percent of the input row count" will catch the fan-out case without requiring you to know the exact expected output size.
We surface these assertions as observable metrics on each pipeline step in Nava Labs. When we are querying across a PostgreSQL source and a BigQuery warehouse simultaneously, we track row counts through each transform stage so the fan-out case does not disappear into the aggregate output.
Pattern three: infrastructure reliability assumptions
The third pattern is assuming your sources are more reliable than they actually are. This shows up most often in connectors to third-party APIs, Kafka topics, and change-data-capture streams.
Source APIs have rate limits, maintenance windows, and partial failure modes that your ingestion logic needs to handle explicitly. A connector that assumes every API call succeeds will fail silently when it gets rate-limited, or it will retry infinitely when the source is in a degraded state. Neither behavior is acceptable in production.
The more subtle version of this problem is partial delivery. An API that returns paginated results may successfully return pages 1 through 8 and then time out on page 9. If your ingestion logic marks the job complete after iterating through all available pages, it will silently drop the records that would have been on pages 9 through 20. Your pipeline recorded a successful run. You are missing a month of data from the final pages of every daily pull.
The mitigation is checkpointing at the source level. Before marking an ingestion run complete, verify that the record count from the source matches the record count you ingested. For paginated APIs, track the total count returned in the response header against the sum of records across pages. For CDC streams, verify the committed offset against the expected offset range.
This sounds straightforward, but the implementation is easy to get wrong. We have seen ingestion jobs that check the connector exit code rather than the record count, or that retry on connection failure but not on partial delivery. The connector exits zero because no exception was thrown. The data loss happens silently between the try and the catch.
What makes these patterns hard to catch before production
All three patterns share a property: they are invisible under normal testing conditions. Unit tests for your transform logic use fixture data with a known schema. The schema will not drift in your test environment. Your fixture data will have exactly the cardinality you expect. Your mock API connector will never return a partial page.
This is not a criticism of unit tests. They are the right tool for testing transform logic correctness. They are the wrong tool for testing the categories of failure that actually cause production incidents. The coverage gap is structural, not a matter of writing more tests.
We are not saying unit tests are irrelevant for data pipelines. They catch logic bugs and they run fast. The point is that unit test coverage in isolation gives teams false confidence that a pipeline is production-ready. The additional coverage you need is integration testing against real or realistic source data, with schema variation introduced, at realistic volumes, and with source API calls that can simulate partial results.
Building in the mitigations from day one
The actionable version of each mitigation:
Schema contract enforcement: define the expected schema for each source at connection time and check it on every run, not just at setup. Store the schema snapshot so you can diff it across runs and see when something changed. A column name change that your pipeline silently accepted six months ago is now findable in the audit trail.
Volume and cardinality assertions: add assertions after every join and aggregation step. Start with row counts. Add cardinality checks on join keys when there is any reason to believe the source might have non-unique keys. These assertions run in seconds and will save you hours of debugging wrong aggregations.
Checkpointing at ingestion: verify record counts at the source before marking an ingestion run complete. For streaming sources, verify committed offsets. For batch pulls from APIs, compare the expected record count from the source metadata to the count you actually received and wrote.
None of this requires a new platform. You can implement all three with SQL assertions and basic monitoring on your existing stack. Where a platform like Nava Labs helps is surfacing these checks as first-class observability signals rather than buried scripts that someone will eventually comment out because they slow down the development loop.
The teams that build these mitigations in from the start spend their on-call shifts doing something other than grepping through connector logs at 2 a.m. The teams that discover them through incidents build them eventually, at much higher cost.
See cross-source queries in action
Nava Labs connects your databases, warehouses, and event streams. Run SQL across all of them without moving data.
Get Early Access