Ingestion code is usually the first thing a data engineering team writes and often the last thing they document. The patterns that cause the most production pain are rarely exotic. They are structural mistakes that felt like reasonable shortcuts during the initial build and become load-bearing bugs once data volumes scale or source API behavior shifts.
These five patterns cover the failure categories we encounter most often. None of them are obscure edge cases. Most teams hit at least three of them in the first year of running a production ingestion layer.
Anti-pattern one: full table replacement without a transaction boundary
The pattern looks like this: truncate the target table, then insert the new records from the source. It is simple to implement and easy to reason about. It is also a guaranteed data loss scenario when anything goes wrong between the truncate and the insert.
The failure mode is not theoretical. A network interruption, a timeout on the insert, or a constraint violation partway through the load leaves the table empty. If the pipeline runs nightly and nobody is watching the row count, the table can stay empty for 12 hours before someone notices. If the table feeds a dashboard that is only checked weekly, longer.
The fix is a staging pattern. Write the new records to a staging table first. Validate row count and schema in the staging table. Then, inside a transaction, truncate the target and insert from staging. If the transaction fails, the previous data is still intact. For warehouses that support atomic swap operations, use those instead: create the replacement table, validate it, then rename it into the target slot and drop the old table atomically.
The staging pattern adds a few minutes to your load job. It eliminates the data loss window entirely.
Anti-pattern two: offset-based pagination without a stable sort order
When pulling records from a paginated API using offset and limit, the correctness of your pagination depends on the sort order being stable across pages. If the underlying data is sorted by creation timestamp and new records are being inserted while you paginate, pages shift under you. Records that were on page 3 when you fetched page 1 may have shifted to page 4 by the time you fetch page 3. You skip them entirely.
This is not a hypothetical scenario. Any API that returns results sorted by a mutable field (updated_at, popularity score, relevance ranking) will shift during pagination. Many APIs that appear to be sorted by creation date are actually sorted by internal ID, which looks stable but is not in systems that reuse IDs or batch-assign them.
The correct approach is cursor-based pagination, where the API returns a token you pass on the next request to continue from exactly where you left off. When cursor-based pagination is not available, keyset pagination using a stable, indexed column (like an autoincrement primary key) is preferable to offset-based pagination. Both approaches are immune to the insert-during-pagination problem.
When you are stuck with an offset-based API that has no cursor support, ingest during a low-traffic window where inserts are unlikely, and validate total record count after ingestion against the count reported by the API on the first request.
Anti-pattern three: treating HTTP 200 as success
A common pattern in hand-rolled connectors is to retry on non-200 responses and mark the job successful on 200. This handles the obvious failure case but misses two common API failure modes.
The first is partial results in a 200 response. Some APIs return a 200 with an empty or truncated results array when they are rate-limited, under load, or when the query matches no records. Your connector gets a 200, parses the empty array, and writes zero rows. The job is marked successful. You have silently ingested nothing.
The second is error details inside a 200 body. REST API conventions vary. Some APIs return {"status": "error", "message": "rate limit exceeded"} with an HTTP 200 response because they use the HTTP status code for transport-level success only. Your connector sees 200, parses the body, and finds an error message in a field it was not checking.
The fix is to validate the response body, not just the status code. Check that the parsed record count matches your expectation before marking the job complete. Log the raw response when the count is unexpectedly low so you can debug the source behavior later.
Anti-pattern four: type coercion at write time rather than at source time
Loading data from a source that returns everything as strings, then coercing types as you write to the destination, concentrates all type conversion errors into the write step. When a coercion fails, the entire batch fails. You debug at the wrong level: instead of looking at the source data to understand why a value cannot be coerced, you are looking at the write error and trying to reconstruct what the source sent.
More importantly, silent coercions can produce wrong values. The string "1,234.56" coerced to a float in many locales becomes 1.0 because the comma is interpreted as the end of the integer portion. In other locales, it coerces correctly to 1234.56. Source data that mixes locale conventions, which is common in any multinational system, will have both forms in the same column. Your coercion will silently produce wrong values for half the records.
The correct approach is to validate and coerce at ingestion time, before writing to the destination, and to reject rows that fail coercion rather than silently substituting NULL or zero. Send the rejected rows to a quarantine table with the original raw values and the coercion error. Review the quarantine table as part of your data quality monitoring. The coercion failure rate is itself a useful signal about data quality at the source.
Anti-pattern five: building connector-specific retry logic instead of a shared retry layer
The first connector you build has retry logic. The second connector has slightly different retry logic because a different engineer wrote it and had a different opinion about exponential backoff parameters. By the tenth connector, you have ten different retry implementations with different maximum attempt counts, different backoff curves, different error classification logic for what constitutes a retriable error versus a fatal error, and different logging formats for retry events.
Debugging a partial failure in connector seven requires understanding connector seven's specific retry implementation. Tuning retry behavior requires touching all ten connectors independently. Adding a global circuit breaker requires retrofitting all ten connectors.
The fix is a shared retry layer that all connectors use. Define a standard interface: the connector produces a callable that makes a single attempt. The retry layer wraps the callable and handles retry scheduling, backoff, jitter, circuit breaking, and logging. Each connector focuses on the source-specific logic of constructing the request and parsing the response. Retry policy is a cross-cutting concern that lives in one place.
This matters more as your connector count grows. At 16 connectors, connector-specific retry logic is already a maintenance problem. At 40 connectors, it is a systemic reliability problem. Build the shared layer before you reach the point where retrofitting it is expensive.
The common thread
What ties these five patterns together is that they all optimize for the happy path at the expense of failure handling. Full table replacement is simpler than staging. Offset pagination is simpler than cursor pagination. Checking HTTP 200 is simpler than parsing the response body. Inline type coercion is simpler than a quarantine layer. Per-connector retry is simpler than a shared retry framework.
Every one of these patterns works correctly under normal conditions. The cost appears during partial failures, at scale, and when source behavior changes in ways you did not anticipate. Ingestion infrastructure that handles the happy path well and the failure path badly is the most expensive kind of technical debt, because the failures it produces are the hardest to diagnose.
See cross-source queries in action
Nava Labs connects your databases, warehouses, and event streams. Run SQL across all of them without moving data.
Get Early Access