Building a connector that works in a demo is straightforward. Building a connector that works reliably at 2 AM on a Tuesday in November, when the source API is returning 429s from a Black Friday traffic spike and your ingestion pipeline has been silently dropping records for the past 90 minutes, requires a different set of design decisions.
Connector reliability is the layer between source system variability and pipeline stability. This article covers the specific patterns that make the difference between connectors that require babysitting and connectors that run for months without operator intervention.
Rate limiting: respect the limit before you hit it
Every production API has rate limits. The naive approach is to make requests as fast as possible, handle 429 responses when they arrive, wait the indicated retry-after period, and continue. This works in development but creates problems in production.
First, 429 responses are not free. Some APIs count 429 responses against the rate limit quota. If your connector hammers the API until it gets 429s and then backs off, it may have consumed a significant fraction of the rate limit just on unsuccessful requests.
Second, concurrent connectors share a rate limit. If your ingestion pipeline runs three concurrent connector instances against the same API (perhaps ingesting from three different endpoints), each instance making requests as fast as possible without awareness of the others, they will collectively hit the rate limit much faster than any one would alone. The result is burst-then-throttle behavior that produces uneven ingestion throughput.
The better pattern is proactive rate limiting. Implement a token bucket or leaky bucket algorithm at the connector level that enforces a maximum request rate below the API's stated limit, leaving headroom for variability. For multi-instance connectors, implement a shared rate limit counter (in Redis or a similar shared store) so all instances collectively stay under the limit rather than each independently approaching it.
When a 429 does arrive, the connector should parse the Retry-After header and wait exactly that long, not a fixed sleep. APIs that set Retry-After are telling you the correct wait time; ignoring it and using an exponential backoff that waits too long wastes time, and waiting too short may trigger another 429 immediately.
API versioning and forward compatibility
Source APIs change. A connector that was written for v2 of an API will at some point face a world where v2 is deprecated and v3 has different field names, different pagination patterns, or different authentication requirements. How gracefully the connector handles this determines how much maintenance the connector requires over its lifetime.
The basic reliability pattern is explicit version pinning combined with version lifecycle monitoring. Pin your connector to a specific API version, not the current default. Register for deprecation notifications from the API provider if they offer them. Build in a mechanism to alert when the connector's pinned version is approaching end-of-life.
The harder problem is additive changes: fields that appear in the response without a version bump, or fields that change from required to optional. A strictly typed schema mapping that rejects any field not in the expected set will break when the API adds new fields. A schema mapping that accepts any field and maps only the ones it knows about is forward-compatible by default. For most analytics connectors, the second approach is safer: unknown new fields are ignored, which is preferable to the connector failing entirely.
When handling field changes, log what the connector did not recognize. Silent ignoring of unknown fields is fine for reliability, but you want a record of what was ignored so you can decide whether those fields contain data your analytics queries should eventually use.
Partial response bodies and incomplete pagination
HTTP 200 does not mean the response is complete. This is one of the most reliable ways connectors fail silently in production. The server responded, the status code was 200, the connector processed the response body, and moved on. What was in the response body might have been a truncated JSON object, a paginated response where the next_page token was missing, or a response where the server returned fewer records than the page size with no indication that more are available.
For JSON responses, validate that the response parses completely before processing. A truncated response body will often parse with missing fields or null values rather than failing to parse entirely, depending on where the truncation occurred. Structural validation (are all expected top-level keys present? is the data array present and non-empty when it should be?) catches truncation that JSON parsing alone does not.
For paginated responses, verify that your pagination logic handles the terminal condition correctly. The most common failure is a connector that stops paginating when it receives an empty next_page token, but the API returns a null next_page field rather than an absent one on the last page. The connector reads null, evaluates it as truthy (depending on language and null-handling code), and requests page 2 indefinitely. Or the opposite: the API always returns a next_page token even on the last page, and the connector follows it into an empty response that it then treats as a pagination error.
Test pagination terminal conditions explicitly against the API's actual behavior, not against the documentation. API documentation and API behavior often disagree on edge cases.
Idempotent writes and the cursor position problem
Connector reliability is not only about reading correctly. It also requires that a retry after a failed write does not produce duplicate records in the destination.
The standard pattern for idempotent writes is to use an upsert (insert-or-update on a primary key) rather than a plain insert. If the connector writes a batch, fails before checkpointing its cursor position, and retries the same batch, the upsert ensures that re-written records overwrite the previous write rather than creating duplicates.
This requires that every record has a stable primary key. For APIs that return records with a native ID field, this is straightforward. For APIs that return records without stable identifiers (some event streams), you need to derive a key from the record content: a hash of the source system identifier, the event type, and the event timestamp is a common pattern. The hash is stable across retries and provides a write-safe key.
Cursor position should be checkpointed after a successful write, not after a successful read. The sequence is: read a page of records, write all records from that page, checkpoint the cursor to indicate that page is complete. If the write fails, the cursor has not advanced, and the retry reads and writes the same page. This is why idempotent writes are essential: the retry path goes through the write on the same records.
Alerting on silence, not just failure
The hardest connector failures to detect are silent failures. The connector is running, it is not throwing exceptions, but it has stopped ingesting data. The source API changed its authentication flow and is now returning empty results instead of an error. The rate limiter is holding the connector in a permanent backoff loop. The cursor position is stuck at a date six weeks ago because of a timezone bug.
Exception-based alerting catches the noisy failures. Silent failures require a separate layer: monitoring the rate of records ingested over time and alerting when that rate falls below a threshold that indicates the connector is no longer making forward progress.
For each connector, establish a baseline expected ingestion rate. A CRM connector might normally ingest 50-500 records per sync. An event stream connector might ingest thousands per minute. Set an alert threshold below which the connector is considered unhealthy: if the CRM connector syncs 0 records for 3 consecutive sync cycles, alert even if no exceptions were thrown. This catches the class of failures that exception monitoring misses entirely.
We are not saying every connector failure is preventable. Source APIs have outages, authentication tokens expire without warning, and new API behavior occasionally breaks assumptions. The goal of these patterns is to fail loudly and quickly when failure is unavoidable, and to reduce the category of failures that go undetected for hours or days.
16 connectors built for production reliability
Nava Labs handles rate limiting, pagination, and idempotent writes across all connectors. Silent failures alert, not silently drop.
Get Early Access