Data contracts are the idea that the team producing a dataset should make a formal commitment about its structure and semantics. The team consuming that dataset should be able to rely on that commitment, detect when it breaks, and alert the producer. In practice, most data contract implementations are wikis. A page somewhere documents what the schema is supposed to look like. No one checks.
The gap between "we have data contracts" and "our data contracts catch schema drift before it breaks downstream" comes down to where and how enforcement happens. This article works through what enforcement actually requires and where the common implementations fall short.
Why documentation-only contracts fail
The failure mode is predictable. A team writes a schema specification in Confluence. The data engineering team links to it from the pipeline README. Six months later, a source system adds two new columns and silently renames a third. No one updates the spec. The consuming pipelines start failing with NULL values in a column that is now named differently. The wiki page still shows the old schema because nobody had a reason to update it.
The underlying problem is that a wiki page has no runtime connection to the actual data. There is no mechanism that compares "what we said the schema would be" to "what the schema actually is today." Until you close that gap with something that runs at ingestion time or pipeline execution time, you have aspirational contracts, not enforced ones.
The second failure mode is granularity. Schema specification at the table level (what columns exist, what types they have) catches renames and column additions but misses semantic drift. A column called event_timestamp might change from UTC to Pacific time without any schema change. The type is still TIMESTAMP, the column still exists, but every metric computed from that column is now off by up to eight hours. Table-level schema contracts do not catch this.
What an enforceable contract needs to contain
An enforced data contract has three parts: a machine-readable specification, a validation step that runs against the actual data, and a routing mechanism for failures.
The specification needs to describe more than column names and types. At minimum, it should include nullability constraints (is this column expected to ever be null?), cardinality hints (is this a high-cardinality free-text field or a low-cardinality enum?), value range expectations (should this field ever be negative? should it ever exceed a plausible maximum?), and temporal freshness expectations (how recent should the most recent record be?).
For semantic guarantees that cannot be expressed in schema alone, add documentation that will travel with the schema file: what does this timestamp represent? Is it event time or processing time? What timezone? When is it populated versus null? These annotations do not run as assertions at validation time, but they create a shared reference that makes downstream breakage traceable.
Where to run validation: at write time, not read time
There are two obvious places to put contract validation: at the point where data is written into the system, or at the point where downstream jobs read it. Most teams instrument the read side first because that is where breakage becomes visible. This is backwards.
Read-time validation catches failures after they have already propagated. If a source system starts sending malformed data at 2 AM, the ingestion pipeline writes it, the validation job notices it when the morning dbt run fails, and the data engineering on-call is paging at 8 AM. The corrupt records have been in the table for six hours.
Write-time validation interrupts the ingestion pipeline when incoming data does not match the contract. The options are: reject and fail the batch (good for critical upstream sources where bad data should stop the pipeline), reject and quarantine (route non-conforming records to a reject table while allowing conforming records to proceed), or warn and continue (log the violation but do not interrupt ingestion). The choice depends on how much downstream damage corrupt data can cause and how tolerant the pipeline is to delayed records.
Implementing write-time validation requires that the validation step is in the ingestion code path, not a separate monitoring job. In practice, this means adding a validation call in the ingestion handler before the write, or using a streaming processor that can apply schema rules before committing records to the destination.
Schema registry vs. inline specification
Two common approaches for storing the specification itself: a centralized schema registry (Confluent Schema Registry for Kafka-based pipelines, or a dedicated catalog like DataHub or Amundsen), or inline specification files stored alongside the pipeline code in version control.
The registry approach works well when multiple teams produce and consume the same schemas and you need centralized governance. The registry is the source of truth, pipelines register their schemas and check compatibility before publishing, and consumers know exactly what schema version they are reading.
The inline approach works better for smaller teams with simpler pipeline structures. Keeping the specification as a YAML or JSON file in the same repository as the ingestion code means the spec evolves with the code, schema changes appear in pull requests alongside the code changes that require them, and there is a natural review process.
We use inline specification files for our own ingestion connectors, stored in the same repository as the connector code. Each connector has a schema.json that declares the expected output schema, and validation against that schema runs at ingestion time. When a source system changes its API response shape, the connector's CI pipeline fails because the connector code is no longer producing output that matches the declared schema. The change must update both the connector code and the schema file, and both must pass validation before merging.
Versioning and backward compatibility
Data contracts need versioning. Upstream systems change, and a contract that cannot evolve will either be abandoned or worked around. The challenge is that downstream consumers may not be able to upgrade at the same time the producer changes.
The standard compatibility levels, borrowed from Confluent's schema registry model, are: backward compatible (new schema can read data written with the old schema), forward compatible (old schema can read data written with the new schema), and full compatibility (both). For most analytics pipelines, backward compatibility is the minimum required: new code consuming old data must still work.
In practice, this means: adding new optional columns is safe. Renaming or removing columns is not. Changing a column type from integer to string is not, even if the values remain the same. Any change that requires coordinated updates to all consumers before the producer can deploy is a breaking change and needs a migration plan rather than a version bump.
What data contracts do not solve
Data contracts are not a substitute for data quality monitoring. They catch schema-level violations: the wrong columns, the wrong types, unexpected nulls. They do not catch semantic errors where the schema is technically valid but the values are wrong. A column that should contain user IDs but is now containing order IDs has the right type and cardinality but completely wrong semantics. That requires distribution monitoring and cross-source consistency checks, not schema validation.
Data contracts also do not solve the organizational problem of ownership. A contract is only as good as the producer's commitment to maintain it and the consumer's commitment to surface violations quickly. If the producing team has no accountability mechanism when they break a contract, the contract will degrade from enforcement boundary to documentation over time. The technical enforcement is necessary but not sufficient without organizational alignment around who is responsible for what.
Schema validation built into your connectors
Nava Labs validates schema at ingestion time across all 16 connectors. Violations surface before they reach your downstream tables.
Get Early Access