What does a data contract specify beyond the column names and types?
answer
- more than a table's DDL
- types cannot say what a number means
- who owns it, and for how long
- schema, semantics, SLA, change policy
basics
~20 sA data contract pins down the schema plus everything a schema cannot say: field meanings and units, the key and grain, freshness and volume expectations, a named owning team, a version, and how breaking changes get announced.
solid answer
~40 sA data contract is the explicit, versioned agreement between a producing team and everyone reading its dataset. It has roughly four parts. **Schema**: field names, types, nullability, and the key plus the grain (what one row represents). **Semantics**: what the values actually mean — is `amount` in integer cents or dollars, is `placed_at` event time or ingest time, can a row be superseded by a correction. **Service levels**: maximum staleness, an expected volume range, availability. **Change policy**: the version scheme, what counts as a compatible versus breaking change, the deprecation window, and who must sign off. It also names an owner — a team, not "the data platform" — so there is somebody accountable when any of the above is violated.
code
yaml · 24 linesdataset: orders.order_placed
version: 2.1.0
owner: team-checkout
grain: one row per placed order
key: [order_id]
fields:
- name: order_id
type: string
required: true
description: Stable across producer retries; safe as a dedupe key
- name: placed_at
type: timestamp
required: true
description: Event time in UTC, not ingest time
- name: gross_amount_cents
type: integer
required: true
description: Integer cents, tax included, shipping excluded
sla:
max_staleness: 15m
expected_daily_rows_min: 20000
change_policy:
breaking_change_requires: consumer_signoff
deprecation_window: 90dgo deeper
Be able to say what a data contract is in one sentence — an explicit, versioned agreement about a dataset between the team producing it and the teams reading it — and name at least schema, meaning and freshness as its parts.
Explain why each part exists: what a grain declaration prevents, why units and event-versus-ingest time belong in writing, and why a version plus a change policy is what separates a contract from a schema dump.
Show that you place the contract where change happens — the producer's repo and CI — and that you can argue which fields deserve a contract at all rather than proposing one for every table in the warehouse.
Own the question of scope and cost: how many datasets justify contracts, what an organisation actually commits to when it signs one, and how ownership is assigned so no dataset ends up owned by the platform team by default.
## The problem a data contract solves Every analytics pipeline depends on tables and event streams produced by teams that did not write the pipeline. A backend engineer renames a column on Tuesday because it reads better in their service; on Wednesday a dashboard is silently wrong and nobody connects the two. The dependency was real, but it was invisible: nothing in the producing team's repository said that seven downstream models select `coupon_code`. A data contract makes that dependency explicit, versioned and reviewable. It is a specification of the dataset as an interface — the same idea as an API contract for a service, applied to a table or an event topic. ## Part one: schema The obvious part. Field names, types, nullability, allowed values for enumerated fields. Two things matter here that raw DDL usually omits: - **The key** — which field or combination uniquely identifies a row, and whether the producer guarantees it is stable across retries. - **The grain** — what one row *is*: one row per order, per order line, per shipment. Grain is the single most consequential fact about a dataset and the one most often left unwritten. ## Part two: semantics Types do not carry meaning. `integer` does not tell you that `gross_amount_cents` is cents rather than dollars, includes tax, and excludes shipping. `timestamp` does not tell you whether `placed_at` is the moment the customer clicked (event time) or the moment the row landed (ingest time), nor whether it is UTC. Nothing in the type system says whether a cancelled order is deleted, updated in place, or emitted as a second row that supersedes the first. These are exactly the details that produce quietly wrong numbers rather than loud failures, so a contract states them in prose next to each field. If your organisation has a metric definition — "active user", "net revenue" — the contract is where the producer commits to the definition the field implements. ## Part three: service levels A correctly-typed but empty table is still an outage. Contracts therefore carry operational commitments: a **freshness** target (maximum lag between when something happened and when it is queryable), an **expected volume** range so a partial or empty load is detectable, and often an availability or completeness expectation. These are the clauses that turn a contract from documentation into something you can alert on. ## Part four: change policy The contract says how it may change. That means a version scheme, an explicit list of what counts as a compatible change (adding an optional field) versus a breaking one (removing a field, narrowing a type, changing units or grain), a **deprecation window** during which a field is marked dead but still emitted, the channel on which change is announced, and who has to approve a breaking change. Without this section a contract is just a snapshot of today's schema. ## Ownership and where it lives A contract with no named owner is a wish. Ownership sits with the producing team, because they are the only ones who can prevent a breaking change rather than detect it after the fact. Practically, the spec file lives in the **producer's** repository next to the code that emits the data, is reviewed in their pull requests, and is published to a catalog or registry when it merges. That placement is what makes the contract enforceable in the producer's own CI instead of being a document the data team maintains alone. ## What a data contract is not It is not a serialization compatibility setting — a registry's compatibility mode is a mechanical check on encoded payloads and is a neighbouring concern, not the whole agreement. It is not a set of pipeline tests that run after the data has landed; those detect a violation, they do not prevent one, and by then the bad rows exist. And it is not a document written by the consuming team and shown to producers hopefully — a contract only functions when the producer has agreed to it and is gated by it. ## How contracts fail in practice The common failure modes are all social. A contract written for every table on day one, most of which nobody reads, dies of maintenance. A contract with sixty fields and no owner is unreviewable. And a contract nothing enforces is a comment. The version that works starts with one high-value dataset, is small enough that a producer will actually read the diff, and is checked automatically.
- Why does the contract file live in the producing team's repository rather than the data platform's?Because that is where change originates. If the spec sits next to the code that emits the data, a rename and the contract diff appear in the same pull request and the producer's own CI can block it. A contract kept in the platform repo can only be updated after the fact, which turns it into documentation of breakage rather than prevention of it.
- What does declaring the grain in a contract prevent that a type declaration cannot?Double counting. If a consumer believes the table is one row per order and the producer starts emitting one row per order line, every sum of revenue inflates while every type and null check still passes. Grain is the assumption every aggregate silently depends on, so stating it makes a grain change a visible, breaking contract change.
It is an OpenAPI spec for a table: the schema is the request shape, the semantics are the field docs, the SLA is the uptime promise, and the change policy is the versioning rule that stops a rename from being a silent outage.
saying these in an interview costs you the question
- Calling a data contract just the table's DDL written down
- Assuming the data team writes and owns the producer's contract
- Leaving grain and units unstated because the types look obvious
- Treating a data contract as documentation with no enforcement
- Believing a schema check alone proves the data is usable