skip to content

What does generating DDL, tests and docs from a data contract spec buy you?

level: middleimportance: nice to knowfreq 26%

answer

  1. one file, many downstream outputs
  2. stops the doc drifting from reality
  3. the only way to change it is to change the spec
  4. DDL, checks and docs from one source

basics

~20 s

Generation makes the contract load-bearing instead of descriptive. If the target DDL, the validation checks and the catalog entry are all produced from one spec, the documented contract cannot drift from the enforced one, because there is only one source.

solid answer

~50 s

A contract that is merely *compared* against reality can drift: someone edits the DDL, forgets the YAML, and the spec quietly becomes fiction. If the spec is instead the **source** — the target table's DDL, the serialization schema for the event, the not-null and range checks that run on landed data, the typed producer or consumer stub, and the catalog documentation are all generated from it — then divergence is structurally impossible and the only way to change the interface is to change the contract, which is exactly the file the CI gate watches. The costs are real: your spec format's expressiveness becomes a ceiling on what teams can model, escape hatches accumulate for the cases it cannot express, and the generator is now bespoke platform code somebody must own. Worth it for a stable core of high-traffic datasets, overkill for a long tail.

code

yaml · 10 lines
yaml
dataset: orders.order_placed
fields:
  - name: order_id
    type: string
    required: true
    description: Stable across producer retries
  - name: status
    type: string
    required: true
    allowed: [placed, cancelled, fulfilled]

go deeper

for a junior

Know the idea: if the table definition and its documentation are produced from the data contract file, they cannot disagree with it, because nobody writes them by hand.

for a middle

Explain the inversion — comparison detects drift after the fact, generation makes it impossible — and list what is typically generated: target DDL, data checks, typed stubs, catalog docs.

for a senior

Argue the tradeoff rather than the feature: which datasets justify a generator, where an escape hatch reopens the drift you closed, and why generating desired state is not the same as migrating an existing large table.

for a principal

Own the build-versus-buy call and the standing cost. A homegrown generator is platform code with an owner, a migration burden when the warehouse changes, and lock-in for every consumer that depends on its output.

## Descriptive contracts drift The weakest version of a data contract is a YAML file that describes a table. Nothing forces the description to stay true. An engineer alters the emitted structure, does not touch the spec, and the contract silently becomes a historical document — which is worse than no contract, because consumers now trust something false. The standard first fix is a comparison: CI asserts that the produced structure matches the spec. That works, but it only catches drift after someone creates it, and it needs the comparison to be implemented for every artefact you care about. ## Generation inverts the dependency The stronger version makes the contract the single source of truth and derives everything from it: - **Target DDL** — the `CREATE TABLE` in the warehouse, including column comments taken from the field descriptions. - **The wire schema** — the serialization schema for an event topic, produced from the same field list (compatibility mode itself is a registry concern next door, but the schema body comes from here). - **Data checks** — not-null, allowed-value and range assertions derived from the declared constraints, plus the freshness and volume probes derived from the SLA clause, wired into whatever runs them on landed data. - **Typed stubs** — a producer class or consumer type so an application that emits or reads the dataset fails to compile when the contract changes. - **Documentation** — the catalog entry, the field glossary, the owner and the SLA, published automatically on merge so what consumers read is always the current version. Once these are generated, drift is not detected — it is impossible. Changing the interface *requires* editing the contract, which is the file the producer's CI gate diffs and classifies. Governance and implementation stop being two things that can disagree. ## A second benefit: the description becomes real Generation also fixes the perennial documentation problem. When a field description is only prose in a wiki, it rots. When it is generated into the warehouse as a column comment and into the catalog page in the same commit, the incentive flips: the description ships with the change or the change does not ship. ## The costs, which an interviewer will probe Be ready with the downsides, because uncritical enthusiasm for codegen is itself a red flag. **The spec becomes a ceiling.** Your generator can only produce what the format expresses. The first team that needs a clustering key, a masking policy, a nested type your format does not model, or an unusual partitioning scheme hits a wall. Either the format grows every quarter, or escape hatches appear — a raw DDL block pasted into the spec — and the guarantee erodes. **The generator is bespoke platform code.** Somebody owns it, versions it, tests it, and handles the migration when the warehouse changes its syntax. That is a standing cost, and on a small team it is often larger than the drift it prevents. **Regeneration is not migration.** Generating a `CREATE TABLE` is easy; turning a contract change into a safe `ALTER` on a table holding a billion rows is a different problem, with backfills and locks. Teams usually generate the *desired* state and still hand-review the migration plan. **Tooling lock-in.** Consumers start depending on generated stubs, and swapping the generator later means touching every consumer. ## Where the tradeoff lands Generation pays off on a small set of high-traffic, long-lived, cross-team datasets where drift is expensive and the shapes are conventional. For the long tail — one team's exploratory tables, short-lived datasets, anything with unusual physical requirements — comparison in CI gives most of the safety for a fraction of the platform investment. Saying that out loud, rather than proposing to generate everything, is what distinguishes a considered answer from a fashionable one.

  • What is the main risk of generating everything from a contract spec?
    The spec's expressiveness becomes the platform's ceiling. The first dataset needing a clustering key, a masking policy or an unusual nested type either forces the format to grow or gets an escape hatch that quietly reopens the drift you were preventing. The generator is also bespoke code somebody must own indefinitely.
  • Does generating a CREATE TABLE from the contract solve schema migration too?
    No. Generation produces the desired state; getting an existing large table from the old state to the new one is a migration problem with backfills, locks and cost. Most teams generate the target DDL and still review the ALTER plan by hand, especially for anything that rewrites data.

saying these in an interview costs you the question

  • Assuming generated artifacts remove the need for producer sign-off
  • Proposing to generate every dataset's DDL from day one
  • Ignoring that the generator itself needs an owner and tests
  • Believing generated CREATE TABLE handles migrating existing data

context