skip to content

When would you let teams auto-generate GraphQL schemas from their tables rather than hand-design them?

level: principalimportance: should knowfreq 36%

answer

  1. Generation defers the cost, it does not remove it
  2. Consumers, lifetime, and whether it is published
  3. Shared namespaces are the sharp edge
  4. Boundary-based policy, not a blanket ban
  5. Re-evaluate when the conditions expire

basics

~20 s

Allow generation where the coupling never has time to hurt: one known consumer, a short-lived or internal surface, a prototype, or a bridge during a migration. Forbid it wherever types are published into a contract or a shared namespace.

solid answer

~50 s

Generation is not right or wrong; it is a way of deferring the design cost, so the decision is about who pays it later. Ask three things about the surface. How many consumers, and how fast can the slowest one redeploy? How long will it live? And is the schema published anywhere shared — a partner, a mobile build, or a composed graph where these types enter everyone's namespace? A single-consumer internal admin surface scores well on all three and generation is a rational trade. A composed graph does not: generated types carry storage vocabulary into a shared namespace, entity keys become storage keys, and fourteen teams generating from fourteen databases produce three names for the same concept. The policy I would write is boundary-based rather than blanket: generate freely behind an internal boundary, hand-design anything published, and never let a generated type cross from one to the other unrenamed.

code

graphql · 15 lines
graphql
# subgraph published by the hiring team
type Job {
  id: ID!
  title: String!
  employerId: ID!
  statusCode: Int!
}

# subgraph published by the search team
type Posting {
  postingId: ID!
  headline: String!
  companyId: ID!
  isActive: Boolean!
}

go deeper

for a junior

You will not be asked to set this policy, but know the premise: generating a schema from tables is fast and makes the storage shape public, so it suits throwaway internal surfaces far better than a lasting contract.

for a middle

Be able to name the concrete costs a generated schema pushes onto consumers — reassembling objects, hardcoding lookup values, and absorbing storage refactors as breaking changes.

for a senior

Show that you evaluate the surface, not the technique: consumer count, upgrade latency of the slowest client, expected lifetime, and whether the types are published anywhere you cannot redeploy alongside.

for a principal

Own the written rule and its exit. Argue the boundary between generated internals and a designed contract, who governs shared vocabulary across teams, and what evidence triggers re-evaluating a generation decision that has outlived its conditions.

## Reframe the question before answering it The interviewer is not asking which technique is better. They are asking whether you can price a trade that a team is already making, usually under delivery pressure, and turn it into a rule other people can apply without you in the room. Generation does not remove the design work; it defers it and moves who pays. On the day, one team saves weeks. Afterwards the cost reappears in three places: every consumer reassembles rows into the objects it actually wanted, every storage refactor becomes a compatibility event, and nobody owns the contract because nobody wrote it. A policy that says "never generate" ignores the real saving; a policy that says "generate everything" ignores the bill. A principal writes the boundary between them. ## The three questions that decide it **How many consumers, and what is the slowest one's upgrade latency?** One consumer in a repository you control means a breaking change is a same-day edit in the same change. A shipped native build means the coupling is permanent for the life of that build. The count matters less than the tail. **How long will the surface live?** An internal reconciliation tool for a three-month data migration accrues almost no coupling debt before it is deleted. A candidate-facing graph will outlive the current storage design, probably twice. **Is it published into anything shared?** This is the sharpest of the three, and it is where teams get surprised. In a composed graph, a subgraph's types enter a namespace everyone else reads. Generated types bring storage vocabulary with them — lookup-table integers, audit columns, soft-delete flags — and once composed they are as public as anything hand-designed. Worse, the identity used to join across services becomes whatever the table's primary key happens to be, so a storage decision quietly becomes a cross-team contract. ## What it looks like at scale Concretely: an organisation running a composed graph of 62 subgraphs across 14 product teams. Three teams generate from their databases because it shipped in two days rather than three weeks, and nobody said not to. The first symptom is vocabulary. One team publishes `Job`, another `Posting`, a third `Requisition`, and they are the same concept with different column names underneath. Consumers cannot tell, so they learn all three, and every new engineer asks which one to use. Composition does not stop this — the names do not collide, so nothing fails; the graph simply gets harder to understand, permanently. The second symptom is churn. A storage refactor in one subgraph now moves a published type, which means a composition check, a client migration, or both. Traced over a quarter, the majority of contract-breaking edits come from the minority of subgraphs that generate — which is the number worth putting in front of leadership, because it converts a taste argument into evidence. The third symptom is ownership. When nobody wrote the schema, nobody reviews it, and there is nothing for a design conversation to be about. The team that generated it will say, honestly, that they did not choose these fields. ## The policy I would actually write Boundary-based, not blanket: * **Generate freely behind an internal boundary.** Prototypes, single-consumer admin and analytics surfaces, migration bridges, anything with a stated deletion date. Require the deletion date to be written down, because "temporary" surfaces are the ones that survive. * **Hand-design anything published** — partner-facing, mobile-facing, or composed into a shared graph. Published means somebody who cannot redeploy with you depends on it. * **Never let a generated type cross the boundary unrenamed.** A generated internal layer with a hand-designed contract in front of it is a perfectly good architecture; the mistake is exporting the generated one directly. * **Set the vocabulary rules centrally.** In a composed graph, the concepts and their identities are a shared asset. Which types exist and what identifies them is not a per-team decision, whether or not a generator is involved. * **Name the review gate and the exception path.** Who signs off on publishing a surface, and who may waive the rule, on what evidence. ## The honest counter-argument, and the exit The strongest case against this is speed, and it is a real case: for a team with one consumer and a deadline, generation genuinely is the correct engineering call, and a policy that forbids it makes the platform team the obstacle. Say that plainly in an interview — a principal who cannot argue the other side is not evaluating a trade, they are enforcing a preference. What makes it safe is the exit. Generation should be reversible: keep the generated surface behind a boundary you can put a hand-designed contract in front of, hold usage evidence per field so you know what the migration would actually cost, and re-evaluate when any of the three inputs change — a second consumer appears, the deletion date passes, or someone proposes composing it into the shared graph. Most of the damage in real organisations comes not from choosing generation but from never revisiting the choice once the conditions that justified it have quietly expired.

  • How would you migrate a table-mirrored schema that is already published to external consumers?
    Incrementally and per surface. Gather usage evidence per field and per caller, add hand-designed concept fields alongside the generated ones, move consumers one at a time, then deprecate the mirrored fields starting with the least used. The external tail sets the pace, so expect the mirrored fields to outlive the migration by a long window; what matters is that nothing new is built on them.
  • What signals tell you a generated schema is now costing more than it saved?
    Contract-breaking edits traced back to storage refactors rather than product decisions; consumers shipping their own reassembly helpers for the same objects; support questions about which of several near-identical types to use; and fields in the published schema that nobody chose to expose. Any two of those together mean the conditions that justified generation have expired.
  • Does composing subgraphs into one graph change the calculus?
    Substantially. A published subgraph's types enter a namespace every other team reads, so generated storage vocabulary becomes shared vocabulary, and the identity used to join across services becomes whatever the primary key happens to be. Nothing fails at composition — the names simply diverge — which is why this needs a policy rather than a build check.

saying these in an interview costs you the question

  • Treats generation as always wrong or always right
  • Publishes generated types into a shared graph namespace
  • Cannot say who owns a contract nobody designed
  • Judges only by delivery speed, ignoring consumer cost
  • Assumes generation removes the storage mapping entirely
  • Never revisits the decision after the conditions change

context