skip to content

Why does a GraphQL schema generated from database tables turn storage refactors into client breakages?

level: middleimportance: must knowfreq 53%

answer

  1. The column is the field
  2. Regeneration is a publication, not a build step
  3. Two shapes need something in between
  4. Rename, split, widen for backfill
  5. A subscription only learns at reconnect

basics

~20 s

Because the generated field names, types and nullability are derived from columns, so renaming, splitting or widening a column produces a different published contract. With no translation layer to absorb the change, storage edits reach clients directly.

solid answer

~50 s

In a generated schema the column *is* the field: its name, its type and its nullability come straight from the table definition, and regeneration republishes whatever the table now says. So a rename removes a field, a split into a lookup table removes one field and adds another, and making a column nullable for a backfill widens the field's type — every one of which is a change a client can notice. A hand-designed schema has two shapes, the contract and the storage, plus a mapping between them that you own; a storage refactor becomes an edit to the mapping and the contract does not move. Generation collapses those into one shape, so there is nothing left to absorb the churn. The practical consequence is that the database team can no longer refactor without a compatibility conversation, which inverts the reason the API layer exists.

code

json · 8 lines
json
{
  "errors": [
    {
      "message": "Cannot query field \"status\" on type \"Application\".",
      "locations": [{ "line": 4, "column": 7 }]
    }
  ]
}

go deeper

for a junior

Understand the core sentence: in a generated schema the column is the field, so changing the column changes what clients are promised. One concrete example, such as a rename removing a field, is enough at this level.

for a middle

Be able to walk through several refactor classes — rename, split, widen for a backfill, denormalise — and explain why a hand-designed schema plus a mapping layer absorbs all of them while a generated one publishes them.

for a senior

Bring the operational reading: which consumers find out at deploy and which at reconnect, how a quiet client turns a validation failure into silent staleness, and what you would monitor to catch it before a user does.

for a principal

Argue where the cost belongs. A mapping layer concentrates the work in one team; generation distributes it to every consumer as coordinated releases. Be ready to defend that trade in front of a team that finds generation dramatically faster today.

## One shape doing two jobs A published GraphQL schema and a database schema answer different questions. The database schema answers "how do we store this efficiently and consistently?" — and the right answers change as volume, access patterns and correctness requirements change. The published schema answers "what may a client ask for, and what will it get back?" — and the right answer there is supposed to stay still, because other people's code depends on it. Generating one from the other collapses two shapes into one. The field name is the column name; the field type is derived from the column type; the field's nullability is derived from the column's. Regeneration is not a build step, it is a *publication*: whatever the table says today becomes what clients are told today. That is the whole mechanism, and it is worth stating that plainly in an interview before listing examples, because the examples are only interesting once the mechanism is clear. ## The refactors that leak Using a job-board graph, with `applications` as the table underneath an `Application` type: * **Rename.** `status` becomes `state` for clarity. Regeneration removes a field and adds another. Every document selecting `status` now fails validation. * **Split into a lookup table.** The status string is normalised into `application_states` with a foreign key. The scalar field disappears; something object-shaped or integer-shaped appears in its place. * **Widening for a backfill.** A column is made nullable so a migration can run in stages. The generated field goes from Non-Null to nullable, and every client that assumed a value now has to handle a hole it never planned for. * **Denormalisation.** A counter is cached on the row for performance. A field appears that nobody designed, and a client starts selecting it, and now the cache is public API. * **Moving storage.** The table moves to another service or another store entirely. There was never a contract to preserve — only a table — so the port is a rewrite of the surface. None of these are database mistakes. They are ordinary, correct storage engineering, and a mirrored schema converts each into an API event. ## A worked failure An employer dashboard on the job-board graph opens a subscription for application status changes and paints a live board of candidates by stage. The storage team normalises the status column into a lookup table — a clean, well-reviewed migration with a backfill. The schema regenerates: `Application.status` is gone, `Application.stateId` has appeared. On the next connect, the dashboard's subscription document is rejected at validation with `Cannot query field "status" on type "Application"`. The dashboard, like a great many clients, treats a failed subscription as a transient network problem: it renders whatever it last had and retries quietly. Nothing turns red. Recruiters keep looking at a board that has stopped moving, and the gap is noticed 41 hours later when a candidate phones about an interview the board still shows as unscheduled. Two lessons sit in that story. The obvious one is the coupling. The subtler one is that a schema change reaches a subscription at *reconnect*, not at deploy, so the blast radius of a regenerated schema is spread across hours and looks like flakiness rather than a release. ## Where the mapping should live The fix is not "never rename a column" — that just pins the database to API vocabulary and makes the coupling bite from the other side, so the storage team inherits the API's compatibility rules. The fix is to have somewhere for the difference to live. In a hand-designed schema there are three artefacts: the published contract, the storage shape, and a mapping layer between them that belongs to the API team and is not published to anyone. When the status column splits into a lookup table, the mapping learns to read the new place, translate the row into the existing `JobApplicationStatus` enum, and hand back exactly what the contract always said. Clients see nothing. When the backfill needs the column nullable, the mapping supplies a defined value for the gap rather than passing a hole through to the contract. When the table moves to another service, the mapping changes shape and the contract does not. That layer is not free — it is code, it needs tests, and it has to be maintained as the storage drifts. But its cost is paid by one team, once, in a place they control, instead of being distributed to every consumer as a coordinated release. ## What is safe, even with generation Be fair about the other side. Adding a column is usually additive and harmless to existing documents, which is why generated schemas feel fine for months: most days you add, and adding does not break. The trouble is that generation also publishes every addition — audit columns, soft-delete flags, internal scoring — so the schema accretes fields nobody designed and now cannot remove, because a client somewhere may have found them. The failure mode is not only the loud break at a rename; it is the quiet growth of a contract nobody ever agreed to.

  • Why is freezing the column names to protect the schema a bad fix?
    It moves the coupling rather than removing it. The database inherits the API's compatibility rules, so a name that turned out to be wrong, a normalisation that would fix a bug, or a move to another store all become contract negotiations. You end up with a storage layer nobody may improve, which costs more over a few years than the mapping layer you were avoiding.
  • Which storage changes stay invisible to clients once a mapping layer exists?
    Renames, splits into lookup tables, denormalisation, index and partition changes, a column widened for a staged backfill, and moving the data to another service entirely. The mapping changes; the published fields do not. What a mapping layer cannot hide is a genuine semantic change — if the concept the field names no longer exists, no translation saves you and it becomes a real contract decision.
  • Is adding a column to a table with a generated schema harmless?
    Usually harmless to existing documents, since additions do not break selections. But it publishes fields nobody designed — audit columns, soft-delete flags, internal scores — and once a client selects one, removing it is a breaking change. The slow damage from generation is not the loud rename; it is a contract that grows fields no one agreed to expose.

saying these in an interview costs you the question

  • Assumes a generated schema stays compatible because it is automatic
  • Treats a column rename as purely a database concern
  • Proposes freezing column names as the permanent fix
  • Cannot name a storage change that is safe for the DB but breaking for clients
  • Thinks nullability changes are invisible to clients
  • Believes a redeploy is when clients notice, ignoring reconnecting consumers

context