skip to content

Why is a GraphQL schema usually designed from client demand rather than from database tables?

level: juniorimportance: must knowfreq 61%

answer

  1. Start from the questions, not the tables
  2. The schema is a contract, storage is not
  3. Foreign keys are not relationships
  4. Something has to absorb storage churn
  5. The spec has no opinion here

basics

~20 s

The schema is the contract clients hold, not a view of storage. Designing it from the questions clients actually ask keeps types stable when tables change, and stops the shape of the database leaking into the API.

solid answer

~50 s

Demand-oriented design starts from the operations consuming surfaces need to send, and works backwards to the types that serve them. On a job-board graph, the saved-search results view needs a job posting with its employer, a salary range and whether the viewer already applied — one coherent type assembled from three tables and a pricing service. A schema mirrored from storage instead exposes rows and foreign keys, so every client reassembles the same object by hand and every client makes slightly different mistakes doing it. The deeper reason is coupling: a mirrored schema promotes column names, column types and normalisation decisions into the public contract, so a storage refactor becomes a client-visible change. None of this is in the GraphQL specification — the spec defines a type system and an execution algorithm and says nothing about where types come from. Demand-orientation is design practice, and it is what interviewers are testing.

code

graphql · 17 lines
graphql
type JobPosting {
  id: ID
  employerId: ID
  title: String
  salaryMinCents: Int
  salaryMaxCents: Int
  statusCode: Int
  isDeleted: Boolean
  createdAt: String
}

type Application {
  id: ID
  jobPostingId: ID
  candidateId: ID
  stageId: Int
}

go deeper

for a junior

Be able to state the default and one concrete cost: the schema is a contract, so designing it from the screens keeps clients stable when tables change. Have an example of storage vocabulary leaking, such as an integer status code.

for a middle

Explain the mechanics of the coupling — foreign keys instead of relationships, lookup-table integers instead of enums, and the fact that a mirrored schema leaves no translation layer to absorb a storage refactor.

for a senior

Show judgement about when generation is acceptable anyway: consumer count, surface lifetime and whether the types are published anywhere shared. Be ready to describe migrating an already-mirrored schema without breaking shipped clients.

for a principal

Own the organisational version: who is accountable for the contract when nobody designs it, how vocabulary fragments across teams that each generate from their own database, and where you would allow generation behind an explicit boundary.

## Two ways to arrive at a schema There are only two honest starting points for a GraphQL schema. You can start at the bottom — take the tables, views or documents you already store, and derive a type per table and a field per column. Or you can start at the top — take the surfaces that will consume the graph, write down the operations each one needs to send, and design the smallest set of types that answers all of them. The first is cheap on the day and expensive forever after; the second costs a design conversation up front and buys a contract that survives the storage underneath it. Nothing in the GraphQL specification prefers either. The specification defines a type system, a document syntax, validation rules and an execution algorithm. It has no opinion about where a type comes from, and it never mentions databases. So when an interviewer asks this question, they are not testing spec recall — they are testing whether you understand what a schema *is*: a published contract, and therefore a design artefact with a long life, not a projection of an implementation detail. ## What demand-oriented actually means Take a job-board graph. A candidate's saved-search results view needs, per result: the job title, the employer's display name and logo, a salary range presented as a range rather than two integers, the posting's status, and whether this candidate has already applied. That is one screen's demand, and it describes one type: A `JobPosting` with an `employer` field, a `salaryRange` field, a `status` field, and a `viewerApplication` field that is null when the viewer has not applied. Four of those five come from different places underneath — the employer from a directory service, the salary from two integer columns and a currency, the status from a state machine, the application from a table keyed by candidate and posting. The client never learns that, and does not need to. Now run the second surface through it. The employer's applicant-review view needs a job posting too, but with an applicant count and a hiring stage breakdown. Demand-oriented does not mean a new type; it means adding the two fields the second surface needs to the same concept. The test that you got it right is that the second consumer could be served by *selecting differently* from types that already exist, rather than by adding a new entry point. ## What a mirrored schema costs A schema derived from tables looks superficially similar and behaves very differently. **It exports foreign keys instead of relationships.** Clients receive `employerId` and must issue a second operation to turn it into an employer, which is precisely the round-tripping GraphQL exists to avoid. **It exports storage vocabulary.** `statusCode: Int`, `stageId: Int`, `isDeleted: Boolean`, `createdAt` on every type — enumerations that live in a lookup table arrive as integers, and every client hardcodes the same mapping from `3` to "closed". The knowledge that should have been in the schema, as an enum, ends up copied into every consumer. **It exports decisions that were never meant to be public.** Normalisation, soft-delete flags, audit columns, sharding keys and the compromises of a five-year-old migration all become fields somebody might select. Once selected, they are load-bearing. **It turns storage work into contract work.** This is the expensive one. A column rename, a split into a lookup table, a type widening for a backfill — each regenerates a different schema, and a different schema is a client-visible event. The team that owns the database discovers that they can no longer refactor it without a compatibility conversation, which is exactly backwards: the whole point of an API layer is that it absorbs that churn. **It leaves nowhere for the mapping to live.** In a demand-oriented design there are two shapes — the contract and the storage — and a translation between them that you own and can change. In a mirrored design there is one shape playing both roles, so there is no translation to change, and the only way to absorb a storage change is to break someone. ## The honest caveats Demand-oriented does not mean "whatever the current design mock says" — a schema shaped screen-by-screen has its own failure mode, where every new page adds a root field and a redesign becomes a schema migration. Demand is *evidence* of what concepts the domain has; it is not a specification of them. The craft is reading several surfaces' demands together and extracting the concept they share. It also does not mean generated schemas are never appropriate. A short-lived internal admin surface with one known consumer may be perfectly well served by generation, because the coupling it creates never has time to hurt. The judgement is about lifetime, consumer count and who pays the cost later — and a junior is expected to know the default, which is: design the contract, then map it to storage.

  • Does the GraphQL specification say anything about how a schema should relate to storage?
    No. The specification defines the type system, document syntax, validation and the execution algorithm, and never mentions databases or persistence. Demand-oriented design is engineering practice built on top of it, not a rule you can cite. Saying so plainly is a good sign in an interview, because a lot of what people attribute to 'the spec' on this topic is convention.
  • If your own web client is the only consumer, does mirroring the tables still hurt?
    Less, but the coupling is still there and it runs both ways: the client hardcodes storage vocabulary, and the database team can no longer refactor without a client change. The risk is that 'only consumer' expires quietly — a second surface, a partner, or a mobile build arrives, and by then the storage shape is load-bearing in code you cannot all redeploy at once.
  • How do you gather demand when no client exists yet?
    Write the operations first. Take the designs or the product description, draft the documents you would expect a client to send, and design types that satisfy them — the document is a cheap, throwaway artefact that exposes missing concepts before any resolver exists. Where there is genuinely no consumer to reason about, that is a signal to build less schema, not to fall back on the tables.

A menu is written from what diners want to order, not from the shelf layout of the walk-in fridge. Reorganising the fridge should never reprint the menu.

saying these in an interview costs you the question

  • Says the schema should mirror the database to stay in sync
  • Assumes GraphQL types must map one-to-one to tables
  • Exposes foreign-key ids and expects clients to join
  • Cannot say why storage vocabulary in a contract hurts
  • Claims the specification requires demand-oriented design
  • Thinks generating the schema removes the need to design it

context