skip to content

How do you decide which GraphQL operations belong behind an edge cache and which stay at the origin?

level: principalimportance: should knowfreq 44%

answer

  1. Per operation, never a global switch
  2. Cost times rate, not hit rate
  3. Count the distinct keys first
  4. A hit means no resolver ran
  5. How long can a wrong entry live

basics

~20 s

Per operation, never as a global switch. Promote reads whose origin cost times request rate is large, whose variables have low cardinality, that carry no viewer-scoped field, and whose staleness budget you can state in seconds.

solid answer

~50 s

Treat the edge as a portfolio decision over individual operations. The prize is origin cost saved multiplied by request rate, so a deep, expensive, high-volume public read is worth promoting and a cheap one at two requests per second is not, whatever its hit rate. Then apply the disqualifiers: variables with near-unbounded values never accumulate hits; any viewer-scoped field in the selection forces either an identity dimension in the key or a split of the document; and an operation whose staleness budget you cannot state in seconds is not ready to be cached. Weigh what a hit costs you as well as saves you — no resolver runs, so per-field authorization, field-usage telemetry and per-request logging cover only miss traffic. Finally, size the blast radius: a wrong entry is served to everyone until it expires, which argues for short lifetimes over an invalidation pipeline you would have to build and operate.

code

pseudocode · 12 lines
pseudocode
function belongs_at_edge(op):
    if op.type != QUERY:
        return false
    if any_field_is_viewer_scoped(op.selection):
        return false
    if staleness_budget_seconds(op) is unknown:
        return false
    if distinct_variable_values(op, window = "7d") > 5000:
        return false

    saved_ms_per_day = origin_cost_ms(op) * requests_per_day(op) * expected_hit_rate(op)
    return saved_ms_per_day > 600000

go deeper

for a junior

Understand the basic trade: a cached response is fast and cheap but may be out of date, and only data that is the same for every caller can be shared safely. Knowing which reads are public and which are personal is the starting point.

for a middle

Be able to reason about hit rate concretely — how many distinct keys the variables produce, how often each is requested within one lifetime — and to spot the viewer-scoped field that disqualifies an operation from a shared cache.

for a senior

An interviewer expects you to weigh what a hit removes as well as what it saves: no resolver execution means no per-field authorization, thinner logs and incomplete field-usage data, plus a bounded but real staleness tail.

for a principal

Own the policy and the review loop: default to origin, promote operations individually against stated evidence, keep lifetimes short instead of building invalidation early, and make sure a schema change cannot silently invalidate an edge rule owned by another team.

## Why this is a per-operation decision "Should we put a CDN in front of the GraphQL API?" is the wrong question, because a GraphQL endpoint is not one thing. It is every read in the schema arriving at one URL, with wildly different costs, cardinalities, viewer-dependence and tolerance for staleness. A global answer is either so conservative that it buys nothing or so aggressive that it leaks. The unit of decision is the operation, and the output is a short list — usually a handful of reads out of hundreds — that earns edge treatment while everything else goes to the origin by default. That framing matters organizationally as much as technically. A per-operation list is reviewable, has an owner, and can be argued about; "caching is on" cannot. ## The prize: cost saved times rate Start with what promotion is worth, which is origin work removed, not hit rate. A music catalogue's `ReleaseDetail` read is the canonical candidate: a 19-level-deep document that walks releases into tracks into credits into contributing artists into their related releases, costing 812 ms at the ninety-fifth percentile and accounting for roughly a third of read traffic. Saving most of that is a real capacity story. The counter-example matters just as much. A read that costs 14 ms and runs twice a second may cache beautifully and still not be worth a single line of edge configuration, because you have added a component, a failure mode and a reasoning burden in exchange for nothing measurable. Cheap and rare stays at the origin however cacheable it looks. ## The disqualifiers **Variable cardinality.** Hits accumulate only when the same key is requested repeatedly within one lifetime. A read parameterised by a free-text search string has effectively unbounded keys and will approach a zero hit rate while filling the cache with entries nobody asks for twice. A read keyed by release id over a catalogue of 240,000 releases sounds equally hopeless until you look at the access distribution: if a few thousand releases take most of the traffic, the head caches and the tail simply misses, which is fine. Estimate the distinct-key count over a window before promoting anything. **Viewer dependence.** Any field in the selection whose value depends on who is asking either forces an identity dimension into the key, collapsing the hit rate toward one entry per viewer, or forces the document to be split so the shared half carries nothing personal. Neither is free, and choosing between them is part of the promotion decision rather than an implementation detail after it. **A staleness budget you can state.** If nobody can say how many seconds out of date this response may be, it is not ready. Catalogue metadata tolerates minutes. A live listener count does not tolerate seconds. The number is a product decision, and it should be written next to the operation. ## What a hit costs you This is the part experienced candidates raise and everyone else forgets: a cache hit means the server never executed the operation. Everything the server would have done, it did not do. Per-field authorization checks live in resolvers, so on hit traffic they do not run — which is exactly why viewer-dependent selections cannot be at the edge, and why the check has to be the *key's* job instead. Per-request logging covers only misses. And field-usage telemetry, the data on which deprecation decisions rest, now describes miss traffic alone. That last one produced the catalogue's most instructive incident. `Track.previewUrl` showed almost no usage in the schema's field-usage data and was duly deprecated and removed — but its traffic came almost entirely from one high-volume operation served from the edge, so the usage numbers had never seen it. A client pinned to a document still selecting that field then began failing validation on every request. The defect was not the removal; it was making a schema-evolution decision on telemetry that a caching layer had silently made incomplete. Wherever an operation is promoted to the edge, its field usage has to be reconstructed from the edge's own request data, or the schema's owners must know the numbers are partial. ## Blast radius, and why short lifetimes beat a purge pipeline A cache multiplies whatever it stores, including mistakes. A wrong entry — a poisoned response carrying errors, a leaked viewer-specific body, a stale price — is served to everyone asking for that key until it expires. So the honest question for each candidate is not "how likely is a bad entry" but "how long can we live with one". That usually argues for short lifetimes rather than for building invalidation. A lifetime is a number you set once and it bounds staleness by construction; a purge pipeline is a system with its own availability, its own ordering problems and its own on-call. Start with a lifetime short enough that you would never need to purge urgently, measure whether the hit rate is still worth having, and only build invalidation when the arithmetic genuinely demands a long lifetime. ## The default, and the review Default to the origin. Promote a specific operation when someone can state its cost, its key cardinality, its staleness budget and its viewer-independence, and can say who notices if it is wrong. Re-review on schema change, because a promoted operation is a contract between two systems owned by different teams: a field added to a selection, or a field that quietly becomes viewer-dependent, can invalidate an edge rule that nobody on the schema side has ever read.

  • How would you set the initial lifetime for a newly promoted operation?
    Short enough that a wrong entry expires before anyone escalates — often tens of seconds — then measure. Hit rate for a fixed lifetime is a function of the request rate per distinct key, so you can compute what a longer lifetime would buy before you risk it. Lengthen only when the measured gain justifies the extra staleness, and stop at the product's stated budget rather than at whatever still looks acceptable.
  • How do you keep field-usage data honest once operations are served from the edge?
    Reconstruct the missing traffic from the edge's request logs: you know which operation each key belongs to, so you can attribute hit counts back to the fields that operation selects. Failing that, mark the affected operations in the schema's usage reporting so nobody reads a low number as an absence of callers. The rule is that a deprecation decision must never rest on miss-only telemetry.
  • A team asks to raise a lifetime from 60 seconds to an hour for a better hit rate. What do you ask?
    What the product's tolerance for stale data on that read actually is; what a wrong entry would look like and who would notice; and what the hit-rate gain really is, since past a certain point extra lifetime buys very little on a key that is requested every few seconds. If the gain is real and the staleness is acceptable, the next question is whether you now need targeted invalidation and who will operate it.

saying these in an interview costs you the question

  • Turns edge caching on for the whole endpoint
  • Picks candidates by hit rate alone
  • Ignores that a hit skips every resolver
  • Trusts field-usage numbers behind a cache
  • Caches free-text search results
  • Reaches for a purge pipeline before a short lifetime

context