How do you budget capacity for a GraphQL subscription feature whose cost is events multiplied by subscribers?
answer
- Executions per second, not sockets
- Size from the hottest entity, not the average
- One event that touches everyone at once
- Thin event plus a cacheable refetch
- Decide the degradation before the peak
basics
~20 sBudget in executions per second, not sockets: peak subscribers on the hottest single entity times that entity's event rate, times the CPU and backend calls one execution costs. Then decide whether the feature can afford that.
solid answer
~50 sSockets are almost never the binding resource; **executions per second and the backend calls they make** are. Size from the worst concentration rather than the average: the hottest entity's peak subscriber count times its peak event rate gives executions per second, multiplied by measured CPU and downstream calls per execution. Add the correlated cases — an event that touches every subscriber at once, and the resubscribe wave after a rolling restart. Then treat the design as a lever, not a constant. A thin "this entity changed" payload that clients follow with a normal query moves the cost onto a request path you can cache, shed and rate-limit, at the price of an extra round trip and weaker ordering. A fat per-subscriber payload gives lower latency and no refetch storm, and costs a full execution per subscriber per event. Choose per feature, write the budget down, and alert on the ratio rather than the totals.
code
pseudocode · 16 linespeakSubscribersPerEntity = 3187 # one flight at sale open
peakEventsPerSecond = 2.4 # seat holds and releases
cpuPerExecution = 0.0018 # seconds, measured
backendCallsPerExecution = 0.4
executionsPerSecond = peakSubscribersPerEntity * peakEventsPerSecond # 7,649
coresRequired = executionsPerSecond * cpuPerExecution # 13.8
backendCallsPerSec = executionsPerSecond * backendCallsPerExecution # 3,060
# correlated worst case: one event relevant to every subscriber at once
burstExecutions = peakSubscribersPerEntity # 3,187 in one tick
# thin-event alternative: cost moves to a cacheable query path
executionsPerSecond_thin = peakSubscribersPerEntity * peakEventsPerSecond
cpuPerExecution_thin = 0.0002
refetchQueriesPerSecond = executionsPerSecond_thin # served from a shared cachego deeper
Be ready to say that the cost of a subscription feature grows with subscribers times events, not with connections, and that a payload which needs a backend call pays for it once per subscriber per event.
Explain the arithmetic and where the inputs come from: measured CPU and backend calls per execution, peak event rate, peak subscribers per entity. Be able to contrast a fat payload with a thin notification the client refetches.
Show that you plan for the shapes that break the average: concentration on one hot entity, an event correlated across every subscriber, and the reconnect wave after a deploy. Name the degradation you would trigger and in what order.
Own the budget as a contract. Interviewers expect a defensible cost model per feature, an argued position on thin events versus per-subscriber payloads, and a mechanism — ceilings, gauges, review rules — that stops an unbounded multiplier being created by teams who do not carry it.
## Size the multiplier, not the traffic The instinctive capacity model for a realtime feature counts connections. For GraphQL subscriptions that model is wrong in a way that is expensive to discover late, because the cost is not per connection, it is per (subscriber, event) pair — each one is an execution. The number to size against: ``` executions/s = peak subscribers on the hottest entity x peak events/s for that entity cost = executions/s x CPU per execution + executions/s x backend calls per execution ``` On an airline seat map at sale open: 3,187 watchers on one flight, 2.4 seat events per second, 1.8 ms of CPU and 0.4 backend calls per execution gives ~7,650 executions per second, ~14 cores and ~3,060 backend calls per second — for one flight, out of the several hundred in the sale. Doing this arithmetic before building is the entire exercise; almost every subscription capacity surprise is someone having multiplied the average instead of the concentration. ## The distributions that actually hurt - **Concentration.** Ten thousand subscribers over a thousand entities is a benign shape. Ten thousand on one entity is the same subscriber count and a hundredfold the fan-out cost, and all of it lands on whichever nodes hold those subscribers. Alert on subscribers-per-entity, not on total subscribers. - **Correlated events.** Most events touch a few watchers; some touch all of them at once. An aircraft swap invalidates the whole seat map for every watcher simultaneously. Your budget has to survive the correlated case, because that is the one that arrives during the sale. - **The resubscribe wave.** Subscription state is node-local and dies with the process, so a rolling deploy converts steady state into a synchronised reconnect-and-resubscribe burst. Capacity for the wave is capacity you must hold idle the rest of the time, unless the client's retry policy spreads it. - **Payload growth over time.** The per-execution cost is not fixed. A product team adds two fields to a subscription payload and quietly multiplies the backend calls by the fan-out. The budget needs an owner and a review point, or it degrades without any single change looking expensive. ## The lever that changes the arithmetic The biggest decision is what the subscription payload carries. **Thin event, client refetches.** The subscription delivers little more than "seat map for QX414 changed", and the client issues a normal query for the current state. The per-event server cost collapses to a near-trivial execution, and the expensive work moves onto the query path — which is cacheable, can be served from a shared response cache, can be rate-limited per client, and can be shed under load without breaking the socket. The costs are real: an extra round trip of latency, a refetch storm from every watcher of a hot entity at the same instant, and a client that must reconcile a snapshot with the notification that triggered it. **Fat payload, per-subscriber execution.** The subscriber gets exactly the fields it selected with no follow-up. Latency is one hop, there is no storm, and every event costs a full execution per subscriber, forever. The honest rule of thumb: as subscribers-per-entity grows, thin events win, because the refetch path has a cache in front of it and the execution path does not. For a small, high-value audience — the 62 accredited agents rather than the 3,187 watchers — the fat payload is usually right. ## Degradation you decide in advance A budget is only real if you have decided what to give up when it is exceeded. Options worth naming: coalesce events per subscriber into one execution per interval so a burst costs a fixed rate rather than a multiple; degrade a hot entity's subscribers to thin events dynamically; cap concurrent subscribers per entity and fall back to polling above the cap; shed the least valuable audience first, since a passenger watching a seat map and an agent completing a booking are not equal. Choosing which of these fires, and in what order, is the part that cannot be delegated to a library. ## The organisational half A 4-person platform team owning the subscription tier cannot review every product team's payload, so the budget has to be expressed as something enforceable rather than as advice: a published cost per subscription feature, a hard ceiling on subscribers per entity, executions-per-delivery and cost-per-execution as gauges the owning team watches, and a rule that a new subscription field costing a backend call needs the fan-out multiplier attached to its review. The alternative — a platform team absorbing an unbounded multiplier created by decisions made elsewhere — is the failure mode this question is really probing for. ## What is specified here None of it. The GraphQL specification defines a subscription's per-event execution for one operation and says nothing about many subscribers, cost, fan-out or degradation. Every number and every lever above is engineering judgement layered on top, which is exactly why it is a principal-level conversation rather than a recall one.
- Why is a thin notification plus a client refetch usually cheaper at high fan-out?Because it moves the expensive work onto a request path that has the tools a stream does not: a shared response cache means one computation can serve thousands of refetches, and the path can be rate-limited or shed without dropping connections. A per-subscriber execution has none of that — every subscriber pays in full, every time. You trade a round trip of latency and a refetch burst for a cost that stops scaling with the audience.
- What would make you keep fat subscription payloads despite the multiplier?A small, bounded audience with high value per subscriber, strict latency requirements, or payloads whose content is genuinely per subscriber so a shared refetch would not be cacheable anyway. Sixty agents completing bookings justify a full execution each; thirty thousand watchers of a public seat map do not.
- Which single metric would you put on the wall for a subscription tier?Peak subscribers on the hottest single entity. It is the multiplier in every cost calculation, it moves for product reasons rather than infrastructure ones, and it predicts the correlated burst better than totals do. Executions per second is the number you budget with; subscribers-per-entity is the number that tells you it is about to change.
- How do you stop the per-execution cost drifting upward after launch?Make the multiplier visible at review time. A field added to a subscription payload that costs a backend call is not one call, it is one per subscriber per event, so the review needs the current fan-out attached to it. Pair that with a measured cost-per-execution gauge per subscription operation, so drift shows up as a trend rather than as an incident.
saying these in an interview costs you the question
- Sizes the tier by concurrent connections alone
- Plans from average subscribers rather than the hottest entity
- Ignores the correlated event that touches every subscriber
- Forgets the resubscribe wave after a rolling restart
- Has no decided degradation path at the ceiling
- Lets product teams grow payloads with no fan-out review