skip to content

One client opens a GraphQL WebSocket and starts hundreds of subscriptions — which limits stop it?

level: seniorimportance: must knowfreq 45%

answer

  1. One request that never ends
  2. Score once, pay per event forever
  3. Count sockets, streams, events, bytes
  4. Per-process caps die at the load balancer
  5. A slow reader fills the outbound buffer

basics

~10 s

Cap active subscriptions per connection and connections per identity, reject duplicates, bound each stream's event rate, and bound the outbound buffer so a client that stops reading is disconnected rather than buffered forever.

solid answer

~50 s

The ordinary controls miss this entirely: an HTTP rate limit counts requests, and a subscription is one request that never ends; a cost limiter scores the document once at subscribe time, while the real cost is that score multiplied by events multiplied by hours. So the limits have to be counted in the units the abuse actually uses. Cap **active subscriptions per connection**. Cap **connections per identity**, in a shared counter — a per-process cap is bypassed by reconnecting through the load balancer. Reject a subscription that duplicates one already running on the connection. Bound the **event rate** per stream. Bound the **outbound buffer**: a subscriber that stops reading does not stop the server producing, and the send buffer for that socket grows until you either drop events, drop the connection, or run out of memory. Finally, close idle connections. None of this is specified.

code

pseudocode · 8 lines
pseudocode
onSubscribe(connection, subscriptionId, document):
    if connection.activeSubscriptions.size() >= 20:
        terminateSubscription(subscriptionId, "subscription limit reached")
        return
    if connection.hasEquivalent(document):
        terminateSubscription(subscriptionId, "duplicate subscription")
        return
    connection.activeSubscriptions.add(subscriptionId, start(document))

go deeper

for a junior

Be ready to say why counting HTTP requests misses subscriptions: one connection is opened once and then carries many long-lived streams as messages.

for a middle

Explain the countable units — subscriptions per connection, connections per identity, events per stream, bytes buffered — and why a cost score computed at subscribe time understates a stream's true cost.

for a senior

Show production judgement: make the identity counter distributed, choose deliberately between disconnecting and dropping for slow consumers, close idle sockets, and name the metrics that reveal abuse before memory does.

for a principal

Own the policy: which limits are contractual and published, how a legitimate high-fan-in client gets a higher quota, and whether subscriptions belong on the same fleet as queries given their different failure mode.

## Why the usual limits do not see this A GraphQL API's abuse controls are almost all shaped around the request/response endpoint, and every one of them has a blind spot for subscriptions. **Rate limiting counts requests.** A subscriber makes one connection attempt and then sends operations as messages on an already-open socket. From the perspective of anything counting HTTP requests per minute, an abusive subscriber is quieter than a normal user. **Cost limiting scores a document once.** When a subscription operation arrives the limiter can score its selection set exactly as it would a query. But the true cost is that score multiplied by the number of events, multiplied by the lifetime of the stream. A document scoring 40 points is trivial as a query and enormous as a subscription that fires 4,300 times an hour. **Deadlines do not apply.** A subscription is designed to outlive any wall-clock budget, so the per-operation deadline that protects the query endpoint has nothing sensible to say here. The consequence is that a server can be perfectly hardened for queries and completely open to a client that opens one socket and starts subscribing. ## Count the units the abuse actually uses Four countable resources, each needing its own cap. **Active subscriptions per connection.** The subprotocols let a client run many independent streams over one socket, each with its own identifier, and they specify no limit. Twenty is a generous ceiling for a real user interface; a warehouse operations dashboard watching every aisle might legitimately want 47 and should be made to justify it. Enforce at subscribe time, before the stream starts. ``` onSubscribe(connection, subscriptionId, document): if connection.activeSubscriptions.size() >= 20: terminateSubscription(subscriptionId, "subscription limit reached") return if connection.hasEquivalent(document): terminateSubscription(subscriptionId, "duplicate subscription") return connection.activeSubscriptions.add(subscriptionId, start(document)) ``` **Connections per identity.** A per-connection cap alone is defeated in the obvious way: open more connections. The counter therefore has to be keyed on the authenticated identity and shared across server instances, because a per-process counter is trivially bypassed by reconnecting until the load balancer hands you a different node. This is the cap people most often forget to make distributed. **Duplicate streams.** The same subscription document with the same variables, started five times under five identifiers, costs five executions of every event. Rejecting equivalents on a connection is cheap and rarely inconveniences a legitimate client. **Event rate per stream.** A subscription on a high-churn source — every stock movement in a distribution centre — can produce thousands of payloads a minute for a client that cannot use them. A per-stream rate cap that conflates or drops intermediate events keeps one busy source from saturating one socket. ## The unread stream This is the failure that surprises people, because the client is not doing anything obviously abusive: it simply stops reading. A browser tab is backgrounded, a mobile connection degrades, a script opens sockets and never consumes them. The server keeps producing. Under the hood, the socket's flow control means the server cannot push bytes the peer is not acknowledging, so the unwritten payloads accumulate in the server's outbound buffer for that connection. Nothing in the server's own logic stops that growth. Memory per connection becomes unbounded, and because it is per connection, a few hundred stalled subscribers can exhaust a process that handles thousands of healthy ones without noticing. Three policies, and you must choose one deliberately: * **Bound and disconnect.** Cap the buffer, and when it is exceeded terminate the connection. Predictable memory, and it puts recovery in the client's hands. * **Conflate or drop.** Keep only the latest payload per stream, or drop intermediate ones. This changes the API's contract — the client can no longer assume it sees every event — so it must be documented, not silently introduced. * **Degrade to a signal.** Send "something in this scope changed" and let the client refetch. Constant-size payloads, and the client resyncs from a query. **Idle connections** are the mirror image: a socket with no active subscriptions and no traffic is pure overhead. The subprotocols provide keepalive messages, so a server can distinguish a live peer from a dead one and close what is neither subscribing nor answering. ## Nothing here is specified Worth saying plainly, because interviewers probe it: the GraphQL specification says nothing about subscription limits, and the WebSocket subprotocols define message flows, not quotas. There is no standard field for "maximum subscriptions", no negotiated limit, no defined error for exceeding one. Every cap in this list is a server policy you implement and document. ## What to measure You cannot cap what you do not count. Track open connections per identity, active subscriptions per connection as a distribution rather than a mean, events emitted versus events actually delivered — the gap is your dropped or buffered volume — and the high-water mark of outbound buffers. The first sign of abuse is almost always a long tail in subscriptions-per-connection, visible long before memory becomes a page.

  • Why is a per-process cap on connections per user insufficient?
    Because reconnecting usually lands on a different node. Each process sees one connection and allows it, while the user accumulates as many as there are instances behind the load balancer. The counter has to live somewhere shared and be keyed on the authenticated identity, with a TTL so a crashed process does not leak the user's quota forever.
  • What are the tradeoffs of dropping events for a subscriber that has stopped reading?
    Dropping bounds memory but silently changes the contract: the client can no longer assume it observed every event. That is fine for state-shaped streams where the latest value supersedes earlier ones, and wrong for event-shaped streams where each occurrence matters. Decide per subscription, document it in the schema, and prefer disconnecting over silently losing events you promised to deliver.
  • Does either WebSocket subprotocol for GraphQL define a maximum number of subscriptions?
    No. Both define message flows and identifiers, not quotas — there is no negotiated limit field and no standard error for exceeding one. A server that caps subscriptions is implementing its own policy and must choose how to signal it, which is why clients need documentation rather than a specification reference to handle the case.

A gym that charges per visit is defenceless against one member who walks in once and never leaves; you have to start counting lockers, not turnstile clicks.

saying these in an interview costs you the question

  • Says the HTTP rate limiter already covers subscriptions
  • Applies the query cost budget once and considers the stream bounded
  • Caps subscriptions per process rather than per identity
  • Assumes a client that stops reading stops costing memory
  • Silently drops events without changing the documented contract
  • Claims the subprotocol negotiates a maximum subscription count

context