How would you decide how many minutes-long, held-open streaming responses one instance should carry, and how would you shed them safely?
answer
- occupancy, not requests per second
- measure one stream's real cost
- find the limit that binds first
- bound the lifetime so it drains
- shed at the door, not mid-stream
basics
~10 sMeasure what one open stream costs on the real execution model, find the limit that binds first, cap admission below it, bound every stream's lifetime so the fleet stays drainable, and bound per-stream buffers.
solid answer
~40 sHeld-open responses are an occupancy problem, not a throughput one, so start by measuring one: a connection and its buffers, framework per-connection state, whatever the execution model pins between writes, and whatever the handler holds — a cursor, a pooled connection, a subscription. Find the limit that binds first, which is usually a pool or worker count rather than memory, and cap admission below it with a clear retryable refusal. Then bound the lifetime, because streams do not rebalance and a deployment cannot drain what never ends; a maximum duration after which the server closes cleanly turns occupancy into a rolling problem. Add a bounded per-stream queue so one slow consumer cannot take the instance, keep the path alive against intermediary idle timeouts, and document the lifetime and any skipping to consumers.
go deeper
The idea to hold onto: a response held open for minutes occupies resources that whole time, so the question is how many can be resident at once rather than how fast each one is.
Be able to enumerate what one open stream pins — connection, per-connection buffers, whatever the execution model holds between writes, plus the cursor or pooled connection the handler keeps — and note the cost differs sharply by execution model.
Show the operating mechanisms: measured per-stream cost, an admission cap below the binding limit, bounded per-stream queues, a maximum lifetime, keep-alive writes against idle timeouts, and per-connection metrics.
Own the tradeoffs: whether the feature needs held-open streams at all, fan-out from one producer versus a query per subscriber, isolation on a separate deployment, and what lifetime and completeness guarantees you are willing to put in the contract.
An endpoint that holds a response open for minutes is a different kind of object from a request that finishes in milliseconds. Its cost is not measured in requests per second but in **occupancy**: how many are resident at once, what each one pins, and for how long. Capacity for these is a design decision, and the shedding plan is part of the design rather than an afterthought. ## Start by measuring what one stream costs Do not reason from first principles about this; hold N streams open on one instance and measure. The per-stream cost is the sum of: - **A connection**, with its kernel buffers and its slot against any connection or descriptor limit on the box and on every intermediary in front of it. - **Framework state** — the per-connection response buffer, and whatever queue the slow-reader policy allows to build. - **Whatever the execution model pins.** This varies enormously: models that dedicate a worker per in-flight request pay a large fixed cost per open stream, while models that suspend a cheap continuation between writes pay very little. The same code can differ by orders of magnitude in ceiling depending on which one it runs on. - **Whatever the handler holds** — a cursor, a pooled connection, a subscription, a cached snapshot. This is the cost people forget, and it is frequently the binding one, because pools are small. Then find which limit binds first. It is usually the pool or the worker count, not memory, and the answer changes the design rather than the sizing. ## The ceiling is only half the answer Three forces make a static number insufficient: 1. **Intermediaries close idle connections.** A path with proxies and load balancers has its own idle timeouts, and a stream that produces nothing for a while is indistinguishable from a dead one. Keeping the path alive means writing something periodically, and those periodic writes are themselves load — at a high stream count, a keep-alive interval is a throughput decision, not a detail. 2. **Long streams pin old instances.** A deployment cannot drain what does not end. Without a bound, a rollout either waits for streams that may never finish or cuts them mid-flight, which is the same outcome with worse timing. 3. **Streams do not rebalance.** Once established, a stream stays on the instance that accepted it for its whole life, so scaling out only helps *new* streams. An instance that filled up during an incident stays full. All three point at the same countermeasure: **bound the lifetime.** Give every stream a maximum duration after which the server closes it cleanly and the client establishes a new one. That turns an unbounded occupancy problem into a rolling one — rollouts drain, load rebalances, and leaks are bounded by the lifetime. It is cheap when the client can resume from a position it already knows, and that resumability is worth designing in for exactly this reason. ## Shedding, in the order it should be tried | Mechanism | What it protects | What it costs | |---|---|---| | Admission cap per instance | The ceiling itself, deterministically | Some clients are refused and must retry elsewhere | | Bounded per-stream queue | Memory, against slow consumers | Skipped records the contract must disclose | | Maximum stream lifetime | Drainability and rebalancing | Periodic re-establishment by clients | | Closing the least valuable streams | Headroom under pressure | Requires a priority notion the system must define | Refusing a stream at admission with a clear, retryable status is far better than accepting it and degrading everyone, because the refusal is visible and the degradation is not. The only rule that matters in ordering these: shed at the door before you shed in the middle. ## The design questions to raise - **Does this need a held-open stream at all?** If updates are infrequent and the consumer tolerates delay, a cheaper interaction may carry the same information at a fraction of the occupancy. Reserve held-open streams for cases where latency genuinely is the feature. - **One producer feeding many streams, or one query per stream?** Fan-out from a single source keeps per-stream work small; per-stream querying multiplies backend load by subscriber count and is the usual cause of the second outage. - **Should these live on a separate deployment?** Long-lived streams and fast request/response traffic have opposite operational profiles — one wants long drains and high connection counts, the other wants quick rollouts. Isolating them stops a stream surge from starving ordinary traffic, at the price of another thing to run. - **What is promised to the consumer?** Maximum lifetime, whether records may be skipped, and what a clean end looks like all belong in the interface contract, because every shedding mechanism above is visible to clients. ## The summary a lead should be able to give Measure the per-stream cost on the real execution model, find the limit that binds first, cap admission below it, bound every stream's lifetime so the fleet stays drainable and rebalanceable, bound the per-stream queue so one slow consumer cannot take the instance, and write all of that into the contract. A number without those mechanisms is a guess that will be wrong the first time traffic is unusual.
- Why does scaling out not relieve an instance already full of streams?Because an established stream stays on the instance that accepted it for its whole life. New capacity only takes new streams, so an instance that filled during an incident stays full. Bounding the lifetime is what makes the fleet rebalance, since every stream eventually re-establishes and can land elsewhere.
- What makes long-held streams awkward for deployments?A rollout drains in-flight work, and a stream that never ends cannot drain. You either wait indefinitely or cut clients off mid-flight. A maximum stream lifetime shorter than the drain window resolves it, and gives clients a defined moment to reconnect rather than an abrupt failure.
- Why are periodic writes on an otherwise idle stream a capacity concern?Intermediaries close connections that look idle, so a stream with nothing to say must still write periodically to keep the path open. Multiply that interval by the number of open streams and it becomes real, continuous load — the interval is a throughput decision once the count is large.
- When should long-lived streams run on their own deployment?When their operational profile starts fighting ordinary traffic: they want high connection counts and long drains, request/response traffic wants fast rollouts and quick recycling. Isolating them prevents a stream surge from starving normal requests, at the cost of another deployment to operate.
saying these in an interview costs you the question
- Picks a stream limit from intuition instead of measuring one
- Assumes autoscaling relieves instances already holding streams
- Lets streams run unbounded, so deployments cannot drain
- Degrades every consumer instead of refusing new streams at the door
- Ignores intermediary idle timeouts on quiet streams
- Gives each stream its own backend query and multiplies load by subscribers