skip to content

Would you standardise your organisation on Server-Sent Events for live feeds rather than WebSockets per team, and what does that commit you to?

level: principalimportance: should knowfreq 33%

answer

  1. most live features are one-way
  2. one default, one authorization story
  3. the exception needs a written trigger
  4. whoever leaves the default owns the rest
  5. a superset makes every team pay

basics

~10 s

Usually yes, with a written exception trigger. One default transport means one authorization story, one capacity model and one runbook; the commitment is owning that exception process and the products it genuinely fails.

solid answer

~40 s

Standardising is a bet that most live features are one-way, which is usually true, and it buys a single authorization path, one set of client behaviour every team already gets for free, one capacity model and one incident runbook. What you commit to is honesty about the exceptions: publish the trigger — measured upstream message rate, a per-message latency budget, binary payloads in volume, an existing subprotocol — and require any team that trips it to bring its own reconnect, resumption, authorization and capacity plan rather than inheriting nothing. The failure mode to avoid is standardising on the socket instead because it is the superset: then every product owns re-opening, resumption and mid-connection expiry, and most of them will get one of the three wrong.

go deeper

for a junior

Recognise that a team-wide default exists for a reason: one transport means one way to authorize, one way to reconnect and one way to debug a live feed.

for a middle

Explain what a default saves in practice — inherited re-open behaviour, a shared capacity model, one authorization path — and why the exception has to be written down.

for a senior

Show the cost landing on a genuinely duplex product, and what that team must own if it leaves the default: reconnection, resumption, authorization after open and expiry.

for a principal

Make the bet explicitly. State the empirical claim about the median live feature, set the default from it, publish a measurable exception trigger with obligations, and say when the decision gets re-opened.

## What the question is really asking No specification answers this. It is a lead's bet about which costs the organisation carries for the next several years, and an interviewer is listening for whether you can name those costs on both sides rather than declare a favourite. The bet rests on one empirical claim: **most features described as live are one-way.** Dashboards, status boards, feeds, job progress, incrementally produced output, alerts — none of them has upstream traffic beyond the occasional action that rides perfectly well as an ordinary request. If that claim holds in your product line, a one-way default is the smaller organisation. ## What standardising buys - **One authorization story.** Every live feed is authorized by the rule that already authorizes everything else, and security review sees one pattern rather than one per team. - **Behaviour every team inherits.** Re-opening after a drop and resuming from the last event `id` are protocol behaviour, so a team that never thought about them still gets them, which is exactly the population of teams you are designing for. - **One capacity model and one runbook.** A single way to size a fan-out tier, a single set of symptoms during an incident, one drain procedure, one dashboard shape. - **One conversation with infrastructure.** Whatever must be true of the path for an incremental response to work is established once, for everyone. - **Cheaper mobility.** An engineer moving between teams meets the same transport, and a review in one team transfers to another. ## What it costs, and who pays The cost is not spread evenly: it lands entirely on the team whose product genuinely is duplex. Denied the socket, that team builds a duplex protocol out of one request per message, discovers the latency, and either ships something worse or works around the standard quietly. **A standard with no exception process is a standard people route around**, and a routed-around standard gives you the costs of both worlds. So the commitment is the exception process, and it has to be written before anyone needs it: 1. **Publish the trigger, in measurable terms.** Sustained upstream message rate, a per-message latency budget a round trip cannot meet, binary payload volume where the encoding tax is material, or an application subprotocol both ends already speak. 2. **Require evidence, not conviction.** A team claiming the exception brings the measured rate from a prototype, not an argument about the future. 3. **Make the exception carry its own weight.** Whoever leaves the default owns reconnection, resumption, authorization after the open and expiry inside a long connection — written down, reviewed and tested, because none of it is inherited. 4. **Bound it.** The exception covers one product and one surface, not the team's whole portfolio, and it is revisited when the feature changes shape. ## The mirror-image mistake The tempting alternative is to standardise on the socket, on the grounds that it is the superset and covers every case. Consider what that commits every team to: - Writing and testing its own re-open policy, in every product, correctly. - Designing resumption — what a client asks for after a gap, and how far back the server can answer. - Handling a credential that ages out inside a connection that is still perfectly healthy. - Losing the per-request logging, metrics and rate limiting that infrastructure provided for free, and rebuilding enough of it to satisfy a review. Most teams will get one of those four wrong, and the failures show up late — during the first rolling deploy, or the first time a credential lifetime is shortened. A superset that most of the estate does not need is not a simplification. ## What honest disagreement looks like There are organisations where the default should be the socket: collaborative editing products, interactive tooling, game-adjacent surfaces, anything where the representative feature is genuinely symmetric. The test is not what any individual prefers but **what the median live feature in this product line does**, and a lead who cannot state that median is not ready to set a default. ## How to present the decision - State the empirical claim and say how you would check it: list the live features that exist today and mark which have real upstream traffic. - Set the default from that list, not from the most interesting feature on it. - Publish the exception trigger and the obligations that come with taking it. - Say when you will revisit — a new product line with a different shape is a reason to re-open the decision, and pretending otherwise is how a sensible default calcifies into a rule nobody can explain.

  • What belongs in the written trigger for taking the exception?
    Measurable conditions only: a sustained upstream message rate, a per-message latency budget that a request round trip cannot meet, binary payload volume where the encoding tax is material, or an application subprotocol both ends already speak. Each is something a prototype can demonstrate. Anything phrased as an expectation about the future is an argument, not a trigger.
  • Why not simply standardise on the socket, since it covers both cases?
    Because it moves work into every product rather than removing it. Each team then owns re-opening, resumption, authorization after the open and mid-connection expiry, and loses the per-request observability infrastructure was providing. Most teams get one of those wrong, and the failures surface late — at the first rolling deploy or the first shortened credential lifetime.
  • When should a one-way default be revisited?
    When the median live feature in the product line stops being one-way — a move into collaborative editing, interactive tooling or anything symmetric — or when exceptions stop being exceptional and several teams have taken one. Both are signals that the empirical claim underneath the default no longer holds, and a default nobody can justify is worse than no default.

saying these in an interview costs you the question

  • Standardises on a transport without any exception process
  • Picks the socket as the default because it is the superset
  • Sets the default from the most interesting feature rather than the median one
  • Lets a team take the exception without owning reconnect and expiry
  • Treats the decision as permanent and never states when to revisit