skip to content

On a WebSocket that multiplexes many channels, why must the server authorize each subscribe message when the upgrade was already authenticated?

level: middleimportance: should knowfreq 44%

answer

  1. who versus what
  2. the protocol has no channels
  3. principal from the connection, not the message
  4. a secret name is not a lock
  5. permission changes reach live subscriptions

basics

~20 s

Authenticating the upgrade establishes who is connected, not what they may read. Each subscribe or publish names a channel, so the server must check that principal against that channel when the message arrives, and end subscriptions whose access is withdrawn.

solid answer

~50 s

RFC 6455 has no notion of a channel: frames carry an opcode, a length and a payload, and channels are something the application or subprotocol puts in message payloads. The handshake binds a **principal** to the connection; it says nothing about which of many channels that principal may read. So every subscribe or join is a request for a resource and gets its own check on the server: take the principal recorded at the handshake (never an id the client wrote into the message), ask the policy whether it may read that channel, and on refusal reply with an error and register nothing. A long random private-channel name is not a control, because names leak through logs, links and client code. Publishes need their own write check. And since the decision is cached in the subscription, a permission change must find and remove existing subscriptions.

code

pseudocode · 20 lines
pseudocode
on subscribe(conn, msg):
    principal = conn.principal
    if not policy.canRead(principal, msg.channel):
        send(conn, {type: "error", ref: msg.id, reason: "forbidden"})
        return
    subscriptions.add(conn, msg.channel)
    index.add(principal, msg.channel, conn)
    send(conn, {type: "subscribed", ref: msg.id})

on publish(conn, msg):
    if not policy.canWrite(conn.principal, msg.channel):
        send(conn, {type: "error", ref: msg.id, reason: "forbidden"})
        return
    fanOut(msg.channel, {from: conn.principal, body: msg.body})

on accessRemoved(principal, channel):
    for conn in index.lookup(principal, channel):
        subscriptions.remove(conn, channel)
        index.remove(principal, channel, conn)
        send(conn, {type: "unsubscribed", channel: channel, reason: "access removed"})

go deeper

for a junior

Recall that the WebSocket handshake only establishes who is connected, and that each channel a client asks for needs its own yes or no from the server.

for a middle

Explain the subscribe-time check step by step: principal taken from the connection, channel resolved, policy asked, nothing registered on refusal, and the separate write check on publish.

for a senior

Show the production judgment: reject secret channel names as a control, route permission-change events to the node holding the sockets, and end one subscription without dropping the connection.

for a principal

Weigh where the decision lives: checked once at subscribe with event-driven invalidation, re-evaluated on a timer, or checked per delivery, against channel fan-out size and how fast access must disappear.

## Two questions, asked at two different times A WebSocket server faces two separate authorization questions, and a design that answers only the first is one of the commonest holes on this protocol. 1. **Who is on this connection?** Answered once, on the HTTP request that opens the socket, by whatever credential the handshake carried: a cookie, a ticket, a token. If the check passes, the server answers `101 Switching Protocols` and records a **principal** against the connection. 2. **May this principal read, or write, this channel?** Asked every time a message names a channel, long after the handshake is over, and answered by nothing unless the server asks it. The first answer does not imply the second. A nurse who may open the ward console may watch ward 7 and not the oncology ward; a tenant who may connect may read its own channels and nobody else's. ## The protocol has no channels RFC 6455 defines frames (an opcode, a payload length and a payload, plus a masking key on frames sent by the client) and the messages assembled from them. Nothing in a frame names a channel. The specification leaves "application-level protocols layered over the WebSocket Protocol" to subprotocols negotiated through `Sec-WebSocket-Protocol`, and what it says about client authentication (Section 10.5) concerns the handshake: the server may use any mechanism a generic HTTP server can, such as cookies, HTTP authentication or TLS authentication. So a **channel** (a room, a topic, a feed, a stream key) is something your application or subprotocol invented, carried as a string inside a message payload. When one socket multiplexes many channels, every subscribe or join message is a request for a resource, and it deserves exactly the authorization check that an HTTP `GET /channels/ward-7` would get. An open socket changes nothing about that. ## The check at subscribe time The server-side handler does four things, in this order: 1. Takes the principal **from the connection**, the one recorded at the handshake, and never from a user id, tenant id or role the client wrote into the message. 2. Parses the channel name the message asks for and resolves what it refers to. 3. Asks the policy: may this principal read this channel, now? 4. On a refusal, replies with an application-level error that references the request and **registers nothing**; on success, registers the subscription and only then starts delivering. A refused subscribe does not normally need to close the connection, because the client's other subscriptions are legitimate. RFC 6455 does define close code `1008` (Policy Violation) for an endpoint that received a message violating its policy, and closing with it is a reasonable implementation choice for a client that keeps probing channels it may not read. Codes 4000-4999 are reserved for private use by prior agreement between applications. ## Why an unguessable name is not a control Giving a private channel a long random name such as `ward-7-c1f9e2a4...` feels like a lock. It is a bearer secret with none of a credential's handling: - it appears in server and client logs, error reports and debugging output; - it gets pasted into links, tickets and screenshots, and still works after it is copied; - it often ships in client code, or in a list the client received for another purpose; - it cannot be withdrawn from one person without renaming the channel for everyone. Possession of a name is not permission. Randomness is a fine defence against **enumeration** when it sits on top of a real check; it never replaces one. ## Publishes are a separate decision A server that checks subscriptions and then accepts any publish from an authenticated connection lets any user inject messages into channels they may only read, or may not read at all. Readers who passed their own check then receive the forged message and trust it. | Message | Question the server asks | What a missing check allows | |---|---|---| | subscribe / join | may this principal **read** this channel? | reading another tenant's stream | | publish / send | may this principal **write** to this channel? | posting into a read-only or foreign channel | | presence or member-list query | may this principal see **who** is here? | learning who watches a private channel | Fan-out to subscribers is not a check on the sender: it decides who receives, not who was allowed to speak. For the same reason, stamp the sender with the server-known principal rather than trusting a sender field from the payload. ## When permissions change A subscribe-time decision is a decision **cached in the subscription**. When a principal loses access to one channel (removed from a team, a patient transferred off the ward) the connection and its credential remain valid, so delivery continues unless the server acts. - Keep an index from principal and channel to live subscriptions, held wherever the sockets live, and route each permission-change event to the node that holds them. - On the event, remove the affected subscription, send the client an application message saying it ended and why, and leave the connection and its other subscriptions alone. - As a backstop for missed events, re-evaluate long-lived subscriptions on a timer. Checking the policy on every delivered message is always current, but costs one decision per message per subscriber on a busy channel. Losing the credential itself, rather than one channel's permission, is the separate mid-connection expiry problem. ## What an interviewer is listening for - That **authenticating the connection** and **authorizing each channel** are named as two separate decisions. - That the principal comes from the connection, never from the payload. - That an obscure channel name is rejected as a control. - That publishes are checked as well as subscriptions. - That a permission change reaches subscriptions that already exist.

  • Should a refused subscribe close the whole WebSocket connection?
    Usually not. Reply with an application-level error that references the request and register nothing; the connection's other subscriptions are legitimate. RFC 6455 defines close code `1008` (Policy Violation) for an endpoint that received a message violating its policy, and closing with it is a reasonable choice for a client that keeps probing channels it may not read, since it ends everything that client had open.
  • Why not check the policy on every message delivered to a subscriber instead of at subscribe time?
    Because it costs one policy decision per message per subscriber, which on a busy channel with thousands of readers is the most expensive place to put it. It is always current, though. Most designs check at subscribe time, invalidate on permission-change events through a principal-to-subscription index, and re-evaluate long-lived subscriptions on a timer as a backstop for missed events.
  • Where does the principal used in the subscribe check come from?
    From the connection: what the server recorded when the handshake was authenticated. A user id, tenant id or role inside the subscribe message is client input like any other and proves nothing. The same principal should stamp the sender of any message the connection publishes, instead of a sender field from the payload.

saying these in an interview costs you the question

  • Believes authenticating the upgrade authorizes every channel the client later subscribes to
  • Treats a long random private channel name as an access control
  • Trusts a user or tenant id the client puts inside the subscribe message
  • Checks subscriptions but lets any authenticated connection publish to any channel
  • Assumes a subscription granted once stays valid until the client unsubscribes
  • Relies on the client application hiding channels the user may not open