Your grid sessions fail mid-run with a router-level 404, not a WebDriver error. Why?
answer
- which component actually answered
- the reply encodes where, not only what
- a table built at startup can change
- a load balancer may pick another instance
- a lost route is not a lost session
basics
~20 sThe rewritten reply bound the session to a route, and the router can no longer resolve it. That is a routing failure, answered by the router itself, so you get its own message rather than a WebDriver session error.
solid answer
~50 sWhen a router rewrites the reply, it hands you an identifier that carries **routing state**, and that state is valid only while the router's own table still resolves it. Ggr — unmaintained by its own README — builds its routes map at startup from its quota files, keying each host by a digest of its address. Every later command is looked up by the prefix it sliced out of the path; a miss is logged as `ROUTE_NOT_FOUND` and answered by Ggr's own error handler with `route not found` and HTTP 404, not by any browser. Its documentation names the usual cause outright: quota-file inconsistency between instances, so a session created through one instance cannot be resolved by another behind the same load balancer. The tell is the *layer* of the error — a router's body and status, arriving mid-run, on a session that was healthy a moment earlier.
code
go · 16 lines// Ggr, proxy.go - the prefix in the session id no longer resolves.
sum := r.URL.Path[head:tail]
proxyPath := r.URL.Path[:head] + r.URL.Path[tail:]
h, ok := routes[sum] // built at startup from this instance's quota files
if ok {
// ...rewrite host and path, forward upstream, return...
}
log.Printf("[%d] [-] [ROUTE_NOT_FOUND] [-] [%s] [%s]\n", id, remote, proxyPath)
r.URL.Host = listen
r.URL.Path = paths.Err
// ...and the handler that path lands on:
func err(w http.ResponseWriter, _ *http.Request) {
reply(w, errMsg("route not found"), http.StatusNotFound)
}go deeper
Notice which component answered a failure. A message and status that mention no session at all usually came from something in front of the browser, not from the browser itself.
Be ready to explain why an identifier a router rewrote can stop resolving: the routing table it is read against is configuration, and configuration changes while sessions are still running.
Diagnose by correlation and layer. Check whether failures began at a reload or deployment, whether they affect a fraction of commands, and whether the responding body is the router's own rather than a WebDriver error.
Decide how routing configuration is rolled out across a fleet. Live sessions depend on it, so treat identical tables as a release invariant and prefer draining a host to removing it while sessions are still open.
## Read the layer, not just the failure When commands on a live session start failing, the first useful question is which component answered. A browser that has gone away produces a WebDriver error with a recognised code and a session-shaped message. A router that cannot work out where to send the request produces something else entirely: its own body, its own status, and no reference to a session at all. Ggr makes this concrete. Its proxy handler slices the host prefix out of the request path and looks it up. On a miss it logs `ROUTE_NOT_FOUND`, points the request at its own error path, and that handler replies with the message `route not found` and HTTP 404. Nothing about that answer came from a browser, a hub or a session. For a bidding-screen suite it looks like a sudden, total loss of a session that was fine on the previous command. ## Why a rewritten reply can expire The reply you were given is not a plain name. Because the router rewrote it, it encodes **where** as well as **what**, and the encoding is only meaningful against the router's live table. Anything that changes that table can invalidate an identifier already in flight: - The table maps a digest of a host's address to that host, so an instance can resolve only the digests its own configuration taught it. - Ggr's map only ever gains entries — its `appendRoutes` adds and never deletes — so a running instance does not forget a host mid-session; what it never learned is the problem. - A load balancer can send the next command to a different instance, which resolves the prefix only if its files list the same host. - A restart re-reads those files from scratch, so a host dropped from the configuration is simply absent from the new process's map while sessions issued by the old one may still be live. - A host renamed, re-addressed or moved to a different port hashes to a different digest, and so becomes a different key entirely. Ggr's own operational documentation names the load-balancer case and the never-learned case plainly: `ROUTE_NOT_FOUND` usually means quota-file inconsistency between instances, and its guidance for running several instances behind a load balancer states that it is very important for every instance to carry exactly the same files, because otherwise some of them will answer 404 when a request arrives carrying a host prefix they do not know. ## Telling the causes apart | what you see | what it usually means | |---|---| | router body and status, mid-run | the route was lost, not the session | | WebDriver session error, mid-run | the browser or its session ended | | failures only on some commands | requests are landing on different instances | | failures start after a restart or deploy | the new process never learned that host | The most useful signal is correlation. If the failures begin at a deployment, a configuration reload or a scaling event rather than at a timeout boundary, the routing table is the suspect. If they affect a fraction of commands and shift around, the fleet of routers is inconsistent rather than broken. ## What to do about it 1. **Capture the reply.** Record the returned session identifier and which address keys came back, at session start, and mask the user info out of any address you log — the rewrite that built it copied the front door's credentials in. Without that record you cannot tell a lost route from a lost session. 2. **Keep the routing tables identical.** Every instance a load balancer can pick must resolve every identifier any of them issued. Treat their configuration as one artefact, deployed together. 3. **Drain before you remove a host.** A host dropped from the configuration is absent from every instance that starts afterwards, so stop new sessions landing on it and let the live ones finish first. 4. **Distinguish the errors in the harness.** A router-level status is not a flaky browser; classify it separately so it does not vanish into a general retry policy. 5. **Do not treat a retry as a fix.** If the route is genuinely gone everywhere, re-issuing the same command cannot succeed; if it sometimes succeeds, the fleet is disagreeing with itself and the retry is hiding that. Either way the recovery is a new session. 6. **Never reconstruct the identifier.** The client's only valid move is to return the string it was handed, so a repair attempt that rebuilds it makes the diagnosis worse. ## The property worth carrying away A rewritten reply is a **coupling**, not just a convenience. The moment an intermediary hands you an identifier or an address of its own making, you depend on that intermediary's internal state for the life of the session, and the strength of the coupling is invisible from the client. The same is true of a hosted provider, where you cannot read the table at all: a session is reachable for as long as the provider's own routing says it is, and when that stops being true the failure looks like whatever layer decided it, not like a browser problem. So the habit that survives both cases is: record what the reply gave you, classify failures by the layer that answered, and replace the session rather than retry the command.
- How would you tell this apart from the browser simply having gone away?By who answered. A browser or hub that ended the session returns a WebDriver error with a session-shaped code and message; a router that cannot resolve the route returns its own body and status with no session reference at all. Ggr's is `route not found` with HTTP 404. Correlating the start of the failures with a deployment or reload rather than a timeout boundary settles it.
- Why can a fleet of routers make this intermittent rather than total?Because each request can land on a different instance. A router that keeps all its routing state in the identifier is stateless by design, so any instance can serve any session — but only if it holds the same table. When the tables diverge, the instance that issued the identifier still resolves it and the others do not, so the same session fails or succeeds depending on which one the load balancer picked.
- What is the correct recovery once a route is lost mid-run?Start a new session. The identifier cannot be resolved by the instance that must route it, so if the whole fleet has lost the route a retry is futile, and if only some instances lost it an intermittent success is the fleet disagreeing with itself — worse than a clean failure. Classify the router-level status separately from browser errors so a general retry policy does not hide it.
A session identifier the router rewrote is a cloakroom ticket: the number means something only to the cloakroom that issued it, and only while that cloakroom's own book still has the matching row. Hand it to a second counter working from a different book and you get a shrug, not your coat.
saying these in an interview costs you the question
- Reading a router's status as a flaky browser session
- Retrying the same command after the route has gone
- Assuming every instance behind a balancer shares its configuration
- Believing an identifier stays valid for the session's whole life
- Reconstructing the identifier to work around the failure
- Diagnosing from the failing command instead of the reply