How does a routing data-access layer decide to send a unit of work to a replica, and when is that decision fixed?
answer
- decided at the boundary, not per statement
- the read-only marking is the signal
- chosen before the connection is taken
- one connection, one server, whole unit
basics
~20 sThe read-only marking declared on the unit is the usual routing signal, and it is read before the unit takes a connection. That connection then points at one server for the unit's whole life, so nothing can re-route it later.
solid answer
~50 sRouting happens at the boundary, not per statement. When a unit of work starts, the layer looks at what the unit declares - above all whether it is read-only, sometimes an explicit route hint - and picks which datasource to borrow a connection from. From then on the unit is bound to one connection and therefore one server, so discovering half way through that a write is needed means failing, or ending the unit and starting a fresh one against the primary. Two consequences follow. The read-only marking now carries two jobs at once, an in-layer optimisation and a routing instruction, so a careless mark is no longer just a lost write but a read served from a lagging server. And an inner boundary that joins an outer unit inherits the outer connection and its target, whatever the inner marking says.
go deeper
Remember the shape: the layer picks a server when the unit begins, based on what the unit declared, and keeps that connection until the unit ends. Reads can go to a replica; writes go to the primary.
Explain why the decision has to precede the first statement - one unit binds one connection, and a connection points at one server - and why statement sniffing cannot substitute for a declared marking.
Bring in the consequences: the marking now carries a staleness promise as well as an optimisation, joined inner boundaries inherit the outer target, and a read path that grows a write has to be restructured rather than re-routed.
Own the contract. Decide what the routing inputs are, whether hints are allowed alongside the marking, how request-scoped pinning interacts with them, and how a team learns that switching routing on changed the meaning of every existing read-only mark.
## The decision point A routing layer sits between the application's units of work and a set of datasources - typically one primary that accepts writes and one or more replicas that receive its changes asynchronously. Its job is to answer one question at one moment: **when this unit starts, which datasource should its connection come from?** The moment matters as much as the answer. A unit of work borrows a connection when it begins (or at its first statement, in layers that defer the checkout) and holds it until it ends. A connection is attached to exactly one server. Therefore: - the routing decision must be made **before the first statement runs**; - it is **fixed for the whole unit**; - there is no "move this unit to the primary now" once it has started. ## What the layer routes on In practice the inputs are all things declared at the boundary, because that is all the layer knows before any query has run: - **the read-only marking** on the unit - by far the most common signal; - **an explicit route hint** naming primary or replica, for cases the marking cannot express; - **request-scoped state**, such as a flag saying this request has already written and must be pinned to the primary; - **the tenant or shard identity**, where routing selects among per-tenant datasources rather than among replicas. What it deliberately does not route on is the content of the statements. Sniffing the first statement and calling a `SELECT` read-only is tempting and wrong: the connection has already been chosen by then, and the first statement in a unit says nothing about whether the fifth one writes. | approach | when it decides | fails when | |---|---|---| | declared read-only marking | before the connection is taken | the marking is dishonest | | explicit route hint | before the connection is taken | hints drift out of date with the code | | statement sniffing | after the connection exists | a later statement in the unit writes | ## Why the marking now means two things Before routing, read-only was a local optimisation: skip the snapshot, skip the dirty check. With routing wired to it, the same marking also says **this may be served by a server that is behind the primary**. Those are different promises, and a team that adopted the marking purely for the memory saving may not have audited its paths for the second one. This is worth stating explicitly in an interview, because it explains a whole class of incidents: a read path that was correctly marked read-only for years becomes stale-data-sensitive the day routing is switched on, without a single line of that path changing. ## Nesting and joining Units nest. When an inner boundary joins an outer one rather than starting its own, it runs inside the outer unit's transaction, on the outer unit's connection, and therefore on the outer unit's server. An inner unit marked read-only inside an outer read-write unit is not routed anywhere - it is already on the primary. Layers usually treat a read-only marking on a joining inner boundary as advisory or reject it outright, precisely because the connection is no longer theirs to choose. The reverse - an inner write joining an outer read-only unit - is worse, because the unit is on a replica that will refuse the write or, if the marking never reached the transaction, lose it. ## The unit that turns out to need a write Sometimes a read path discovers work: a missing derived row, a repair, a stamp. Options, roughly in order of preference: 1. **Leave the read unit alone and do the write afterwards** in its own short read-write unit, accepting that the two are not atomic. 2. **Do not mark the path read-only at all** when a write is genuinely part of the use case; it belongs on the primary as one unit. 3. **Fail fast** if the layer detects the write, rather than letting it disappear. What does not work is hoping the layer will notice and re-point the connection mid-unit. ## What routing does not decide Whether replicas should exist, how many, and how the fleet is laid out are architecture questions above this layer. What this layer owns is narrower and entirely mechanical: **given a unit's declared attributes, choose a datasource before the first statement, and keep the choice for the unit's lifetime**. Getting that contract clear is what makes the stale-read and mis-pointed-connection failures diagnosable rather than mysterious.
- Why not just inspect the statements and route reads automatically?Because the connection is already chosen by the time a statement exists, and the first statement does not predict the rest of the unit. A unit that starts with a query and later updates a row would then be stranded on a replica. Content-based routing can only work per statement outside a transaction, which gives up the consistent view a unit is there to provide.
- What happens to a read-only inner boundary inside an outer read-write unit?It joins the outer unit, so it runs on the outer transaction and the outer connection, which is on the primary. The inner marking cannot re-route anything and is at best advisory for tracking. That is usually the behaviour you want: work that is part of a write use case should see that write's own uncommitted changes.
- How does a request that has already written avoid being routed to a replica afterwards?By carrying request-scoped state that the routing layer reads at the next boundary - a flag set when a unit writes, which forces subsequent units in that request onto the primary. Without it each unit is routed on its own declared attributes alone and cannot know a write just happened.
saying these in an interview costs you the question
- Thinks the layer inspects each statement and re-routes as it goes
- Believes a unit can be moved to the primary after it has started
- Ignores that a joining inner boundary inherits the outer connection
- Treats the read-only marking as only an optimisation once routing exists
- Assumes routing can tell in advance that a read path will write