Would you route every read-only unit of work to a replica by default or make replica routing opt-in per unit, and what does each choice cost?
answer
- two markings, two different promises
- default-on reclassifies existing paths
- opt-in leaves offload unclaimed
- freshness belongs to the endpoint contract
basics
~20 sRouting every read-only unit by default maximises offload but makes correctness depend on each marking also being stale-tolerant, which it never meant. Opt-in is safe and under-used; most systems vet a few paths and pin after writes.
solid answer
~40 sThe two markings are different promises. Read-only says this unit will not write; stale-tolerant says its caller can live with data a little behind. Making the first imply the second reclassifies every existing read path without anyone reviewing it, and the failures are silent and load-correlated. Opt-in inverts that: nothing moves off the primary until someone judged it safe, which is correct but leaves offload unclaimed and quietly rots as paths are copied. I would default to the primary, add an explicit stale-tolerant marking a path opts into, pin the rest of any request that has written, and treat freshness as a property of the endpoint's contract. Then measure: replication delay against a per-endpoint budget, with a fallback to the primary, and per-unit logging of which datasource served each read.
go deeper
Recall that sending reads to a replica trades freshness for load relief, and that a read taken right after a write is the case most likely to show the difference to a user.
Be able to argue both defaults: what default-on buys immediately and what it silently assumes about existing markings, and why opt-in is safer but tends to be under-used.
Describe the operating detail - request-scoped pinning, per-unit logging of which datasource served a read, replication delay measured against a budget, and a fallback path - and how you would roll the policy out incrementally.
Own the framing: freshness is a per-endpoint contract, not a repository detail. Decide the default, the second marking, the budget mechanism, capacity headroom for the fallback, and how the distinction is kept alive as the team turns over.
## Why this is a real decision and not a preference Once the routing layer exists, someone must choose what the read-only marking means. Two candidate policies: - **default-on**: every read-only unit goes to a replica unless something pins it back; - **opt-in**: everything goes to the primary unless a path explicitly declares it may be served from a replica. The technical difference is one line of configuration. The organisational difference is large, because default-on **retroactively changes the meaning of every existing marking in the codebase**, made by people who were promising "this will not write" and not "this may be stale". ## The two promises, kept apart | property | what it asserts | what it protects | |---|---|---| | read-only | no write happens in this unit | change-tracking cost, and correctness of routing | | stale-tolerant | the caller accepts data slightly behind | the user's own-write visibility | Most read paths satisfy both. The ones that do not are precisely the ones users notice: the re-read after a save, the confirmation screen, the poll immediately after a submit, the check-then-act read that decides whether to create something. Conflating the two properties means those paths get moved without review. ## What each policy actually costs **Default-on** - buys the maximum offload immediately, with no per-path work; - makes correctness depend on a marking that was never audited for freshness; - fails silently and intermittently, worst under load, when lag is largest; - concentrates the blast radius: one policy switch changes every read at once. **Opt-in** - keeps the primary as the default, so nothing changes correctness until someone opts a path in; - leaves offload unrealised until the work is done, which may be never; - drifts, because a new path copied from an opted-in one inherits a judgement nobody re-made; - needs a review habit, otherwise the marking spreads by imitation. ## The policy I would actually run 1. **Primary by default.** The safe direction, since the failure of being too fresh is a load problem you can see, and the failure of being too stale is a correctness problem you cannot. 2. **A separate, explicit stale-tolerant marking** that a path opts into, independent of read-only. Two markings, two promises, reviewable one at a time. 3. **Automatic pinning after a write**, request-scoped, so opted-in paths later in a writing request still land on the primary. This is what makes opting a path in a local decision rather than a whole-request analysis. 4. **A per-endpoint staleness budget**, stated in the endpoint's contract, so "how stale may this be" is answered where the product requirement lives. 5. **A lag-aware fallback**: when measured replication delay exceeds the budget, route back to the primary and accept the load rather than serve data outside the contract. 6. **Per-unit observability**: which datasource served each unit, and delay at read time. Without this the stale-read incident is unfalsifiable. ## What makes the trade-off decidable The honest input is the read mix. Offloading pays when a large share of reads are both heavy and genuinely stale-tolerant - reports, exports, search, dashboards, historical views. It pays very little when reads are small point lookups in interactive flows, where the primary was not the bottleneck and the freshness risk is highest. Measuring which bucket the traffic is in beats arguing about defaults, and often shows the whole benefit sits in a handful of endpoints - which is itself an argument for opt-in, since opting in five paths captures most of the win at a fraction of the risk. A second input is the shape of the lag. A fleet whose delay is a steady few milliseconds behaves very differently from one that spikes to seconds during batch work, and a policy chosen against the average will be wrong exactly during the spike. ## Failure modes to name explicitly - **Marking drift**: read-only spreads for its tracking benefit and silently enrols paths into replica routing. - **The pin that does not travel**: asynchronous continuations lose request scope and are routed away again. - **Cache ahead of routing**: a shared cache not keyed by datasource serves a stale copy no routing policy can intercept. - **Replica capacity as a hidden dependency**: once a share of reads has moved, the primary is no longer sized to take them back, so the lag-aware fallback needs headroom planned for, not assumed. ## Where the decision sits The repository method cannot know the freshness requirement; the endpoint does. So the durable form of this policy is that **each endpoint declares its staleness budget**, and the data-access layer implements exactly what is declared. That converts an argument about defaults into a per-endpoint contract that can be reviewed, tested and monitored - and it survives the arrival of engineers who never saw the incident that prompted it.
- What evidence would move you from opt-in to default-on?A read mix dominated by heavy, genuinely stale-tolerant traffic, a replication delay distribution whose tail stays inside the budgets endpoints have declared, per-unit routing observability already in place, and pinning proven to survive every asynchronous hop. Absent those, default-on is a bet that every past read-only marking was also a freshness judgement.
- How do you keep the two markings from collapsing back into one?Give them separate names and separate review expectations, and make the stale-tolerant one require a stated budget rather than being a bare flag. Then check it: an assertion in tests that a path with no declared budget never routes away from the primary keeps the distinction from decaying into imitation.
- What breaks when replication delay spikes and the fallback sends reads back?The primary suddenly absorbs traffic it has not carried for months. If capacity was quietly re-planned around the offload, the fallback becomes the outage. Headroom for the full read load has to be an explicit part of the policy, or the fallback must shed lower-value reads instead of returning all of them.
saying these in an interview costs you the question
- Treats read-only and stale-tolerant as the same property
- Switches routing on globally without auditing existing read paths
- Plans no fallback for when replication delay exceeds the budget
- Assumes the primary can always absorb the traffic back
- Leaves the freshness decision to whoever wrote the repository method