skip to content

A user's two requests hit different instances, each holding its own in-process copy, and disagree — which reads may still be served that way?

level: middleimportance: must knowfreq 72%

answer

  1. filled separately, at separate moments
  2. as many versions as instances
  3. routing decides which one you see
  4. may two callers differ right now?

basics

~20 s

Each instance filled its copy independently, so the fleet holds as many versions of the entry as it has instances and nothing in the arrangement reconciles them. Only reads where two callers getting different answers is harmless belong in that placement.

solid answer

~50 s

Every instance filled its own copy at its own moment, from whatever the system of record held then, and the copies are private structures in separate processes with no link between them — so the number of versions of one entry that the fleet can hold at once is the instance count. A load balancer is free to send one user's successive requests to two of them, which is exactly how the disagreement becomes visible. The test for whether a read belongs here is not how important the data is but whether two callers may legitimately be given different answers at the same instant: a reference value that changes weekly, or a figure nobody compares across requests, is fine. A value the user just influenced, or one another system will reconcile against, is not — that read wants one shared answer, which costs a hop and a dependency on the tier being up.

go deeper

for a junior

Remember that every instance keeps its own copy, filled at its own moment. Two requests answered by two instances can honestly return two different values.

for a middle

Explain the count: the number of versions the fleet holds is the number of instances, and the arrangement provides nothing that reconciles them. Then say which reads tolerate that.

for a senior

Show the production shape: it passes testing, appears only for users whose requests bounce between instances, and gets worse as you scale out under load.

for a principal

Make the criterion reviewable — whether two callers may legitimately be given different answers — so teams stop choosing the placement on read frequency alone.

## Why the two answers differ Each application instance holding an **in-process copy** filled it on its own: it asked the system of record at some moment, kept what came back, and has been answering from that ever since. Instance A filled the entry at 10:00:03; instance B filled the same entry at 10:04:41; the underlying value changed at 10:02. Both instances are behaving exactly as designed and they hold different values. Which one a caller sees is decided by whatever routed the request, and nothing routes a single user consistently to one instance unless someone deliberately made it do so. The load-bearing part is what is absent: **nothing in the arrangement reconciles the copies**. They are private data structures in separate processes with no channel between them. Making them agree after a change is a mechanism somebody has to design, build and operate; it is not something the placement gives you. ## How many versions of the answer exist The count is the instance count, and it moves with the fleet: - Ten instances can hold ten different values for one entry at the same moment; two hundred instances can hold two hundred. - Because each copy is filled independently, the spread between the oldest and the newest version is bounded by how long a copy is kept, not by how often the underlying data changes. - Instances that start later fill from current data, so a scale-out event can widen the disagreement rather than settle it. - A user routed repeatedly to one instance never notices. The same system therefore looks correct in testing and wrong in production, for the subset of users whose requests bounce between instances. - Adding instances to handle load means the arrangement is at its most divergent exactly when it is busiest. ## Which reads this placement can still serve The useful test is not "is this data important" but **"may two callers be given different answers at this instant without anyone being harmed?"** | The read | Per-instance in-process copy | One shared tier | |---|---|---| | A reference value that changes weekly | fine | fine | | A page of a list the user is stepping through | risky: consecutive pages can come from different copies | fine | | A value the user just changed and expects back | wrong: routing decides what they see | one answer, though not necessarily a current one | | A figure another system will compare against yours | wrong: your two instances do not agree with each other | defensible | ## Which reads it cannot serve - Anything a single caller will compare across two requests: a refresh, a paging click, a retry after a timeout. - Anything two instances will act on in a way that produces conflicting effects outside the process. - Anything a second system or a human operator will reconcile against, where "which instance answered" is not an acceptable explanation for a discrepancy. ## Three questions that settle the placement 1. May two callers legitimately be given different answers at the same instant? If not, the copy cannot be per-instance. 2. Will one caller compare two answers across requests? Routing turns that into a comparison between copies. 3. Does the size of the copy, times the largest instance count you will run, fit inside each instance's own memory limit? If not, the arrangement is unaffordable at peak however correct it is. ## What varies across stores, and what does not The divergence above is a property of the **arrangement** and holds whatever software is holding the copies. What does vary is the software itself: some products in this class run only as a server reached over the network, while others can be embedded inside the application process, so "in-process or shared" is sometimes a deployment choice within one product and sometimes a choice between two different pieces of software. What a store does when it runs out of room also varies — some reclaim entries, others refuse the write — so an answer that asserts one behaviour as the model is describing one store rather than the class.

  • Every entry has a short lifetime. Does that make the per-instance copies agree?
    No. A lifetime bounds how long any one copy can be wrong; it does not synchronise the copies with each other. Two instances that filled the same entry seconds apart also drop it seconds apart, so there is always a window in which they disagree. Shortening it narrows the window and raises the load on the system of record, but it never produces one answer.
  • If all instances read one shared tier instead, can a user still see a value move backwards between requests?
    Not because of the placement: every instance now reads the same entry, so successive requests return the same value whichever instance serves them. What the shared arrangement does not promise is that the entry matches the system of record — it makes the fleet equally right or equally wrong, which is why the next question is how current the value is, not who read it.
  • Does pinning each user to one instance fix the disagreement?
    It hides it for that user while the pin holds, which can be enough for a display value. It does not remove the divergence: the fleet still holds one version per instance, other systems and other users still see different answers, and any deploy, scale-in or instance failure moves the user to a copy filled at a different moment.

Ten cashiers each working from a price list they copied off the board at whatever moment their shift started, against ten cashiers who all glance up at the one board on the wall. The private lists are instant to read and quietly drift apart, so two customers in two queues are quoted two prices; the board gives everyone the same price at the cost of looking up, and if the board is covered nobody can price anything at all. Neither arrangement promises that the price is the one head office set five minutes ago.

saying these in an interview costs you the question

  • Blames the load balancer instead of the placement of the copies.
  • Says one copy must be corrupt because the values differ.
  • Assumes copies converge on their own because they hold the same data.
  • Treats a short lifetime as a way to make the copies agree.
  • Judges the placement by how hot the read is rather than by who may disagree.