skip to content

Your fleet autoscales between 12 and 200 instances — how should that change the rule your teams use to choose an in-process copy?

level: principalimportance: should knowfreq 41%

answer

  1. instance count is now an input
  2. memory and divergence both scale out
  3. the rule must hold at peak
  4. frequency is not the criterion
  5. record the placement per read

basics

~10 s

Instance count becomes an input nobody fixes, so a per-instance copy's memory and its divergence both peak exactly when traffic does. A reviewable rule has to hold at 200 instances, not at 12.

solid answer

~50 s

With a fixed fleet, both costs of the in-process placement are constants somebody once checked. With an elastic one they are functions of a number the scaling policy chooses: memory is the copy's size times the current instance count, and the number of versions of an entry the fleet holds is also the instance count. Both peak at the moment of heaviest traffic, which is the worst moment to discover either. So the rule teams apply cannot be "this read is hot enough to hold locally" — frequency says the read is common, not that instances may answer differently. State it instead as two conditions that must both hold at the top of the scaling range: two callers being given different answers at the same instant is acceptable for this read, and the copy fits inside each instance's own memory limit with the request-handling heap intact. Everything else reads the shared tier and pays the hop.

go deeper

for a junior

Remember that scaling out adds a whole copy per new instance. The dataset does not grow, but the fleet's memory for it does.

for a middle

Be able to state both curves: bytes times instance count, and versions in flight equal to instance count, both peaking at heaviest traffic.

for a senior

Check the copy against the per-instance limit at the maximum fleet size, and make sure its bound is on memory rather than on a count of entries.

for a principal

Turn it into a rule a review can apply: disagreement acceptable for this read, and the copy fits every instance at peak. Record the decision per read, not per service.

## What autoscaling changes about the question With a fixed fleet, both costs of an in-process copy are constants that somebody checked once at design time: a known number of copies, a known amount of memory, a known number of versions in flight. Autoscaling turns the instance count into a parameter chosen at runtime by a policy that has no idea a copy exists. Every consequence of the placement is multiplied by that parameter, and the parameter is largest when the system is busiest. ## Two curves that peak together | As instances go from 12 to 200 | In-process copy | One shared tier | |---|---|---| | Memory held for the set | grows with instance count | flat: one copy | | Versions of an entry in flight | grows with instance count | one | | Request load on the tier | none | grows with instance count | | Per-read cost | a local lookup | a hop | Read across the table and the trade becomes a choice about *what you want to scale with the fleet*. The in-process placement scales memory and divergence; the shared placement scales request load on one component and leaves memory and agreement alone. Neither is free, and the elastic fleet makes the difference between them much larger than it looked at 12 instances: - Memory: a 500 MB copy is 6 GB of fleet memory at 12 instances and 100 GB at 200, for one unchanged dataset. - Divergence: at 200 instances there are up to 200 versions of an entry at once, so the probability that a given user's two requests land on two differently filled copies rises exactly when the fleet is widest. - Per-instance fit: the copy has to fit inside the instance limit that the scaling policy assumes when it decides how many instances a given load needs. A copy that grows with usage quietly reduces the headroom that policy was counting on. ## Stating the rule so it can be reviewed A rule of thumb that lives in people's heads will be applied at today's fleet size. Write it as conditions that are checkable in a design review: 1. **Disagreement is acceptable for this read.** Two callers, or one caller across two requests, being given different values at the same instant causes no incorrect decision and no visible anomaly. If that cannot be asserted plainly, the read does not go in a per-instance copy at any fleet size. 2. **The copy fits at the top of the scaling range.** Its bounded size plus the request-handling heap fits inside one instance's memory limit at peak, and the bound is enforced by something rather than hoped for. 3. **The bound is on bytes, not on entries**, wherever entry sizes vary by orders of magnitude — a count-based bound does not constrain memory. 4. **The placement is recorded with the read, not with the service.** One service can legitimately hold some reads in-process and take the hop for others; a blanket answer for the whole service is how both mistakes get made. ## What the rule must not be - **Not "is this read hot".** Frequency says the read is common. It says nothing about whether instances may answer differently, and that is the question the placement actually turns on. - **Not "does it fit today".** At 12 instances almost anything fits; the arithmetic that matters is the one at 200. - **Not "the shared tier adds latency".** That is an argument about the hop, and it is answered by measurement rather than by placement doctrine — but it is never a reason to put a read that must agree across instances into per-instance copies. ## What still has to be checked against the specific store The shared side of this comparison is a real piece of software with its own behaviour, and none of it should be assumed from the placement: - **At the memory ceiling, stores differ**: some reclaim entries to make room, others refuse the write and return an error to the caller. Whether a shared tier under pressure silently loses entries or starts failing writes changes what your callers experience at exactly the load where you scaled to 200. - **Across a restart, stores differ**: some retain something, others come back empty, and a design should say which behaviour it needs rather than assume the one it has seen. - **Connection load is part of the arithmetic.** Going from 12 to 200 instances multiplies the number of clients talking to one component, and that is a capacity question about the shared placement in its own right.

  • Does the shared placement have an equivalent problem as the fleet grows to 200?
    Yes, but a different one: its memory stays flat while the request and connection load from the fleet grows with the instance count. That makes it a capacity and concurrency question about one component rather than a memory question spread across the fleet — and it is a question you can answer by measuring one place instead of every process.
  • Why record the placement per read instead of per service?
    Because the criterion is a property of the read, not of the service. One service will legitimately hold a weekly reference value in-process and take the hop for a value users compare across requests. A blanket decision either pays the hop for reads that never needed agreement, or serves reads that do need it from copies that cannot provide it.
  • The team wants a per-instance copy bounded by entry count. What is the objection?
    A count does not bound memory when entry sizes vary. Ten thousand entries is a few megabytes of small records and many gigabytes if some entries are large documents, and the same bound can be comfortable on one day and fatal on another. Bound it in bytes, so the number you check against the instance limit is the number that actually gets allocated.

saying these in an interview costs you the question

  • Chooses the in-process placement because the read is hot.
  • Sizes the copy against today's instance count instead of the maximum.
  • Bounds the copy by entry count when entry sizes vary widely.
  • Applies one placement decision to a whole service rather than per read.
  • Assumes the shared tier behaves the same at the ceiling as the one they last used.