skip to content

Your in-memory store reports a 95% hit ratio while the calling service's own metrics show 70% - how can both be right?

level: middleimportance: must knowfreq 64%

answer

  1. a ratio needs a vantage point
  2. different denominators, same traffic
  3. the server never sees what never arrived
  4. cumulative counters hide a live collapse
  5. the gap between the two is itself signal

basics

~10 s

Both figures are right because they count different populations of requests: the server counts only lookups that reached it, while the caller also counts lookups that never arrived at all.

solid answer

~50 s

A hit ratio is meaningless without a vantage point, because the server and the caller are not counting the same requests. The **server-side** figure counts only lookups that actually arrived: it cannot see a lookup the caller answered from its own in-process copy, one that failed while waiting for a connection, or one that timed out before the reply came back - and on a partitioned deployment each node counts only its own share. The **caller-side** figure counts every lookup the application made, and typically scores a timeout or a connection failure as a miss, because the application had to go to the origin either way. The gap between the two numbers is therefore itself a signal - it is roughly the traffic that never completed a round trip - and neither number is per-population unless somebody labelled it that way.

go deeper

for a junior

Remember that a hit ratio is only meaningful with a vantage point attached, and say which one you mean whenever you quote one. 'As counted at the server' and 'as counted by the caller' are different numbers from the same traffic.

for a middle

Explain the mechanics: which requests each counter includes, why the denominators differ, and why counters that accumulate from process start must be read as a change over a window rather than as a level.

for a senior

Demonstrate that you keep both vantage points and treat the gap as its own signal, and that you label the caller-side ratio per population instead of trusting one instance-wide average across unrelated workloads.

for a principal

Make the vantage point part of the operating contract: which ratio each team is accountable for, where it is computed, and the rule that no instance-wide average is ever used to judge a single team's workload.

## A ratio is not a number, it is a number plus a vantage point A hit ratio is the fraction of lookups that found an entry. Every part of that sentence is ambiguous until you say **who counted** and **which lookups**. The same traffic produces at least three defensible ratios, and quoting one without its frame is the single most common way this signal misleads. | Vantage point | Which lookups it counts | What it cannot see | |---|---|---| | **As counted at the server** | every lookup that arrived at this server process | lookups answered by an in-process tier in the caller and never sent; lookups that died waiting for a connection; lookups that timed out before the reply arrived; lookups that went to a different node | | **As counted by the caller** | every lookup the application code made, including ones that never completed | which server, which node, or which population served it, unless the caller labelled that | | **Per logical population, wherever computed** | one named group of entries - sessions, rendered pages, lookup results | everything outside that group, which is exactly the point of computing it | ## Why 95% and 70% can both be honest Work through the arithmetic of the gap. Suppose the application makes a hundred lookups: 1. **Twenty-five never reach the server.** They fail while waiting for a connection from the caller's pool, or they time out, or a short-lived in-process copy answered them. The caller records twenty-five outcomes the server has no record of. 2. **Seventy-five arrive.** Of those, roughly seventy-one find an entry - the server divides seventy-one by seventy-five and reports about 95%. 3. **The caller divides by a hundred.** Whatever the caller does with the twenty-five - most commonly scoring them as misses and going to the origin - its ratio lands near 70%. Neither party is wrong. They are dividing by different denominators, because they are watching different populations of requests. ## Other gaps between the same two numbers - **The window differs.** Many of these counters are cumulative from process start, so a ratio computed from them is a lifetime average that keeps looking fine for hours after a live collapse. Read them as a change over a window, not as a level. - **The mix differs.** A single server-side ratio blends every caller and every population behind one instance. One team's sessions, which should essentially never miss, can hold the number up while another team's lookups miss constantly. - **The nodes differ.** On a partitioned keyspace, each node reports its own ratio. Averaging those gives the mean of the nodes' ratios, not the ratio of the traffic - the two are only equal when every node takes identical volume. - **The definition differs.** Whether a lookup that found an entry whose deadline had already passed is scored as a hit or a miss, and whether operations other than plain lookups count at all, are per-store properties rather than universal ones. ## Where stores in this class differ Not every store in this class publishes these counters, and among those that do, what is inside them varies - some count every read-shaped operation, some count only simple lookups, and some expose the pair only per logical grouping of entries. Some expose nothing and leave the caller as the only vantage point that exists. Before comparing a figure from one deployment to a figure from another, confirm that both count the same events; otherwise you are reading a difference in definitions as a difference in behaviour. ## Reading the signal without being misled - **Always say the frame out loud**: "hit ratio as counted at the server" or "as counted by the caller". A bare percentage in an incident channel costs ten minutes every time. - **Keep both**, and watch the gap. A widening gap with both ratios otherwise stable means more traffic is failing before it completes a round trip, which is a connection-side story rather than a content story. - **Label the caller-side ratio by population.** The instance-wide number is the least actionable of the three, because it is an average over unrelated workloads. - **Alert on the change, not the level.** A ratio that falls by a third within a few minutes is a signal regardless of where it started. What value a given design *should* see, and what to do about a poor one, is a cache-design question rather than a signal-reading one - the number tells you something moved, not what the target was.

  • The server-side and caller-side ratios have been stable for weeks, and now the gap between them is widening while both stay flat. What does that suggest?
    That a growing share of lookups is failing before it completes a round trip - dying in the caller's pool, timing out, or never being sent - since those are exactly the requests the caller counts and the server cannot. The content of the tier has not obviously changed; the path to it has. Which part of that path is a separate investigation.
  • Why is the mean of ten nodes' hit ratios not the hit ratio of a partitioned keyspace?
    Because each node's ratio is weighted by its own traffic, and averaging the ratios throws that weighting away. A quiet node with a poor ratio drags the mean down out of proportion to the requests it served, and a hot node's good ratio is counted once rather than in proportion. To get the traffic's ratio you must sum hits and sum lookups across nodes, then divide.

saying these in an interview costs you the question

  • Quotes a hit ratio without saying who counted it
  • Assumes the server can see lookups that timed out before arriving
  • Averages per-node ratios to get the ratio of a partitioned keyspace
  • Reads a cumulative-since-start ratio as the current one
  • Treats one instance-wide ratio as describing every caller behind it