skip to content

A service's reads were hitting the host's cached file pages until a co-tenant began streaming a large file — why did they get slower?

level: seniorimportance: nice to knowfreq 31%

answer

  1. one host, one cache pool
  2. read once, never read again
  3. hot pages pushed out
  4. a ceiling caps, it does not reserve
  5. microseconds become milliseconds

basics

~20 s

One host keeps one pool of cached file pages. A single pass over a large file fills it with pages nobody will read again, pushing the service's hot pages out, so its reads now reach the storage device.

solid answer

~40 s

Cached file pages are a host-wide pool, not a per-container allocation. A sequential stream touches each page once and never returns to it, but the cache generally cannot tell a page that will be reused from one that will not, so a large one-pass read displaces a neighbour's hot pages. The neighbour's reads then reach the storage device: microseconds become milliseconds, and the effect compounds because those new device reads join the same queue the streamer is already filling. Its memory ceiling never helped, because a ceiling is an upper bound on what the workload may hold, not a guarantee of what is kept for it — and cached pages are the first thing a host gives back under pressure.

go deeper

for a junior

The key fact is that the memory used to cache file contents belongs to the whole host, not to one container. Another workload's reading can push your data out of it.

for a middle

Explain why a one-pass sequential read is uniquely destructive to a cache built on recency, and why a declared memory ceiling — being a cap rather than a guarantee — cannot protect anything from that.

for a senior

Recognise the signature: a latency step change tied to a neighbour's schedule, with the victim's own usage unchanged. Then reason about the compounding device queue and propose fixes in the bulk reader rather than new limits on the victim.

for a principal

Ask the design question underneath: which services have a latency commitment that silently depends on shared, evictable memory, and whether that dependency is acceptable at the density the fleet is targeting.

## One host, one pool of cached pages When any process on a host reads a file, the host keeps those pages in memory so the next read of the same data does not touch the device. That pool is **one pool for the whole host**. Containers on the host do not each get a private one; the boundary around a container gives it separate views of the filesystem, the process table and the network, but the memory used to cache file contents is host-wide and reclaimable. This is what makes cached pages the sharpest example of a resource that no declared ceiling divides. ## Why a one-pass reader is the worst tenant for it A workload that streams a large file has a peculiar access pattern: - It touches an enormous number of distinct pages. - It touches each of them **exactly once**. - It derives almost no benefit from any of them being retained. A cache keeps what has been used recently, on the assumption that recent use predicts future use. That assumption is exactly backwards for a one-pass stream: every page it brings in is the *least* likely page on the host to be read again, and yet it is the most recently used. So a single pass over a file larger than the pool will, in the general case, walk the entire neighbour's working set out of memory and replace it with data that will never be touched again. The result on the neighbour is abrupt rather than gradual. Its reads were being answered from memory; now they are answered by the device. | Where a read is answered | Typical cost | What changed | |---|---|---| | Cached pages in host memory | Microseconds | Nothing — the steady state | | The storage device, queue empty | Sub-millisecond to milliseconds | The page was evicted | | The storage device, queue deep | Many milliseconds | The page was evicted and the co-tenant is also queueing reads | ## Why the ceiling did not help Three separate reasons, and candidates usually have only the first: 1. **A ceiling is a cap, not a reservation.** It is the maximum a workload may hold. Raising it does not cause anything to be kept for that workload, so a higher memory ceiling protects no cached page. 2. **Cached pages are reclaimable, and reclaim is host-wide.** Because a cached page can be re-read from the device, it is the cheapest thing for a host under pressure to give back — which means it goes first, before anything drastic happens to any process. 3. **The neighbour is not using more of its own memory.** Its resident memory never moved, so nothing in its accounting changed. Everything the workload owned it still owns; what it lost was something it never owned in the first place. ## The second-order effect The damage is not one step. Once the service's pages are gone, it begins issuing device reads it was not issuing before — and those reads join the very queue the streaming co-tenant is already keeping deep. So the victim's latency is hit twice: once for having to reach the device at all, and again for waiting behind the workload that evicted it. This is why the symptom is usually described as a step change at a fixed time of night rather than a gentle drift. ## What actually helps - **Have the bulk reader declare its intent.** A reader can signal that data it streams once need not be retained, so its pages are dropped rather than left occupying the pool. Where that is available it is the cheapest fix by a wide margin. - **Cap the bulk reader's read-ahead and concurrency** in the workload itself, which limits both how fast it can flood the pool and how deep it makes the device queue. - **Move the bulk work out of the interactive window.** The cheapest contention to fix is the kind that simply does not overlap. - **Keep the two classes off the same hosts.** Which host a workload lands on is a placement decision, but recognising that a cache-dependent service and a large sequential reader are a bad pair is this reasoning. - **Make the service less dependent on a cache it does not control.** Anything whose latency budget assumes a cache hit on shared, evictable memory has a dependency on its neighbours' behaviour, and that is worth knowing before an incident rather than during one. ## The shape to recognise A latency step change that lines up with a neighbour's schedule, on a service whose own processor usage and resident memory did not move, whose read latency jumped by roughly the distance between memory and a storage device. That combination is a cache that was taken away, and nothing in the workload's own declared numbers will ever show it.

  • Would giving the affected service a larger memory ceiling protect its cached pages?
    No. A ceiling is an upper bound on what a workload may hold, not a floor that is kept for it, and cached file pages live in a host-wide pool rather than inside the workload's own allocation. Raising the number changes nothing about what the host reclaims when it needs memory back. What helps is preventing the eviction — through the bulk reader's behaviour or by not sharing the host.
  • Why does the victim often end up making the contention worse?
    Because eviction converts its memory reads into device reads. Requests that never touched storage now queue against the device, and that queue is already deep because the streaming co-tenant is filling it. The victim therefore adds to the pressure that is hurting it, which is why the latency step is usually larger than the plain memory-to-device gap would suggest.

saying these in an interview costs you the question

  • Believing each container has a private file cache a neighbour cannot evict
  • Claiming a higher memory ceiling would keep the service's cached pages resident
  • Assuming a one-pass sequential read is harmless because it never re-reads anything
  • Concluding that unchanged memory usage rules memory out as the cause
  • Thinking the cache keeps pages for whichever container is paying for them
  • Expecting a cache to recognise on its own which pages will never be reused