skip to content

A gateway's p99 doubles nightly while it and a co-located batch job stay inside every processor and memory ceiling — what are they fighting over?

level: seniorimportance: must knowfreq 58%

answer

  1. limits certify only two axes
  2. everything else is undivided
  3. cached pages, queues, links, bandwidth
  4. the cost arrives as waiting
  5. own usage falls while latency rises

basics

~20 s

Everything the ceilings do not divide: one pool of cached file pages, one storage device and its queue, one network link, and the shared path to memory. Ceilings partition two axes and certify nothing about the rest.

solid answer

~40 s

The two declared ceilings partition exactly two things: the share of processor time each workload may consume, and the resident memory each may hold. A host has far more shared surface than that. Both workloads read through **one storage device with one queue**, transmit through **one network link**, pull through **one path to memory** and share the **processor cache** behind it, and both make the shared kernel do work on their behalf. A nightly batch that streams data fills the device queue and evicts cached pages, so the gateway's reads that used to be served from memory now wait on hardware. Neither exceeded anything it declared, which is exactly why 'we were inside our limits' is not an alibi here: limits certify the two accounted axes and say nothing about contention on the rest.

go deeper

for a junior

Take away one fact: the limits on a container cover processor time and memory only. The disk, the network and the cached file pages on that host are shared by everything running there.

for a middle

Be able to list the unpartitioned resources and say how each one is taken — a queue filled, a cache evicted, a link kept busy — and explain why the victim's usage graph stays flat while its latency climbs.

for a senior

Demonstrate the diagnostic reasoning: correlate the tail with the neighbour's schedule rather than your own traffic, explain why staying inside the ceilings proves nothing, and propose changes to the bulk workload rather than reaching for another limit.

for a principal

Treat this as the reason a density target needs a latency guardrail. Decide which workload classes may never share a device or a link, and be explicit that the cost of the wrong pairing lands on the tail of whichever class the business cares about most.

## What a ceiling actually divides A declared processor ceiling caps the share of processor time a workload may consume in each accounting period. A declared memory ceiling caps the resident memory it may hold. Both are enforced by the kernel's resource-accounting and limiting mechanism, and both are real. What they are not is a machine boundary. They divide **two** accounted quantities. Everything else on the host is shared with no divider at all, and a workload can consume an unlimited share of it while remaining, truthfully and provably, inside every number it declared. ## The shared surface no ceiling partitions - **Cached file pages.** The host keeps one pool of file data in memory. A workload that reads a large file once fills that pool with pages nobody will read again, pushing out a neighbour's frequently read pages. The neighbour's reads then go to the device. - **Storage device bandwidth and queue depth.** One device, one queue. A batch job that issues deep, parallel reads puts its requests in front of a latency-critical workload's small ones. Waiting behind a deep queue is measured in milliseconds; being served from memory is measured in microseconds. - **Network link and transmit queue.** A bulk transfer fills the outbound queue, and a small request's packets wait behind it. The link is not oversubscribed in bytes per second; it is oversubscribed in *queueing delay*, which is what latency-sensitive traffic actually feels. - **Memory bandwidth and the shared processor cache.** Even on separate cores, two workloads pull through the same path to memory and share a level of processor cache. A streaming workload evicts a neighbour's working set out of that cache, so the neighbour executes fewer instructions per cycle for the same work — its *throughput per unit of processor time* falls without its usage changing. - **Kernel work done on a co-tenant's behalf.** Interrupt handling, packet processing, filesystem work and the contention on shared kernel data structures are all done on one shared kernel, and the cost of doing them for a busy neighbour is frequently not charged anywhere useful. ## Why the victim's own metrics look innocent This is the part that makes the symptom hard to believe. The gateway's dashboards show **processor usage flat or falling** and **memory well under its ceiling**, while its tail latency doubles. That is not a measurement failure; it is what the mechanism predicts: 1. The cost lands as **waiting**, not as executing. Time blocked on a device read or a transmit queue is not time on a processor, so it does not appear as usage. 2. The slower the gateway gets, the *less* processor time it consumes per second, because more of each request is spent blocked. 3. Memory usage does not move either, since the thing that was taken away — cached pages — is not counted as the workload's own resident memory in the way its ceiling is. So the shape to recognise is: **latency up, own usage flat or down, no ceiling exceeded, and the timing correlates with someone else's schedule rather than with your own traffic.** | Resource | Divided by a declared ceiling? | How a neighbour takes it | What the victim sees | |---|---|---|---| | Processor time | Yes | Cannot exceed its own share | Throttling, if anything | | Resident memory | Yes | Cannot exceed its own share | Its own process ended, if anything | | Cached file pages | No | Reads a large file once | Reads that used to hit memory now hit the device | | Device queue depth | No | Issues deep parallel reads | Small reads wait behind a long queue | | Network queueing delay | No | Sends a sustained bulk transfer | Small requests delayed, no packet loss needed | | Memory bandwidth and processor cache | No | Streams through a large working set | Same work takes more processor time | ## What 'we were inside our limits' is worth It is a true statement about two accounted axes and no statement at all about contention. Both workloads can be simultaneously innocent by their own accounting and jointly responsible for the tail. The reasoning an interviewer wants is that you **stop treating the ceilings as proof** and start asking which unpartitioned resource moves on the same schedule as the symptom. Two consequences follow that are worth saying out loud. - **The effect compounds.** Once the gateway's cached pages are gone, it starts issuing device reads of its own — into the same queue the batch job is already filling. The victim becomes an additional source of the pressure hurting it. - **The cure is rarely another limit.** The productive levers are about the shared resources themselves: run bulk work when the interactive class is quiet, cap the concurrency and read-ahead of the bulk workload at the application level, and keep classes with incompatible demands off the same hosts. Deciding *where* workloads land is a placement question, but recognising that the two classes should not share a device is this reasoning. The one-line answer: **ceilings divide processor time and memory; they divide nothing else, and a co-tenant's tail cost is paid entirely out of everything else.**

  • Why does the victim's own processor usage often go down rather than up while it is being hurt?
    Because it is blocked, not busy. Time spent waiting on a device read, a transmit queue or a memory stall is not time spent on a processor, so usage measures it as idleness. The requests take longer and each one consumes less processor time per second of wall clock, which is why a falling usage graph beside a rising tail is a strong hint at contention rather than at load.
  • What would you change so this stops happening, without buying more hosts?
    Attack the shared resource rather than adding another ceiling. Move the bulk work out of the window where the interactive class is busy, cap its read and transmit concurrency in the application so it cannot fill a queue, have it signal that one-pass data need not be cached, and keep a class whose tail is the product off hosts that run bulk work at all. Only the last is a placement decision; the rest are workload changes.
  • How would you demonstrate the co-tenant is the cause rather than merely coincidental?
    Correlate the symptom with the neighbour's schedule rather than with your own traffic: the tail should move when the batch starts and stops, not when your request rate changes. Pausing or rescheduling the batch for one night is the decisive test, and a host-by-host comparison between hosts that carry the batch and hosts that do not gives the same evidence without touching production behaviour.

saying these in an interview costs you the question

  • Concluding that because both stayed under their limits, neither can be the cause
  • Believing a processor ceiling prevents one workload from slowing another down
  • Assuming device and network bandwidth are divided between containers like processor time
  • Thinking each container has a private file cache that a neighbour cannot evict
  • Insisting that a rising tail must mean some limit was exceeded
  • Reading flat processor usage on the victim as proof it is not resource-starved