skip to content

A page nginx proxies takes 400 ms to generate and receives about 2,000 requests per second. Explain how an nginx microcache built on `proxy_cache_valid 200 1s` absorbs that load, and what still lets a burst of requests reach the upstream each time the entry expires.

level: seniorimportance: nice to knowfreq 42%

answer

  1. one second is a long time under load
  2. the average is flat, the edge is not
  3. a lock for new, stale for expired
  4. who serves the waiting requests
  5. background refresh needs two directives

basics

~20 s

A one-second cache collapses 2,000 requests per second per URL into roughly one upstream request per second. The gap is expiry: when the entry goes stale every concurrent request is sent upstream unless proxy_cache_use_stale updating, with proxy_cache_background_update, serves the old copy while one request refreshes it.

solid answer

~50 s

Microcaching trades one second of staleness for two orders of magnitude of origin load. With `proxy_cache_valid 200 1s`, all requests for that URL inside the window are HITs, so the backend sees roughly one request per second per key instead of two thousand — and since generation takes 400 ms, uncached traffic would need about 800 concurrent workers. Most dynamic backends emit `Cache-Control: no-cache` and a `Set-Cookie`, so you also need `proxy_ignore_headers` plus `proxy_hide_header Set-Cookie`, and the URL must be genuinely anonymous. The remaining hole is the expiry instant: nginx treats an expired entry as EXPIRED and sends *every* waiting request upstream. `proxy_cache_lock` does not help there — it only serialises requests for an entry not yet in the cache. What closes it is `proxy_cache_use_stale updating` with `proxy_cache_background_update on`: one request refreshes while everyone else is served the stale copy.

code

nginx · 18 lines
nginx
proxy_cache_path /var/cache/micro keys_zone=micro:10m max_size=1g inactive=1m use_temp_path=off;

location / {
    proxy_pass http://app;
    proxy_cache micro;
    proxy_cache_key "$scheme$host$request_uri";
    proxy_cache_valid 200 1s;

    proxy_ignore_headers Cache-Control Expires Set-Cookie;
    proxy_hide_header Set-Cookie;

    proxy_cache_lock on;
    proxy_cache_lock_timeout 5s;
    proxy_cache_use_stale updating error timeout http_500 http_502 http_503 http_504;
    proxy_cache_background_update on;

    add_header X-Cache-Status $upstream_cache_status;
}

go deeper

for a junior

Understand the basic trade: caching a dynamic page for even one second means the backend renders it once per second instead of thousands of times, at the cost of data up to a second old.

for a middle

Explain why a dynamic backend's own headers must be overridden for this to work at all, and name the directive pair that does it safely — proxy_ignore_headers together with proxy_hide_header Set-Cookie.

for a senior

Show that you know where the remaining load spike lives. Separate the cold-start case (proxy_cache_lock) from the expiry case (proxy_cache_use_stale updating with background update), and add the error and timeout parameters so the cache also covers origin failure.

for a principal

Decide whether the proxy tier is the right place for this at all. Weigh a one-second shared cache against splitting personalised fragments out of the page, moving the shortening to a CDN, or fixing the 400 ms render, and state what staleness the business has agreed to.

## The arithmetic At 2,000 requests per second with 400 ms of generation time, an uncached origin needs roughly 800 requests in flight simultaneously (Little's law: 2000 × 0.4). Almost no application server pool is sized for that, so the queue grows, latency climbs, and the pool collapses. Cache the response for **one second** and, for a single hot URL, the origin sees about one request per second — a 2,000-fold reduction — while no client ever sees data more than a second old. That is the whole idea of microcaching: it is not aimed at content that is safe to cache for ten minutes, it is aimed at content nobody thought was cacheable at all. ```nginx proxy_cache_path /var/cache/micro keys_zone=micro:10m max_size=1g inactive=1m use_temp_path=off; location / { proxy_pass http://app; proxy_cache micro; proxy_cache_valid 200 1s; proxy_ignore_headers Cache-Control Expires Set-Cookie; proxy_hide_header Set-Cookie; } ``` The `proxy_ignore_headers` line is not optional in practice: a dynamic page almost always declares itself uncacheable and sets a session cookie, so without it nothing is stored. And because you are ignoring `Set-Cookie`, you must hide it — otherwise the stored copy replays one visitor's session identifier to everyone. Microcaching is therefore only safe on URLs that render identically for every anonymous visitor; logged-in traffic must be excluded before it reaches this location. ## The gap at the expiry instant Once per second the entry becomes stale. nginx's default behaviour for a stale entry is to proxy the request — and it does that for *every* request that arrives while the refresh is in flight. With 2,000 rps and a 400 ms regeneration, that is roughly 800 simultaneous upstream requests, once a second, for exactly the URL you were protecting. The cache flattens the average and leaves a spike. Two directives are relevant, and they cover different cases: - **`proxy_cache_lock on;`** allows only one request at a time to populate a **new** cache element — one that is not stored yet. Others wait, bounded by `proxy_cache_lock_timeout` (default 5s) and released early by `proxy_cache_lock_age`, after which another request is allowed to try. This is the cold-start case: a fresh deploy, an evicted key, the first hit after a purge. It does **not** apply to an entry that exists and has expired. - **`proxy_cache_use_stale updating;`** covers the expiry case. While one request is updating an expired entry, everyone else is served the stale copy and logged as `STALE`/`UPDATING`. Combined with **`proxy_cache_background_update on;`** (nginx 1.11.10 and later), the refresh happens in a background subrequest so even the triggering client gets an immediate stale response rather than waiting 400 ms. Together: ```nginx proxy_cache_lock on; proxy_cache_lock_timeout 5s; proxy_cache_use_stale updating error timeout http_500 http_502 http_503 http_504; proxy_cache_background_update on; ``` The extra `error timeout http_5xx` parameters buy something beyond the herd: when the origin is failing outright, nginx keeps serving the last known-good copy instead of propagating the failure. For a one-second microcache that window is short, but during a deploy or a database stall it is often the difference between a degraded page and an error page. ## What it costs - **Staleness.** Bounded by the validity, so a second — usually acceptable for a homepage or a listing, never for a checkout confirmation. - **Applicability.** Only URLs with real concurrency benefit. A long tail of URLs each getting one request per second gains nothing and just churns the cache; keep `inactive` short so the tail evicts itself. - **Correctness risk.** Everything in the personalisation discussion above applies with less margin for error, because you deliberately disabled the header that would have protected you. - **Observability.** Watch the distribution of `$upstream_cache_status`. Healthy microcaching is overwhelmingly HIT with a thin trickle of UPDATING; a visible band of EXPIRED means the stale-while-updating path is not configured, and BYPASS means your exclusion conditions are firing more than you thought. ## When not to reach for it If the same content can safely live for a minute, cache it for a minute — microcaching is the answer when the content genuinely changes constantly and the load is concentrated. If personalisation makes a shared copy impossible, the fix is to split the page: cache the anonymous shell and fetch the personalised fragment separately, rather than pushing a per-user document into a shared store.

  • Why does proxy_cache_lock not solve the once-per-second stampede?
    Because it only governs a cache element that does not exist yet. It serialises the first population of a new key, so a cold cache or an evicted entry sees one upstream request instead of many. An entry that is present but expired is a different path: nginx proxies it as EXPIRED for every waiting request. The directive for that case is proxy_cache_use_stale updating.
  • What does proxy_cache_background_update add on top of proxy_cache_use_stale updating?
    Without it, the request that triggers the refresh waits for the upstream while the others are served stale. With it, nginx starts a background subrequest to refresh the entry and answers the triggering client from the stale copy immediately, so no user pays the regeneration cost. It requires the updating parameter on proxy_cache_use_stale to be meaningful.
  • How would you decide whether one second is the right validity?
    Work from the concurrency you need to remove and the staleness the content tolerates. Origin load falls to roughly one request per validity window per hot key, so even one second usually suffices when the arrival rate is high. Lengthen it only as far as the product accepts stale content; if a minute is acceptable, this is not a microcaching problem at all.

saying these in an interview costs you the question

  • Thinks proxy_cache_lock prevents the expiry stampede
  • Assumes a one-second cache is too short to help
  • Enables it on pages that render per-user content
  • Ignores Set-Cookie without hiding it downstream
  • Expects stale-while-updating without configuring use_stale

context