skip to content

A self-hosted Next.js app running `next start` in a container shows CPU spikes and slow image responses for the first minutes after every deploy, then settles down. How does the built-in image optimizer explain that pattern, and what levers do you have?

level: seniorimportance: should knowfreq 44%

answer

  1. transforms happen on request, not at build
  2. a fresh container starts cold
  3. widths times qualities times formats
  4. the cache lives on disk under .next/cache
  5. AVIF encoding is the expensive one

basics

~20 s

next/image transforms images on request and caches them on disk under .next/cache/images, so a fresh container recomputes every size and format variant. Cut the number of variants, persist or share the cache, or move transforms to an external loader.

solid answer

~50 s

The optimizer is a request-time service, not a build step: the first time anyone asks for a given source URL at a given width, quality and output format, your server decodes and re-encodes the image, then writes the result to the on-disk cache under `.next/cache/images`. A newly deployed container starts with that cache empty, so early traffic pays for every transform at once — hence the spike that fades as the cache warms. Multiply that by every distinct width in your configured size lists, every quality value, and both output formats, and one source image can be dozens of transforms. The levers are: emit fewer variants (accurate `sizes` so huge widths are never requested, trimmed `images.deviceSizes` and `images.imageSizes`, one quality value), keep the cache across restarts by mounting it on a volume or putting a shared CDN in front of `/_next/image`, raise `images.minimumCacheTTL`, mark already-optimized assets `unoptimized`, or hand transformation to an external service through a custom loader.

go deeper

for a junior

Know that image optimization runs on the server when the image is requested, not during the build, and that the results are cached rather than recomputed every time.

for a middle

Explain the cache key — source URL, width, quality, output format — and why a fresh deploy means a cold cache and a burst of expensive transforms.

for a senior

Diagnose before tuning: confirm the slow route is the optimizer, inspect which widths and qualities are actually requested, check whether the cache survives restarts and whether replicas share it, then pick levers with their trade-offs named.

for a principal

Own the strategy: decide whether image transformation belongs in the app process at all, weigh a persistent or shared cache against an external image service, and set the URL-versioning convention that lets you cache aggressively without stale content.

## Where the work happens This assumes a current App-Router Next.js (Next 16) self-hosted with `next build` and `next start`. Nothing about image optimization happens during `next build`. `/_next/image` is a runtime route, and each request to it is parameterised by four things: the source URL, the width `w`, the quality `q`, and the output format chosen from the request's `Accept` header. The first request for a given combination performs a real decode-and-re-encode using the native image library bundled with Next; the result is written to disk under `.next/cache/images` and served from there afterwards. That design explains the symptom exactly. A fresh container has an empty cache directory, so early traffic is a burst of cold transforms — CPU-bound, on the same process serving your pages — until the popular combinations are all cached, after which the endpoint is essentially a file read. ## The multiplication problem Count the combinations a single hero image can produce: - one entry per candidate width Next generates from `images.deviceSizes` (viewport-scale widths) and `images.imageSizes` (small fixed sizes); - times each distinct `quality` value used anywhere in the codebase; - times each output format actually negotiated — AVIF where enabled, WebP, and the original for browsers that take neither. AVIF is the sting: it typically produces smaller files than WebP but costs substantially more CPU to encode. A page of forty product photos, each with a full width ladder, is a large number of cold transforms concentrated in the minutes after deploy. ## Confirming the diagnosis before changing anything - Check whether the slow requests are `/_next/image` specifically, and whether their latency is high on first hit and low afterwards for the same query string. Cold-cache behaviour is distinctive. - Look at how many distinct `w` and `q` values appear in your access logs. A long tail of widths usually means missing or careless `sizes` props. - Inspect the size and contents of `.next/cache/images` in a running container, and check whether that path survives a restart or is inside the ephemeral layer. - Note the replica count. With N replicas and no shared cache, the same transform is performed up to N times. ## The levers, roughly in order of leverage **Generate fewer variants.** Accurate `sizes` values stop the browser from ever requesting a 3840px candidate for a card thumbnail. Trimming `images.deviceSizes` and `images.imageSizes` to the breakpoints your design actually uses shortens the ladder directly. Standardising on a single `quality` value across the codebase collapses a whole axis of the product. **Stop losing the cache.** Mount `.next/cache` (or at least the images subdirectory) on a persistent volume so the cache survives redeploys and restarts. This is the single change that removes the post-deploy spike, though be deliberate about eviction — the directory grows. **Put a shared cache in front.** A CDN or reverse proxy caching `/_next/image` responses means a given variant is transformed once for the whole fleet rather than once per replica, and repeat requests never reach your origin. `images.minimumCacheTTL` raises the floor on how long a transformed result is considered fresh, which matters when the upstream response carries a short cache lifetime. **Skip the optimizer where it adds nothing.** Assets that are already correctly sized and encoded — a small logo, a pre-generated sprite, an SVG — gain little from a transform. The `unoptimized` prop renders the original URL and removes those requests entirely. **Move the work off the app process.** A custom loader (`images.loader: 'custom'` with `images.loaderFile`, or the per-component `loader` prop) makes you generate the image URL yourself, pointing at a dedicated image service. Your Node process stops doing encoding altogether. You give up Next's built-in negotiation and take on the external service's own cost model and access controls. ## Trade-offs worth naming out loud Every lever trades something. Trimming the width ladder means some devices get a slightly imperfect fit. Longer TTLs mean a replaced image at the same URL takes longer to propagate — content-addressed or versioned URLs are the usual answer. `unoptimized` gives up the responsive srcset entirely, so it is only correct when the asset is already right. A persistent cache volume needs a growth and eviction policy, and a shared volume across replicas needs to tolerate concurrent writers. What should not be on the list is "turn optimization off globally" as a first move: that trades a CPU problem for a bandwidth-and-LCP problem for every user. Measure which variants are actually being generated first — the fix is usually that you were producing far more of them than the design needs.

  • Why does adding replicas make the post-deploy spike worse rather than better?
    Because each replica keeps its own on-disk cache. The same width-and-format variant is transformed independently on every instance it is requested from, so total CPU work scales with replica count instead of being shared. A cache in front of the origin, or a shared cache volume, is what converts N cold transforms back into one.
  • You suspect too many variants are being generated. How do you confirm it rather than guessing?
    Look at the distinct `w` and `q` values hitting `/_next/image` in your access logs, and compare that set against the widths your design actually renders at. A long tail of large widths on small components points straight at missing or inflated `sizes` props; several quality values point at inconsistent usage across the codebase. Then trim the configured size lists to match reality.
  • What breaks if you raise `images.minimumCacheTTL` aggressively?
    Replacing an image at the same URL stops propagating promptly — cached transformed variants keep serving the old bytes until the TTL lapses. The standard mitigation is to make image URLs content-addressed or versioned, so a new image is a new URL and cache lifetime becomes irrelevant to correctness. Without that, long TTLs and editable images conflict.

saying these in an interview costs you the question

  • Thinks next build pre-generates all image variants
  • Assumes the cache survives a container replacement by default
  • Reaches straight for images.unoptimized globally
  • Ignores that each replica maintains its own cache
  • Believes AVIF is free because the files are smaller

context