Why does a route whose HTML is rendered per request cost more server work as its traffic grows?
answer
- cost per view, not per page
- nothing amortised between visitors
- upstreams feel it too
- in-flight slots scale with response time
- arrival rate times duration
basics
~20 sEvery view is its own render: one route match, one set of data reads, one tree rendered to a string, one serialisation. That work is repeated per visitor rather than amortised, so compute, upstream load and concurrency all track traffic almost linearly.
solid answer
~40 sThe unit of cost is the view, not the page. Serving a body that already exists is mostly bytes off a disk or a cache; producing one per request means running the router, the route's data work, the renderer and the serialiser again for every visitor, including two visitors who would have received identical HTML. Three resources scale with that: CPU on the rendering tier, load on every upstream the route reads, and concurrency — each in-flight render holds a slot for as long as its response is held, so slow upstreams raise the number of simultaneous renders even at constant traffic. The usual lever is to stop repeating identical work: a cache in front collapses many views into one render, but only for responses that are genuinely the same for everyone.
go deeper
Hold on to the core fact: the server does the whole job again for every visitor, so twice the visitors means roughly twice the work, even when the output is identical.
Name all three growing resources — compute, upstream reads and concurrent in-flight renders — and explain why a held response ties up a slot for its entire duration.
Show operating instinct: watch in-flight count alongside CPU, expect the shared data source to fail before the stateless tier, and prefer shedding load to queueing without limit.
Reason about what the marginal cost per view buys. Decide which routes deserve a per-view spend, and put a guardrail in place so hot routes cannot quietly acquire one.
## The unit of cost is one view When HTML is produced per request, nothing is shared between two visitors. Each view pays for the full chain again — route match, the route's data work, rendering the component tree to a string, serialising the data for the client, compressing and writing the response. Ten thousand views of the same URL means ten thousand renders, even if all ten thousand outputs are byte-identical. Compare that with serving a body that already exists somewhere. That path is dominated by I/O and bandwidth: the marginal cost of the next view is transferring the bytes again, not manufacturing them. | Resource | Serving an existing body | Rendering per request | |---|---|---| | CPU per view | Small and roughly constant | Proportional to tree size and work in the render | | Upstream load | None after the body exists | One set of reads per view unless deduplicated | | Memory | Buffers for the transfer | Live objects for every in-flight render | | Concurrency | Bounded by bandwidth | Bounded by workers or slots held for the response duration | | Effect of a traffic spike | More bandwidth | More compute *and* more load on the data sources | ## The three things that actually grow 1. **Compute.** Rendering is CPU-bound work: building a tree, walking it, producing and escaping strings. Doubling the views doubles that work. Because rendering usually cannot be interrupted halfway, saturation appears as a queue in front of the renderer, and queueing time is added to every visitor's wait, not only the ones over the line. 2. **Upstream pressure.** The route's data reads happen per view too, so a popular page turns into sustained load on the databases and services behind it. This is often the part that breaks first: the rendering tier scales out easily, the shared database behind it does not. 3. **Concurrency.** This is the term people forget. A held response occupies a slot for its whole duration, so the number of simultaneous renders is roughly *arrival rate multiplied by response time*. If an upstream slows from 100 ms to 1 s at unchanged traffic, in-flight renders grow tenfold and the tier can exhaust its workers without any increase in views. ## What changes the arithmetic - **A cache in front.** A shared cache or CDN serves many views from one render, which is the single biggest lever available. It only applies where the response is identical for every recipient and the headers permit storing it, which per-user output rules out. - **Caching inside the render.** Memoising an expensive read or fragment across requests lowers upstream load and CPU without changing the mode, at the cost of a freshness window you now have to reason about. - **Deduplicating within one render.** Two components asking for the same record should produce one upstream call per request, not two. - **Shrinking the render itself.** Paginating long lists, avoiding expensive formatting per row, and keeping the serialised payload small all reduce a per-view cost that repeats forever. - **Bounding the work admitted.** Limits on concurrency and queue depth turn a would-be collapse into shed load, which is usually the better failure. ## Where the hosting shape changes what you notice On a long-lived server process you see the cost as CPU utilisation and a queue: the machine is always there, and it degrades gradually until it does not. On a platform that starts a short-lived process per request, the same cost appears as instance count and billed execution time, so the arithmetic becomes explicit on the invoice, and cold starts add a fixed penalty to traffic that arrives after an idle period. On an edge runtime the renders spread across many locations, which helps compute but multiplies the number of places hitting the data source. The per-view property is identical in all three; only the way it shows up in monitoring and billing changes. ## The judgment this is really about The honest framing is not that per-request rendering is expensive, it is that **it has a marginal cost per view that other output paths do not**. That cost buys freshness and personalisation. It is well spent on a page whose content genuinely differs per visit and wasted on a page that renders the same bytes a million times, which is why the first question a reviewer should ask about a hot per-request route is what part of its output actually varies. What to do once you know the answer — leaving the mode, caching it, or splitting the varying part out — is a design decision beyond the mechanics described here.
- Why can in-flight renders grow without any increase in traffic?Because concurrency is arrival rate multiplied by how long each response is held. A dependency slowing down lengthens every render, so more of them overlap at the same arrival rate. That is how a slow upstream exhausts workers on a tier whose request volume never moved.
- Which tier usually breaks first under a traffic spike on a per-request route?The shared data sources behind it. The rendering tier is typically stateless and can be scaled horizontally in minutes, while the database or upstream service the route reads on every view is shared, harder to scale, and now taking one read per visitor.
- Does a cache in front remove the per-view cost?It collapses many views onto one render for as long as the stored copy is valid, so the marginal cost per view drops sharply. It does not apply where the output differs per user, and it introduces a staleness window, so it converts a cost problem into a freshness decision rather than eliminating it.
saying these in an interview costs you the question
- Thinks adding rendering machines fixes upstream load
- Ignores concurrency and counts only CPU per render
- Assumes a cache in front works for personalised pages
- Believes cost tracks unique pages rather than views
- Treats a cold start as the main per-view cost