When the out-of-band render of a prerendered page fails, what is served next, and why can that failure go unnoticed?
answer
- a failed publish publishes nothing
- the old copy keeps its job
- success status, frozen content
- alert on copy age, not status codes
basics
~20 sThe last good copy stays in place, because a failed render is not published. Visitors keep getting fast, valid, steadily older HTML with a success status, so only render telemetry and the age of the stored copy reveal that anything is wrong.
solid answer
~50 sA refresh is a publish, and a failed publish publishes nothing. If the render throws, its data source is down, or it never completes, implementations typically leave the existing stored copy untouched rather than replacing it with an error page or clearing it — which is the right default, because discarding the last good HTML would turn a refresh problem into an outage. The consequence is that the failure is **invisible from the outside**: responses stay fast, carry a success status and contain a perfectly valid page, and the only thing that changes is that the content stops moving. Nothing in request metrics moves either. You detect it by emitting the outcome of every render attempt and by tracking the age of each stored copy, and by alerting on age rather than on status codes.
go deeper
Remember the safe behaviour: if the new render does not work out, the page you already had keeps being served. Visitors see a working page, just an older one than they should.
Explain why the response looks healthy — a stored copy is being served successfully — and therefore why the age of that copy, not the status code, is the thing that signals a problem.
Demonstrate the operational answer: emit render outcomes, track per-path copy age, alert on age, and be able to say which surfaces should stop trusting a very old copy instead of continuing to serve it.
Name the risk in business terms. This model fails by serving content that is quietly wrong rather than by going down, so the staleness ceiling and who is accountable for it belong in the surface's definition, not in a config default.
## A refresh is a publish The out-of-band render has one job: produce a complete document and replace what is stored for that path. If it does not produce a complete document, there is nothing to publish, and implementations typically take the conservative branch — **keep what is there**. The reasoning is straightforward: - The stored copy is known to be renderable output; the failure says nothing about its validity, only about its age. - Replacing it with a failure would convert a background problem into a user-visible outage on a page that was working a second ago. - Deleting it would push every subsequent request down the no-copy path, which is strictly worse than serving something slightly old. So the failed render is discarded, the old copy stays in service, and the path is refreshed again the next time it is due. ## What the visitor sees Nothing. Specifically: - **The same fast response**, served from storage with no data access. - **A success status**, because a stored page is being returned successfully. - **Valid, complete HTML** — last week's HTML, but structurally fine. - **Content that has stopped changing**, which only a human comparing it against reality can notice. This is the asymmetry that makes the failure mode dangerous. Rendering inside the request fails loudly: the visitor gets an error and your error rate moves within seconds. Refreshing out of band fails silently, and the damage accumulates as content drift rather than as downtime. ## Why ordinary monitoring does not catch it | Signal | Does it move when refreshes fail? | |---|---| | Error rate for the path | No — stored copies are served successfully | | Time to first byte | No — possibly faster, since no render competes for CPU | | Request volume | No — demand is unrelated to the copy's age | | Uptime and health checks | No — the application is serving correctly | | Age of the stored copy | **Yes** — it grows without bound | | Render attempts succeeded versus failed | **Yes** — directly, if you emit it | The two rows that move are the two nobody instruments by default, which is why this failure is routinely discovered by a person noticing that a page still shows last month's banner. ## Bounding the damage 1. **Emit the outcome of every render attempt** — path, duration, success or failure, and the error. This is the signal that names the cause; without it you learn only that something is old. 2. **Track the age of each stored copy** and expose the worst age per surface. Age is the consequence a visitor actually experiences, so it is the number worth alerting on. 3. **Alert on age against the policy for that surface**, not on status codes. "No successful refresh for this path in an hour" is the alert that would have fired. 4. **Decide in advance what should happen when a copy is very old.** Continuing to serve it is usually right; for a surface where an old fact is harmful, the alternative is to stop trusting the copy and render in the request, accepting the latency and the possibility of an error. 5. **Make failures retry.** A single failed attempt should not take a path out of the refresh cycle, and a path stuck failing should be visible as such. ## When keeping the old copy is the wrong default Most of the time "serve something rather than nothing" is right. It stops being right when the *content* of the stale page carries an obligation — a price, an availability claim, a legal notice, a published deadline. There, silently serving an old page is worse than serving an error, because it produces commitments the system behind it will not honour. That is a judgment made per surface, and the rendering model cannot make it for you; what the model owes you is the visibility to notice that the copy has gone old. ## Where frameworks differ Frameworks vary in how much of this they hand you. Some report render failures through their own logging and expose nothing about copy age; some retry on a schedule of their own; some record a failure marker so a subsequent request can be handled differently. Partial publication of a half-finished render is the behaviour to check for specifically — most implementations replace the stored copy only after a render completes, but an implementation that streams output directly into storage has to answer what happens when the stream ends early. When you cannot determine the behaviour from documentation, it is worth establishing by experiment, because it is the difference between an old page and a broken one.
- Why is keeping the old copy a better default than serving an error?Because the failure is about the copy's age, not its validity. The stored HTML still renders correctly, so serving it keeps the page working while the refresh problem is fixed. Replacing it with an error, or clearing it, would turn a background failure into user-visible downtime on a page that was fine a moment earlier.
- What would you instrument to catch this class of failure?Two things: the outcome of each render attempt, with the path and the error, and the age of each stored copy with a worst-case per surface. Alert on age against the surface's policy. Request-level signals are the wrong place to look, because they stay healthy throughout by design.
- Can a failed render leave a partly written page in storage?Most implementations replace the stored copy only once a render has completed, so a failure part-way leaves the previous copy intact. An implementation that writes output as it is produced has to answer this explicitly. It is worth confirming for your system, since the difference is between serving an old page and serving a truncated one.
The vending box still holds yesterday's edition when the press jams. It looks stocked and it takes your coin; the only sign of trouble is the date on the paper.
saying these in an interview costs you the question
- Thinks a failed refresh clears the stored copy.
- Expects the error rate to rise when refreshes fail.
- Assumes a half-finished render is published as-is.
- Believes an old page always beats an error response.
- Says health checks would surface the stalled refreshes.