skip to content

How would you decide, across an application's streamed routes, which work may sit in a deferred region and which must resolve before the shell?

level: principalimportance: should knowfreq 40%

answer

  1. sort work by what it can change
  2. response-shaping above, region-filling below
  3. first-byte budget versus response authority
  4. make the rule structural, not cultural
  5. measure shell and last chunk separately

basics

~20 s

Sort route work by what its result can change. Anything that can change the response itself — status, headers, cookies, destination — resolves before the shell; work that only fills a region may be deferred.

solid answer

~40 s

I would make it a correctness rule rather than a performance preference. The dividing line is the commit point: once the shell flushes, only the markup inside placeholders is still writable. So **response-shaping work** — does this exist, may this visitor see it, does the session rotate, is this cacheable, what goes in the document head — must resolve above the flush, and I accept the first-byte cost for it. **Region-filling work** — feeds, counts, recommendations — goes below, where the worst case is one degraded box. Then I make the rule hold structurally: shared enclosing levels or a pre-resolution hook own the checks, tests assert response status rather than markup, and we track time-to-shell separately from time-to-last-chunk so the budget for the awaited work is visible.

go deeper

for a junior

Take away the sorting question: ask what a piece of work decides. If it could change the status, a header or a redirect, it has to finish before anything is sent.

for a middle

Be able to place concrete work on the right side of the line, and explain that deferring costs nothing at first byte while awaiting costs everyone on the route.

for a senior

Show how you would find the violations that already exist and keep the awaited set cheap — hoisting shared checks, splitting a cheap decision from the expensive detail.

for a principal

Own the enforcement and the budget: structural placement, tests that assert the response, a first-byte budget with a named owner, and two separate latency numbers so the tradeoff cannot be optimised away.

## Frame it as a correctness boundary, not a speed dial Deferring work in a streamed route looks like a performance decision — move slow things below the shell, first byte improves. It is really a decision about **where in the response the work still has power**. Above the flush, a result can change the status, the headers, the cookies, the destination and the markup every consumer receives. Below it, a result can change one region's bytes and nothing else. Choosing where work sits is choosing which of those two kinds of authority it needs. Stated that way, the policy writes itself and stops being a matter of taste. ## The sorting rule | kind of work | example | placement | worst case if misplaced | |---|---|---|---| | decides the response | existence, permission, session rotation, cacheability | above the flush | a success page served to someone who should have been refused | | must be in the served markup | document title and description, canonical links, anything a script-free consumer reads | above the flush | consumers that read HTML without running script see the fallback | | fills a region | feeds, counts, recommendations, related items | deferred | a degraded box on an otherwise correct page | | refines a region | an expensive confirmation on top of a cheap check made above | deferred | a region that shows less than it could | The middle row is the one teams get wrong most often, because it does not feel like a response decision. It is one: that markup either exists in the bytes or it does not, and no script runs for some of the consumers that matter. ## Paying for what stays above The rule costs first-byte time, so the follow-through is to make the awaited work cheap rather than to move it: 1. **Hoist once, not per route.** A check that runs in a shared enclosing level runs a single time per request for a whole section, instead of appearing in every route's chain. 2. **Split cheap decision from expensive detail.** A token claim or a cached membership flag can decide *may they see this page* above the flush, while the expensive join that decides *what exactly to show them* is deferred. 3. **Give the shell a latency budget** and treat a breach as a defect. Without a number, every new check quietly buys itself a place above the flush. 4. **Accept the whole-response hold where the route earns it.** Some routes should not stream at all — a page that is entirely one restricted record, for example. That is a legitimate outcome of the rule, not a failure of it. ## Making it survive contact with a team Policies about ordering decay silently, because the decayed version still renders. What actually holds: - **structure over convention** — if the permission check lives in a level every route in the section shares, a refactor that defers a region cannot take the check with it; - **tests that assert the response, not the page** — an unauthorised request should be asserted to produce a refusal status; a test that asserts *the markup says access denied* keeps passing after the check drifts below the flush; - **review prompts that name the smell** — a deferred region whose code branches on whether the visitor is allowed, or whose result feeds the document head; - **two separate numbers in monitoring** — time to shell and time to last chunk. One number hides exactly the tradeoff being managed: a team optimising a single figure will defer the wrong work to move it. ## The tradeoffs to say out loud - **First byte versus response authority.** Every check above the flush delays everyone, including the visitors who would have passed it. That is the deliberate trade, and it should be spent on decisions, not on content. - **Streaming versus shared caching.** A response that streams still ends up stored in full by a shared cache, which means anything personalised in a deferred region is personalised content in a cacheable document. Either those regions stay out of shared-cacheable routes, or the route declares it varies — a header decision, and therefore one that lives above the flush. - **Streaming versus consumers that do not run script.** The swap that fills placeholders is a script step, so deferred content is invisible to anything reading the HTML alone. For routes where that audience matters, deferral is a content decision, not a delivery one. - **Team cost.** A per-region fallback and a retry path is real work. Deferring a region is not free even when it is correct, and a route with six placeholders has six failure states somebody has to design.

  • How do you keep the rule from eroding as features land?
    Put the checks somewhere a refactor cannot move them — a shared enclosing level or a hook that runs before route resolution — and make the tests assert response status rather than rendered markup. A markup assertion keeps passing after a check drifts below the flush, which is exactly when you need it to fail.
  • What would you measure to know the policy is working?
    Time to shell and time to last chunk as separate numbers, so the cost of the awaited work is visible instead of averaged away; plus counts of late client-side navigations and of failed deferred regions. Rising numbers in either count mean decisions have migrated below the commit point.
  • Does a shared cache change what you are willing to defer?
    Yes. A cache stores the completed response, so anything personal in a deferred region becomes personal content inside a cacheable document. Either those regions are kept off shared-cacheable routes, or the route declares what it varies on — and since that is a header, the decision itself must be made above the flush.

saying these in an interview costs you the question

  • Defers anything slow without asking what its result decides
  • Thinks awaiting one check costs the route its streaming
  • Treats the placement rule as style rather than correctness
  • Assumes a client-side fix covers consumers that never run script
  • Tracks only total render time, never time to shell