You own a Next.js App Router route whose main data takes roughly two seconds. How do you decide between streaming a skeleton with loading.tsx and letting the response block until the data is ready?
answer
- streaming reorders bytes, it does not shrink work
- ask what the shell is worth alone
- the first chunk commits the response
- status and redirects decided above the boundary
- fix the slow query before staging the wait
basics
~20 sStream when the shell carries real value the user can act on, and when committing to a 200 response early is acceptable. Block when the page is meaningless without the slow data, or when the response still needs to change its status, headers, or destination.
solid answer
~60 sI start with what the shell actually contains. If the route's fast content is genuinely useful — navigation, a product summary, anything readable — streaming turns two seconds of nothing into an immediately usable page, and the skeleton marks the one slow region. If the slow data *is* the page, streaming just replaces a blank tab with a shimmering one; the user still cannot do anything, and I have added layout shift and a more complex failure story for no real gain. The second axis is what the response still needs to do. Once the shell is flushed, the status code and headers are committed, so a decision made deep in a slow subtree can no longer produce a different status or redirect — it has to be surfaced inside the streamed markup. Third, I check whether streaming is even the right question: if the route could be static or revalidated on a schedule, fixing the rendering mode or caching the query beats streaming around a slow call. Streaming reorders bytes; it never makes the work smaller.
go deeper
Know that streaming shows something sooner but does not finish sooner, and that a skeleton only helps if there is real content around it to read.
Be able to compare the two options concretely: what the user sees at 200ms and at two seconds under each, and why the total work is identical either way.
Show that you check the slow dependency before staging the wait — caching, concurrency, rendering mode — and that you size fallbacks to avoid shifting the page when chunks land.
Own the consequences beyond the page: the response commits its status on the first chunk, late failures bypass HTTP-level monitoring, and a blanket skeleton policy creates duplicate UI the team must maintain. Decide which metric the team optimizes and defend it.
## Frame the decision, not the feature Streaming is cheap to switch on in the App Router — a `loading.tsx` file, or a `<Suspense>` around a subtree — which is exactly why it gets applied reflexively. The judgment being tested is whether you know what you are trading away, because the mechanism is a reordering of when bytes arrive and nothing more. Total server work, total transfer, and time-to-complete are unchanged. ## Axis 1 — does the shell have value? Ask what the user can read or do while the slow region is pending. - **Rich shell**: a dashboard whose nav, header and summary cards are fast while one analytics panel is slow. Streaming is clearly right; the page becomes usable in a couple of hundred milliseconds. - **Empty shell**: a search results page whose entire content is the slow query. Streaming produces a header and a shimmer. The user waits the same two seconds and now watches a placeholder pretending to be progress. Sometimes that is still preferable — it proves the app is alive — but it is a much weaker case, and a rich fallback (last-known results, cached partial data) usually beats an empty one. ## Axis 2 — what the response can no longer do This is the axis candidates most often miss. The moment the first chunk is flushed, the response has committed: its status line and headers are already on the wire. Anything the framework decides later, deep inside a slow subtree, cannot retroactively turn the response into a different status or a redirect — it has to be expressed inside the streamed markup instead. Practical consequences: - A route that must reliably answer with a not-found or a redirect based on the slow data is a poor streaming candidate; resolve that decision **above** the boundary, where the response is still uncommitted, then stream the rest. - Errors that occur after the shell is out surface through the nearest error boundary in the UI rather than as a failed request. Monitoring that watches only HTTP status will report these routes as healthy while users see error UI, so instrumentation has to move client-side or into the render itself. ## Axis 3 — is streaming answering the right question? Before optimizing *when* the bytes arrive, ask whether the two seconds are necessary at all. Common findings: the query is uncached and identical for every visitor; several awaits run sequentially that could run concurrently; the data changes hourly and the route never needed to be rendered per request. Each of those is a larger win than any boundary placement, and a candidate who reaches for `<Suspense>` before checking them is optimizing the symptom. There is also a coupling worth naming out loud: streaming is a property of a per-request render. Deciding to stream implicitly asserts that this route renders per request, which is a rendering-mode decision that should be made deliberately rather than inherited from a `loading.tsx` someone added. ## Axis 4 — the cost of the skeleton itself Skeletons are not free: - **Layout stability.** A fallback whose dimensions differ from the real content shoves the page around when the chunk lands. Fallbacks should be sized like what they replace. - **Perceived jank.** Many independent boundaries mean many independent reveals; a page that assembles in six visible steps can feel worse than one that appears once. - **Maintenance.** Every skeleton is a second, unversioned copy of a layout that drifts from the real component over time. ## Axis 5 — measurement and the honest metric Streaming improves first byte and first paint, and typically the largest-contentful paint when the largest element is in the shell. It does not improve time-to-interactive-with-real-data, and it can make a naive "page load complete" metric look worse because the response stays open longer. Decide up front which metric the team is optimizing and make sure the dashboard reflects the user's experience, not just the one number that moved. ## How I would actually answer "Two seconds is a symptom; I want to know why first. If the query can be cached or parallelized, that is the fix. If it genuinely must be slow per request, I stream — but only if the shell is useful without it, and only after confirming this route does not need to decide its status or destination from that data. Then I size the fallback like the content, keep the number of boundaries small enough that the page assembles in one or two visible steps, and make sure errors after the flush are still visible in monitoring."
- What can the response no longer do once the shell has been flushed?Its status line and headers are already on the wire, so nothing decided later can change them. A not-found or redirect determined inside a slow subtree cannot become a different HTTP status — it has to be expressed in the streamed markup. If such a decision must be authoritative, resolve it above the boundary while the response is still uncommitted.
- How does streaming change what your monitoring sees?Failures after the first chunk cannot flip the status code, so a route can serve error UI while HTTP metrics stay green. Instrument the render and the client error boundaries rather than trusting status codes alone, and alert on the rate of error-boundary renders the same way you would on 5xx.
- Your team wants loading.tsx in every route folder as a standard. Would you agree?Not as a blanket rule. On routes whose page renders fast, the file adds a skeleton the user never needed and a second layout to maintain; on routes whose whole content is slow, it dresses up a wait rather than removing it. I would make it a per-route decision with a shared skeleton kit, not a template default.
- How do you decide the fallback's fidelity?Match dimensions first — that is what prevents the content shoving the page when it lands — and shape second, enough that the user recognizes what is coming. Beyond that, fidelity is a maintenance cost: a pixel-perfect skeleton is a duplicate component that drifts from the real one every time the design changes.
saying these in an interview costs you the question
- Treats streaming as making the page load faster overall
- Adds a skeleton to a route whose only content is the slow data
- Assumes a late failure can still return a 500 status
- Recommends loading.tsx everywhere as a blanket standard
- Never asks why the query takes two seconds