When a meta-framework app moves from a long-running server process to per-request functions, which capabilities can silently break?
answer
- process lifetime is the whole difference
- memory is a cache without guarantees
- post-response work may never run
- streaming needs unbuffered flushing end to end
basics
~10 sEverything that assumed the process stays alive between requests: in-memory caches and regeneration state, work continued after the response, connections held open, and streaming when any layer buffers the whole body before forwarding it.
solid answer
~50 sA long-running process gives you one address space that survives many requests, so a cache in a module-level variable, a regenerated page kept in memory, or a background task started after responding all work. On a per-request function target, an invocation may run on a fresh instance, so none of that carries over — the cache never hits, the regenerated copy is lost, and work after the response may be killed. Two more things bite. A **held-open response** (long polling, an event stream) collides with the platform's maximum invocation duration. And **streaming** only survives if every layer flushes as it writes; if the platform or a proxy in front buffers the body, the user gets one chunk at the end and the page looks slower, not broken. Anything that needs to persist has to move to storage every instance can reach.
go deeper
Remember the core fact: a per-request function may run on a fresh instance, so anything stored in memory between requests cannot be relied on. Prerendered files and static assets are unaffected.
Explain each break and its mechanism: caches and regenerated copies need storage all instances reach, post-response work needs a queue or must finish first, held-open connections hit a duration limit, and streaming needs unbuffered flushing through every hop.
Demonstrate the audit. These failures are silent, so name the concrete checks you would run against the real target — hit rate under churn, a regeneration seen from several instances, time-to-first-byte on a streamed route, a deliberate overrun of the duration limit.
Weigh which routes actually need the other shape. Splitting a handful of paths onto a long-running process is often cheaper than redesigning the app, but it doubles the operational surface — make that trade explicitly rather than by default.
Both shapes run the same server bundle. What differs is the **lifetime of the thing running it**, and almost every surprise in this migration traces back to that single difference. ## What each shape actually guarantees - **A long-running process** is started once and handles many requests in one address space. Module-level variables persist. Work can outlive a response. A connection can be held open as long as both ends want. - **A per-request function** is invoked to produce one response. The platform may reuse an instance for the next request or may not, and it reserves the right to stop the instance once the response is done. Nothing about the next invocation is guaranteed to share anything with this one. Read that as a rule: **on a function target, memory is a cache with no guarantee and no lifetime.** It is not empty — it is unreliable, which is worse, because behaviour that works in testing fails intermittently in production. ## The four things that break 1. **In-memory caches and server-side data caches.** A memoised result or a fetched-data cache in a module variable turns into a hit rate that swings with instance churn. Correctness is usually fine; latency and downstream load are not, because the origin now sees traffic the cache used to absorb. 2. **Regeneration of a prerendered page after a window or on demand.** The regenerated document has to live somewhere. On a process it can live in memory or on local disk; on a function target that copy is invisible to the next invocation, so pages appear to revert to the old version at random. The fix is storage every instance can reach, which is a design decision in its own right. 3. **Work continued after the response.** Flushing an analytics batch, warming a cache, writing an audit row — on a process these simply continue on the event loop. On a function target the instance can be frozen or torn down the moment the response is finished, so the work runs sometimes. 4. **Held-open responses.** Long polling, a server-sent event stream, or a slow upload are bounded by the platform's maximum invocation duration. The connection does not fail cleanly at that boundary; it ends mid-stream, and the client sees a truncated body. ## Streaming is the subtle one Streaming a response as it renders is not a property of the framework alone — it needs **every hop to flush as it writes**: the render code, the runtime, the platform's response layer, and any proxy or CDN in front. Function targets vary here; several will happily accept a streamed body and forward it unbuffered, while others (or a misconfigured proxy) accumulate the whole body and send it at the end. The failure mode is what makes this dangerous: nothing errors. The HTML is identical, the status is 200, and only the timing changed — the early shell that used to reach the browser immediately now arrives with the rest. It shows up as a regression in first-paint metrics long after the deploy, not as an alert. | Symptom after the move | Underlying cause | Where the fix lives | |---|---|---| | Cache hit rate collapses, origin load rises | Per-instance memory is not shared | Move the cache to storage every instance reaches | | A regenerated page flips between old and new | The regenerated copy stayed on one instance | Same: a shared store, not local memory or disk | | Analytics or audit writes go missing | Post-response work killed with the instance | Do the work before responding, or hand it to a queue | | A long-lived connection ends mid-stream | Maximum invocation duration reached | Shorten the connection, or keep that route on a process | | First paint gets later but the HTML is the same | A layer buffered the streamed body | Verify unbuffered flushing end to end, proxy included | ## What does not break Plain request-time rendering is fine — that is the shape's whole purpose. Prerendered documents and hashed assets are fine; they are served by the host before any code runs. Reads from an external database or API are fine. Code that treats each request as self-contained ports essentially unchanged, which is the real lesson: **the more stateless the request path already was, the less this migration costs.** ## Finding out before shipping Build with the target's adapter and exercise the app against the real shape, not the development server: hit a cached route repeatedly and watch origin traffic, trigger a regeneration and request the page from several instances, open a streamed route and check when the first bytes arrive, and let a held-open route run past the platform's duration limit on purpose. Each of these has a silent failure mode, so each needs an explicit check rather than a glance at the deploy log.
- Why is intermittent instance reuse worse than never reusing an instance?Because it hides the bug. If memory were always empty, a cache miss would show up on the first test. With occasional reuse, the cache works in development and on a warm instance, then misses under churn — so the defect appears as unexplained latency and downstream load in production rather than as a reproducible failure.
- How would you keep a route that holds a connection open when the rest of the app moves to functions?Keep that one route on a shape that allows it — a long-running process — and route only its path there, or replace the held-open connection with short polling that completes well inside the invocation limit. Splitting one route across shapes is usually cheaper than reworking the whole app.
- Does the same set of problems appear when moving in the other direction, to a long-running process?Different ones. Code written for functions is stateless and ports easily, but a process now accumulates: unbounded caches grow, leaked references are never cleared by an instance teardown, and one bad request can affect later ones. You gain shared memory and inherit responsibility for its lifetime.
saying these in an interview costs you the question
- Assumes module-level state persists across invocations
- Says the render code alone decides whether a response streams
- Expects work started after the response to always finish
- Thinks a green build means the features still work
- Treats intermittent cache misses as a platform bug