A team self-hosts a meta-framework app that a managed platform used to run — which responsibilities now sit with them?
answer
- the platform was doing hidden jobs
- put a proxy in front of the renderer
- asset work now competes for CPU
- no sticky state, bounded memory
- ready means a render can succeed
basics
~20 sServing static output in front of the render process, absorbing request-time asset work that now competes for the same CPU, keeping instances free of sticky state, bounding memory in a long-lived renderer, and withholding traffic until a render succeeds.
solid answer
~50 sA managed platform quietly did several jobs the render process never saw. Self-hosted, they come back. Put a reverse proxy or CDN in front so hashed assets and prerendered documents are answered there and never reach the renderer — otherwise it spends most of its CPU on file serving. Expect **request-time asset work** (image resizing, font subsetting) to now contend with rendering on the same cores, which is what turns a healthy app into a spiky one after deploys. Run instances that hold no sticky state, so any instance can answer any request. Watch memory: a process that lives for days accumulates in a way a per-request function never did, so caches need bounds. And do not send traffic to an instance until it can actually serve a render, not merely since it opened its port.
go deeper
Know the headline: self-hosting means the static files, the asset transforms and the transport work that a managed host absorbed now run on your machines, and something has to be put in front of the render process.
Explain the mechanics of each shift — why serving files from the renderer wastes CPU, why transformed assets should not be regenerated per instance, and why any state kept in one process breaks as soon as a second instance exists.
Diagnose from symptoms: spikes after deploys point at cold asset work, steady memory growth points at unbounded caches in a long-lived process, and errors only around deploys point at traffic arriving before a render can succeed.
Price the move honestly. Self-hosting buys runtime control, network placement and a cost curve you own, and costs a standing operational burden. Decide it against the team that will carry that burden, not against the hosting bill alone.
Moving from a managed host to your own infrastructure does not change the application; it changes **who does the work the application never asked about**. The renderer is the same code, running in a much less helpful environment. ## The four jobs a managed platform was absorbing 1. **Serving static output.** Hashed assets and prerendered documents are the overwhelming majority of requests a typical app receives. Managed hosts answer them from their own edge, so the render process only ever saw the requests that actually needed rendering. 2. **Request-time asset transforms.** Resizing an image per device, subsetting a font, recompressing a payload — on a managed platform these ran on someone else's capacity. 3. **Transport concerns.** TLS termination, compression, caching headers, HTTP version negotiation. 4. **Traffic gating.** Deciding when a new instance is fit to receive requests, and when an old one should stop receiving them. Self-hosted, each of these lands on your cluster, and three of them land on the same CPU as rendering unless you deliberately move them. ## Put something in front The single highest-value change is a **reverse proxy or CDN in front of the render process**, configured so that: - hashed asset paths and the prerendered document tree are served directly from disk or an object store, never proxied through to the renderer; - long-lived cache headers are set on hashed paths, and short or revalidating ones on documents; - compression happens there, not in the render path; - only the paths that genuinely need rendering reach the application at all. Skipping this is the classic self-hosting mistake. The app "works", latency is mediocre, CPU is high, and nobody connects the two — because from the renderer's point of view it is simply receiving a lot of requests. ## The CPU contention nobody budgets for Request-time asset work is the second trap. Encoding one image is far more expensive than rendering one page, and the two now share cores. The symptom is characteristic: **latency spikes right after a deploy and then settles**, because a cold instance has no transformed assets cached and is generating them all at once while also serving first traffic. | Symptom | Likely cause | Practical lever | |---|---|---| | High CPU, most requests are files | The renderer is serving static output | Terminate those paths at the proxy or CDN | | Latency spikes after each deploy, then settles | Asset transforms regenerating on a cold instance | Persist transformed assets outside the instance, or precompute at build time | | Responses correct but slow under moderate load | Compression and TLS in the render path | Move both in front of the renderer | | Memory climbs for days, then the instance dies | Unbounded in-process caches in a long-lived renderer | Bound cache size and entry lifetime; alert on the trend | | Some users see stale or missing data after scale-out | State that lives on one instance | Make instances interchangeable; keep shared state in a store | ## Instances must be interchangeable A long-running renderer makes it easy to keep things in memory that quietly become **sticky state** — a session map, a rate-limit counter, a regenerated document. It works with one instance and breaks the moment there are two, in the least helpful way: most requests behave, a minority do not, depending on which instance answered. The discipline is that any instance can serve any request, and anything that must be shared lives outside the process. ## Memory is now your problem A per-request function is torn down often enough that leaks rarely surface. A process that runs for days surfaces them all: caches with no eviction, per-request objects retained by a module-level structure, growing buffers. Treat steady memory growth across a flat traffic profile as a defect to investigate rather than a reason to schedule restarts — though a bounded restart policy is a reasonable safety net while you look. ## Do not take traffic too early An instance that has bound its port is not necessarily able to render: the route table may still be loading, configuration may be unvalidated, the first render may need to warm caches. Traffic that arrives in that gap produces errors that look like an application bug and appear only around deploys. The rule is simple — **report ready when a real render can succeed, not when the socket is open** — and the detailed sequencing of startup and shutdown is a discipline of its own. ## What you gain This list reads like a cost, and it is, but it buys real things: full control of the runtime and its dependencies, freedom to run alongside other services on your own network, no constraint from a hosting shape you did not choose, and cost behaviour you can model yourself. The decision is worth making with the list in hand, not from a preference for either side.
- Why does self-hosted latency often spike right after a deploy and then recover?Because a fresh instance starts with nothing precomputed. Request-time asset transforms are regenerated on demand while the same cores are serving first traffic, so CPU saturates briefly and then settles once the common assets exist. Persisting transformed assets outside the instance, or generating them at build time, flattens the spike.
- How do you tell whether the render process is doing work it should not be doing?Break its request log down by path shape. If hashed asset paths or prerendered document paths appear at all, the layer in front is misconfigured — those should terminate before the application. A large share of CPU in image or font work says the same thing about asset transforms.
- Is a nightly restart an acceptable answer to memory growth?As a safety net while the cause is investigated, yes; as the answer, no. Scheduled restarts hide the trend, and a leak that outruns the schedule under peak traffic still takes the instance down at the worst moment. Bound the caches, find what retains the references, and keep the restart as a backstop.
saying these in an interview costs you the question
- Lets the render process serve hashed assets directly
- Ignores that image and font work competes with rendering
- Keeps sessions or counters in one instance's memory
- Assumes a bound port means the instance can serve
- Calls steady memory growth normal for a long-running renderer