A non-document route ships in the same build and deployment as the pages. What does it inherit, and which handlers does that rule out?
answer
- colocated means co-constrained
- the adapter's target shape sets the ceilings
- duration, body size, available server APIs
- fit it, except it, or move it out
basics
~20 sIt inherits the pages' adapter, runtime and per-request ceilings - execution time, body and response size, available server APIs, scaling and release cadence. Handlers needing a long-held connection, long compute or native APIs are what that rules out.
solid answer
~50 sThe endpoint is not independently deployed, so it gets the pages' environment whole: the same adapter and target shape - a long-running server process, a per-request function, or a constrained edge runtime - and with it the same timeout, request and response size ceilings, available server APIs, scaling behaviour and release cadence. That inheritance rules out a family of handlers. A response held open to push events fights a platform that bounds each invocation, which is the ordinary reason a connection tier runs as its own deployment. Long CPU-bound work outlives the per-request timeout. A handler needing raw sockets, a native driver or filesystem writes finds them missing from a reduced runtime. The three ways out are to fit the constraint, to take a **per-route exception** where the framework allows one, or to move the work to a deployment with the shape it actually needs.
go deeper
Remember that the endpoint runs where the pages run, so it gets the same environment and the same limits rather than picking its own.
Name the three hosting shapes - a long-running process, per-request functions, a constrained edge runtime - and what each does to timeouts, available APIs and whether a response can be held open.
Diagnose from symptoms: local-only success points at a runtime difference, failure past a size threshold at a body limit, failure at a round elapsed time at a duration cap. Then choose fit, except, or move out.
Weigh a per-route exception against a separate deployment: the exception is cheap today and splits the platform's assumptions tomorrow, so decide which constraints the team wants to keep uniform.
The defining property of an in-app HTTP endpoint is that it is **not independently deployed**. It is compiled in the same build as the pages, packaged by the same adapter, and run in the same place. Everything convenient about that is inheritance, and everything limiting about it is the same inheritance. ## What it inherits - **The execution environment.** Whatever the pages run on, the handler runs on. Broadly three shapes exist: a long-running server process, a per-request function started and torn down around each invocation, and a constrained edge runtime offering a reduced API surface close to the user. - **The ceilings of that environment.** Maximum execution time per request, maximum request body size, maximum response size, memory, and whether a response may be streamed at all. - **The available APIs.** A constrained runtime typically omits direct filesystem and raw socket access and provides only web-standard primitives; a full server runtime provides the lot. - **The deployment topology.** It is served from wherever the pages are served from, with the same scaling behaviour. It does not choose its own location or its own scaling profile. - **The request interception step.** Whatever the app runs in front of matched requests normally runs in front of the handler too. - **Build and release cadence.** Shipping a change to the endpoint means shipping the pages; rolling the pages back rolls the endpoint back with them. ## What that rules out | A handler that needs… | Why the inherited shape fights it | |---|---| | A connection held open for minutes or hours | a per-request function is bounded by its invocation ceiling, and many edge runtimes cap duration too | | Long CPU-bound work — a large export, a heavy render | the per-request timeout applies; the work outlives the request | | Native server APIs — raw sockets, a native driver, filesystem writes | a constrained runtime does not expose them at all | | Very large uploads or downloads | request and response size ceilings are set by the platform, not by the handler | | Its own scaling or isolation profile | it scales with the pages because it *is* the pages' deployment | | In-memory state surviving between requests | the instance answering the next request may not be the same one | The long-held connection is the case interviewers reach for most often, because it is where the inheritance is most visible. A push-style response — a body kept open while events are written into it — works comfortably on a long-running process and fights a platform that bills and bounds each invocation. That is the ordinary reason a connection tier runs as its own deployment rather than as a route inside the frontend application. ## The three ways out 1. **Fit the constraint.** Rewrite the handler to use only what the runtime provides, keep the work inside the timeout, cap the payload. Often the right answer, and always the cheapest. 2. **Take a per-route exception.** Most meta-frameworks let one route declare a different runtime, a longer timeout, or a different execution mode from the rest of the app. That is a real tool, but it splits the deployment's assumptions: the route no longer shares the environment its neighbours were tested in, and the next reader has to notice the declaration to understand why it behaves differently. 3. **Move it out.** Hand the work to a separate deployment with the shape it actually needs — a queue and a worker for long jobs, a dedicated tier for held-open connections, a service with a full runtime for native drivers. The endpoint that remains, if any, becomes a thin thing that accepts the request, records the intent and answers immediately. ## Diagnosing it in production The symptoms are recognisable: - **Works locally, fails deployed.** Development almost always runs as a long-running process with everything available, so this usually means a reduced API surface or a ceiling that does not exist on your machine. - **Succeeds for small inputs, fails past a threshold.** A request-body or response-size limit. - **Fails at a suspiciously round elapsed time.** A duration cap, whether the platform's or a proxy's. - **A streamed response arrives all at once, or not at all.** Something in front of the handler buffered the body — a proxy, a compression layer, or a platform that collects the whole response before returning it. The check that resolves most of these fastest is to compare the app's configured target with what the handler needs, rather than reading the handler's code again: the code is usually fine and the environment is usually the reason. ## The principle behind all of it Colocation is a packaging decision, not an architectural exemption. The endpoint gets the pages' convenience *and* the pages' constraints, and it is worth knowing which of those constraints your handler is closest to before production is the thing that tells you.
- A handler works locally but fails only once deployed. What do you check first?The difference between the local server and the deployed target. Development almost always runs as a long-running process with the full runtime available, so a handler that fails only in production has usually hit a reduced API surface, a duration cap or a size ceiling that does not exist locally. Compare the configured target against what the handler needs before re-reading its code.
- What is the cost of giving one route a different runtime from the rest of the app?It is a real escape hatch, but it splits the deployment's assumptions. That route no longer shares the environment its neighbours were tested in, so its available APIs, startup behaviour and limits differ from everything around it, and the next reader has to notice the declaration to understand why. Worth it deliberately; bad as an accident.
- Why does a streamed response sometimes arrive all at once?Something between the handler and the client buffered it. A proxy, a compression layer, or a platform that collects the whole body before returning it will hold the chunks until the response completes. Writing in chunks is only half the path; the deployment shape and whatever sits in front of it decide whether those chunks actually leave early.
saying these in an interview costs you the question
- Assumes a handler can hold a connection open on any hosting shape.
- Believes a colocated endpoint can scale independently of the pages.
- Blames the handler's code when the environment imposed the limit.
- Thinks the adapter will substitute missing native server APIs.
- Expects in-memory state to survive between requests on every target.