Next.js middleware executes on every request it is applied to, before the response is produced. What size and CPU limits does that impose on it, and how do you stay inside them?
answer
- it runs before everything, for everyone
- one bundle, with a hard cap
- the build output prints its size
- awaits become site-wide latency
- heavy work belongs one tier down
basics
~20 sMiddleware is one bundle on the hot path, so hosts cap its bundled size and give each invocation a short CPU budget. Keep dependencies tiny and edge-native, avoid embedding data files, and never do heavy computation or slow awaits there.
solid answer
~50 sTwo budgets bind at once. First, size: middleware compiles to a single edge bundle and hosting platforms enforce a hard cap on it — a few megabytes on Vercel, varying by plan — and `next build` prints the bundle's size so you can watch it. Second, CPU and latency: every matched request pays for whatever middleware does before anything else happens, and edge invocations get a short CPU budget rather than a generous one. Practically that means importing narrowly, choosing Web Crypto-based libraries such as `jose` over Node-crypto ones, never embedding a large JSON dataset like a full translations table, and treating every `await` on a network call as latency added to every page load. Anything expensive — a database lookup, a slow KDF, heavy parsing — belongs in the route or Server Component instead. If a decision is stable, compute it once and carry it in a cookie rather than recomputing it per request.
go deeper
Know that middleware runs on every applicable request and must stay small and fast, and that heavy libraries or big data files do not belong in it.
Explain the two separate budgets — a hard bundle-size cap enforced at deploy and a short CPU allowance per invocation — and name the build output as where the size is visible.
Show how you diagnose and control it: bisecting transitive imports, choosing edge-native libraries, bounding outbound calls with AbortController, and moving stable decisions into a cookie rather than recomputing them.
Own it as a platform constraint — a size budget tracked in CI, a rule about what is allowed on the pre-route path, and a clear escalation path for teams whose feature does not fit.
## Why middleware has budgets that ordinary server code does not Middleware occupies a uniquely expensive position: it runs before routing, on every request the project applies it to, and nothing else can start until it returns. A route handler that takes 40 ms affects the people who call that route; middleware that takes 40 ms affects everyone, on every navigation. Hosting platforms therefore constrain it deliberately, and the constraints show up in two independent forms. ## Budget one: bundle size All of `middleware.ts` and everything it imports compiles into one bundle for the Edge Runtime. Platforms cap that bundle's size — on Vercel the limit is a few megabytes and depends on the plan — and exceeding it is a deploy-time failure, not a slow request. You can see the number yourself: `next build` prints a Middleware line with its size alongside the route table, which makes it easy to spot the commit where it doubled. The things that blow the budget are predictable: ```ts // middleware.ts — this pulls a whole dataset into the edge bundle import translations from '@/i18n/all-locales.json' import { format } from 'a-large-date-library' ``` A full locale table, a country/IP dataset, a date library imported wholesale, a validation framework — none of them look expensive in a Server Component and all of them are heavy here. The remedies are boring and effective: import only the function you use so tree shaking can work, prefer small edge-native libraries, and keep data out of the bundle by fetching it or by encoding the decision in a compact form. ## Budget two: CPU time and latency Edge invocations are metered on CPU time and given a small allowance, on the assumption that middleware makes routing decisions rather than doing application work. Two different things can burn it: **Computation.** A deliberately slow function — bcrypt-style password hashing, a large key derivation, parsing a big payload, a catastrophically backtracking regular expression — is wrong here even when it fits in the bundle. The whole point of a slow KDF is to be slow, which is the opposite of what a per-request hook should be. **Waiting.** CPU time and wall-clock latency are separate, and awaits burn the second one. Every `await fetch(...)` in middleware adds its full round trip to the time-to-first-byte of every matched request, and if that call goes to a service that occasionally takes a second, so does your whole site. ## How to keep it small and fast - **Import narrowly.** Named imports from tree-shakeable packages, never a default import of a kitchen-sink library. - **Choose edge-native dependencies.** `jose` is built on Web Crypto and bundles small; `jsonwebtoken` is built on Node's crypto and cannot even load. Library choice is a size decision as much as a capability one. - **Do not embed data.** Locale tables, geo datasets and config blobs belong behind a fetch or in the route, not in the middleware bundle. - **Compute once, carry the result.** If a decision is stable for a session, put it in a signed cookie during the request that established it and read the cookie afterwards, instead of deriving it on every hit. - **Fail open on slow dependencies.** If you must call out, use `AbortController` to bound the wait so a degraded service does not become site-wide latency. - **Watch the number in CI.** The middleware size in the build output is a metric you can assert on; a size regression is far easier to catch at the commit that caused it than after a deploy fails. ## Diagnosing a middleware that got heavy Start from the build output to confirm size, then bisect imports — comment out one and rebuild — because the growth is nearly always a single transitive dependency rather than your own code. For latency, compare time-to-first-byte on a path the middleware applies to against one it does not; a consistent gap across otherwise unrelated routes points at middleware rather than at any individual page. ## The judgment being tested An interviewer is checking whether you understand that middleware is infrastructure, not a convenient place to put shared logic. The candidate who says "it seemed like a nice place for our feature-flag evaluation, and then every page got 80 ms slower" has learned the actual lesson.
- How would you find out that a middleware bundle has grown, before a deploy fails?`next build` prints a Middleware line with the bundle size next to the route table, so the number is visible on every build. Assert on it in CI, or at least review it in the build log for pull requests. Growth is almost always a single new transitive dependency, so catching it at the commit makes the cause obvious.
- Middleware calls an internal service on every request and that service occasionally stalls. What do you do?Bound the wait with an `AbortController` and decide explicitly what happens on timeout — usually fail open and let the request through, since a routing hint is rarely worth site-wide latency. Better still, remove the call: cache the answer in a signed cookie or move the check to the destination, which needs the data anyway.
- Why is a slow password-hashing function wrong in middleware even if it fits in the bundle?Because its slowness is the feature. A KDF is tuned to cost real CPU per invocation, and middleware invokes on every matched request — so you have multiplied a deliberate cost across all traffic and put it in front of the edge's CPU budget. Password verification belongs on a Node surface, invoked once at sign-in.
saying these in an interview costs you the question
- Treats middleware as a general place for shared application logic
- Ignores transitive dependency weight in the edge bundle
- Assumes an await in middleware is free because it is not CPU work
- Embeds a full translations or geo dataset in the bundle
- Thinks size limits only matter at runtime, not at deploy