Which cross-cutting concerns would you keep in the application's middleware chain rather than push to a shared edge layer?
answer
- ask what knowledge the concern needs
- route template and principal live inside
- uniform and early belongs outside
- mint once, adopt everywhere
- guarantees must survive a direct call
basics
~20 sKeep in the chain whatever needs application knowledge: the matched route, the principal, the tenant, the meaning of a failure. Push to the edge what is uniform and must act before traffic arrives. A few concerns belong in both.
solid answer
~50 sThe deciding question is what the concern needs to know. A per-account quota, a route-labelled metric, an authenticated identity and an error response shaped like the rest of the API all depend on knowledge only the application has, so they stay in the chain. Coarse flood limits, transport termination, shared caching and bulk asset serving apply uniformly and are cheaper before traffic reaches an application process, so they belong at the edge. Some concerns legitimately exist at both, with different roles: the edge mints a correlation id and the application adopts rather than regenerates it; the edge caps floods while the application enforces fair share. The failure mode to argue against is an application whose guarantees only hold on the path through the edge, because then anything reaching it directly is unprotected and local runs behave differently from production.
go deeper
Know that some protections live in front of the application and some inside it, and that the ones inside are the ones needing to know who is calling or which route matched.
Apply the criterion to concrete cases: a route-labelled metric and a per-account quota inside, transport termination and bulk asset serving outside, and say why each way round.
Argue the both-layers cases — coarse versus fair-share limiting, minting versus adopting an id — and show what goes wrong when a service is called directly, bypassing the edge.
Own the split as policy: which layer is canonical for each concern, how a change is rolled out across teams, and the invariant that the application stays defensible without the edge in front of it.
## The question that decides it When a team asks whether a cross-cutting concern belongs in the application's middleware chain or in a shared layer in front of it, the productive question is not *which is faster* but: **does this concern need knowledge that only the application has?** Application knowledge means things like the matched route template, the authenticated principal, the tenant, the plan a customer is on, whether a failure is a client's fault or the service's, and what the API's error envelope looks like. A layer in front of the application sees a request and a response; it does not see any of that, and every attempt to make it see them ends in fragile path patterns that become stale the moment a route is renamed. ## A working split | Concern | Belongs at | Why | |---|---|---| | Transport termination, connection limits | edge | below the application entirely; uniform for all services | | Coarse flood limiting by address | edge (and often a chain hook too) | cheapest before a process is involved | | Fair-share quota by principal or tenant | chain | needs an identity only authentication establishes | | Route-labelled metrics and access log | chain | needs the matched route template, not the raw path | | Correlation id | minted at the edge if one exists, adopted in the chain | the join must span both, so exactly one layer mints it | | Bulk static assets, shared caching | edge | uniform, cacheable, no application state involved | | Response compression | one layer, chosen deliberately | doing it in both is waste, and double encoding is a bug | | Authentication to a principal | chain | downstream code needs the identity, not just a verdict | | Error shaping to the API's envelope | chain | only the application knows the contract it publishes | ## Concerns that genuinely live in both A few things are not a choice between layers but a division of labour, and being able to describe that division is what distinguishes a considered answer: - **Rate limiting.** A blunt, generous cap at the edge stops floods before they cost a process anything; a tight, identity-keyed limit in the chain enforces what the product promises. Neither substitutes for the other. - **Correlation ids.** One layer mints, every other layer adopts. The common defect is an application that regenerates the id it was given, which severs the join precisely where a request crossed a boundary. - **Observation.** The edge's count and the application's count will differ — the edge sees requests the application never received. That difference is information, as long as someone knows which number answers which question. ## What breaks when everything moves to the edge - **The application stops being runnable on its own.** Local development and tests exercise a chain with none of the protections production has, so behaviour diverges exactly where it matters. - **Guarantees become path-dependent.** Anything that reaches the application directly — an internal caller, a misrouted health probe, a new ingress someone added — bypasses every protection at once. A guarantee that depends on the route traffic took is not a guarantee. - **Policy changes become cross-team tickets.** Shipping a change to a limit or a header now requires a team that does not own the service, on their schedule. - **Diagnosis crosses an ownership boundary.** The evidence for a failure lives in two systems with different vocabularies and different retention. ## What breaks when nothing does - **Duplication across services**, each implementing the same policy slightly differently, with differences discovered during incidents. - **Drift**, because there is no single place to change a rule and no way to prove all services changed. - **Cost**, since work that could be done once for all traffic is done once per service, per process. ## How to present the decision A strong answer names the criterion first — *does it need application knowledge?* — then works through the ambiguous cases rather than the easy ones: where the correlation id is minted, which layer's request count is canonical, whether compression happens at the edge or in the chain, and how the fair-share limit relates to the flood cap. Finish on the invariant worth defending: **the application should still be correct and defensible when called directly**, with the edge making it cheaper and more uniform rather than making it safe. That framing is the difference between an architecture and a deployment accident.
- If an edge layer already mints a correlation id, what should the application's hook do?Adopt it after a cheap shape check, and generate one only when it is absent. Regenerating an id that was already minted severs the join at exactly the boundary crossing you most want to follow. The rule is that one layer mints and every other layer adopts, and it needs to be stated once for all services.
- The edge and the application report different request counts. Which is wrong?Neither, usually. The edge counts requests it received, the application counts requests it processed, and the difference is traffic refused, cached or dropped before reaching a process. The problem is only that nobody has declared which number is canonical for which question, so the two get compared as if they measured the same thing.
- Is it acceptable to rely on the edge for authentication?Only as a first filter. Downstream code needs the principal, not merely a verdict, and an application whose protection disappears when it is called directly has a path-dependent guarantee. The usual arrangement is that the edge may reject obvious garbage early while the chain still establishes identity for itself.
saying these in an interview costs you the question
- Decides by performance intuition rather than by what the concern must know
- Labels edge metrics with raw paths because the route is unknown there
- Assumes the application is safe because the edge protects it
- Regenerates a correlation id the edge already minted
- Compresses at both layers and calls the double encoding a client bug