skip to content

Your platform wants every authorization decision made at the API gateway so services stay simple - what do you accept, and which rule would you still enforce inside each service?

level: principalimportance: should knowfreq 47%

answer

  1. the gateway sees method, path, credential
  2. gateway-only makes the network the boundary
  3. coarse admission duplicated on purpose
  4. operation and object rules stay in the service
  5. decide now how a wrong answer is noticed

basics

~20 s

Accepting gateway-only authorization makes the network the trust boundary: anything reaching a service by another path is unguarded. Keep coarse admission at the gateway for cheap refusal and blast-radius control, and keep the operation-level and object-level rules inside each service.

solid answer

~50 s

A gateway can decide only what a method, a path and a credential contain, so a gateway-only policy silently redefines every rule it cannot express as 'trust the caller'. Worse, it makes the trust boundary the network: a job, a consumer, another team's service, or an operator with a port open reaches the service with no rule applied. The durable arrangement is deliberate, named duplication - coarse route admission at the gateway, because refusing unauthenticated traffic early is cheap and keeps load off services; operation-level and object-level rules inside the service, because only it can express them. Say which rule is duplicated and why, keep one resource-action vocabulary across both, and decide up front how a wrong answer will be noticed - a gateway-only model's mistakes are invisible for as long as nobody happens to probe the internal path.

code

json · 12 lines
json
{
  "edge_admission": [
    { "route": "/buildings/*/tickets/**", "methods": ["GET", "POST"], "require": "authenticated" },
    { "route": "/admin/**",             "methods": ["*"],           "require": { "roles": ["platform-admin"] } },
    { "route": "/health",               "methods": ["GET"],         "require": "none" }
  ],
  "unmatched": "deny",
  "not_expressible_here": [
    "ticket:close requires the managing-agent authority",
    "this principal manages the building on this ticket"
  ]
}

go deeper

for a junior

Understand the scope of what an edge component can decide: it sees a credential, a method and a path, so any rule about a specific stored record is beyond it.

for a middle

Explain why one endpoint that performs several operations defeats route-level rules, and what the service must therefore enforce for itself.

for a senior

Argue the topology risk concretely: name the callers in your own system that reach a service without passing the edge, and say what each of them is currently authorized by.

for a principal

Own the whole decision: what the single decision point costs in availability and ownership, whether the model can express next year's rule without a route per customer, what will detect a wrong answer eighteen months on, and what the migration costs while both enforcement points are live.

## The attraction, stated fairly One place to look, one team to own it, one file to audit, and application teams who never have to think about authorization. On a platform with twenty services and four languages, that is a real operational win and it is why the proposal keeps coming back. It is also why it needs answering with costs rather than with disapproval. ## What the gateway can and cannot express A gateway sees a credential, a method, a path, headers, and sometimes a body it is unwise to interpret. From that it can decide: - is the caller authenticated at all; - may this class of caller reach this route family; - coarse admission by an authority claim such as `roles` or `entitlements` carried on the credential. It cannot decide anything that depends on the stored record - which agent manages the building a ticket was raised against - without fetching the record itself. A gateway that starts loading domain data has become a second service with its own copy of the data model, its own staleness and its own deploy cadence, and the original service still has to load the same row. It also cannot see which operation is being attempted when one route fans out. If one endpoint closes, reassigns or re-prices a ticket depending on the body, a route rule must be the union of the three, and the weakest rule wins for all of them. ## The part that is not about expressiveness Even for rules a gateway *can* express, gateway-only enforcement makes a claim about topology: **that every call arrives through the gateway**. Over a system's life that claim decays: - the nightly escalation job and the e-mail consumer call the service directly; - a second service calls it in-cluster, because that is faster; - an operator opens a tunnel during an incident and never closes it; - a new route is added to the service and the gateway's rule set is updated a week later, or never. At that point the effective rule is 'anyone who can reach the port'. Nothing fails; there is simply no refusal anywhere. ## The arrangement to argue for | Decision | Owner | Reason | |---|---|---| | Is the caller authenticated | Gateway **and** service | Cheap early refusal; the service is not only reachable through the gateway | | May this caller reach this route family | Gateway | Coarse, credential-only, and the cheapest place to shed load | | May this principal perform `ticket:close` | Service | Only the service knows which operation the request became | | May this principal close *this* ticket | Service, where the row is loaded | The rule's input is the record | | What the refusal looks like | Shared contract | Two components refuse; callers should not learn two vocabularies | That first row is the duplication worth buying, and it should be written down as a decision rather than discovered as redundancy by the next engineer who deletes one copy. Refusing unauthenticated traffic at the edge protects service capacity and shrinks the blast radius of an unauthenticated flood; refusing it again in the service is what makes the topology claim unnecessary. ## Costs to put in the decision, not discover later - **Availability coupling.** Everything behind a single decision point inherits its availability. Ask what happens to the product when the gateway's rule distribution is stale or its evaluation is unavailable, and make the answer explicit rather than emergent. - **Ownership.** A central rule set becomes a file every team edits and no team owns. Decide who reviews a change that widens a route, and how a service team learns that the rule guarding their endpoint moved. - **Expressiveness in two years.** If next year's requirement is 'a contractor may close only the tickets assigned to them', a gateway-only model answers it with either a new claim on every credential or a route per case. Both are how you end up with a rule per customer. - **Detection.** This is the one teams skip. A wrong answer here is silent: an over-permissive route admits traffic and nothing logs a surprise. Decide now what would notice - a periodic assertion that every service route is covered by a rule, an alert when a service sees a call that did not come via the gateway, a refusal recorded with enough context to reconstruct it. Eighteen months is the usual gap between the mistake and the discovery, and the discovery is usually external. - **The migration, with both live.** Moving rules out of a gateway is not a switch. Both enforcement points run for months; during that window the effective rule is the intersection, so widening on one side does nothing and narrowing does everything - which is the safe direction, and worth saying out loud to whoever expects the migration to be observable. ## What to say 'The gateway keeps coarse admission and I will pay for that duplication on purpose. Everything that needs the operation or the record stays in the service, because the gateway cannot express it without becoming the service. And I want one thing in place before either: an assertion that every route into a service is covered by some rule, because the failure mode we are choosing between is not a bad decision - it is no decision, made silently.'

  • What breaks first when the gateway is the only enforcement point?
    Not the gateway - the assumption under it. The first caller that reaches a service without passing the edge, usually an internal job or a neighbouring service, runs with no rule applied and nothing records that anything unusual happened. The defect is not a wrong decision; it is the absence of one, which is why it survives review and monitoring alike.
  • How would you notice, eighteen months later, that a route rule was too permissive?
    Build the detection into the decision. Assert periodically that every route each service exposes is covered by a rule and that the coverage list matches the service's own route table; alert when a service handles a call that did not arrive via the edge; and record refusals with enough context to reconstruct the call. Without at least the coverage assertion, the usual discovery channel is a customer.
  • During a migration from gateway-only to service-side rules, both are live. What is the effective rule?
    The intersection: a request must satisfy both. That makes narrowing on either side immediately effective and widening on one side invisible, which is the safe asymmetry but confuses everyone expecting the migration to be observable. Plan to narrow the edge last, and keep one resource-action vocabulary so the two rule sets can actually be compared.

An estate with one gatehouse and no locks on the flats is secure exactly as long as everyone arrives by the gate. The delivery entrance, the contractor's key and the fire door do not care how good the gatehouse is.

saying these in an interview costs you the question

  • If the gateway allows it, the service can trust the caller
  • Internal traffic is safe because it is inside the network
  • Services should contain no authorization logic at all
  • The gateway can enforce object-level rules with enough path variables
  • Every rule in one place means every rule is audited
  • A too-permissive route rule will show up in monitoring