How would you keep failure precedence uniform across an edge gateway and dozens of independently built backend services?
answer
- uniformity beats any single choice
- the weakest endpoint is the oracle
- edge makes it common, service makes it true
- inherit the shape from a template
- price the support cost openly
basics
~20 sWrite the order down as a contract - credential, then rights, then existence - settle credentials once at a shared edge, and verify every endpoint against it. The one service that answers differently becomes the enumeration oracle.
solid answer
~50 sPrecedence is observable API behaviour, so across a fleet the property that matters is **uniformity** rather than any particular choice. Three moves make it real. First, write the contract: the order of the checks, which resource classes hide existence, what a hidden answer looks like down to body and headers, which paths are anonymous, and where the true reason is recorded. Second, put the shared decisions in shared code - the edge terminates credentials and emits the same refusal everywhere, and a common response component gives every service the same failure body. Third, verify per endpoint rather than per team, because the policy is only as strong as its weakest surface: nineteen services answering not-found and one answering forbidden hands an attacker a confirmation oracle. What you trade is diagnosability, client affordance and support cost, and those have to be priced openly rather than discovered later.
go deeper
Understand that the order in which a request is refused is part of what an API promises, and that different services answering differently is a problem rather than a detail.
Be able to explain why one non-conforming endpoint undermines a fleet-wide masking policy, and what parts of a refusal have to match for the policy to hold.
Show how you would make conformance automatic: a shared failure-response component, a service template with the chain already ordered, and probes of the negative paths in delivery.
Own the tradeoff and the standardisation. Decide which resource classes mask, price the support and integration cost, and design so that a team gets the agreed behaviour without having to remember it.
On one service, failure precedence is an ordering question. Across a fleet it becomes an organisational one, because the property you are trying to hold is not "this endpoint is correct" but "no endpoint disagrees". ## Uniformity is the security property A policy that hides whether records exist is only as good as its least conforming endpoint. If most surfaces answer not-found for a record the caller may not see, and one answers forbidden, that one is an oracle: an attacker uses it to confirm identifiers and then uses them everywhere else. The same is true of ordering. If one service checks existence before rights, its status codes enumerate its data regardless of what the other services do. This is why the fleet-level question is not "what is the best order" but "how do we make one order inevitable". ## What the contract has to pin down 1. **The order of the stages** - credential, coarse rights, binding, lookup, fine-grained rights - stated as behaviour, not as a code structure. 2. **Which resource classes hide existence**, by class rather than by endpoint, so the decision is reviewable and consistent. 3. **The exact shape of a refusal**: status, body fields, which headers are present, and that the hidden case is byte-identical to the absent one. 4. **Which paths are anonymous** - probes, metrics, credential issuance, public assets - as an explicit, short list. 5. **Where the true reason lives**: log fields, a correlation id returned to the caller, and metrics that separate masked refusals from genuine absences. 6. **Who may deviate, and how** a deviation is recorded, because some surfaces legitimately need a different answer. ## Where each decision belongs | Decision | Best home | Why there | |---|---|---| | Credential validity and the `401` shape | Shared edge, re-checked in the service | One answer for every caller; the service must not trust reachability | | Coarse rights per route | The service | Only the service knows its own operations | | Existence masking per resource class | The service, after the lookup | Only the stage holding the record can make the two cases identical | | Failure body format | Shared library or template | Identical shape is what makes masking believable | | Anonymous path list | Shared configuration, reviewed | It is the hole in the credential boundary | Note the first row. An edge that enforces the credential is not sufficient on its own: traffic can reach a service by another route - internal calls, a mesh path, a misconfigured ingress, a test harness - so each service still refuses on its own. The edge makes the common case uniform; the service makes it true. ## Make the shape inheritable, not memorable - Ship a service template that already has the chain wired in the agreed order, so a new service inherits it without a decision. - Put the failure-response construction in one shared component, so the body cannot drift per team. - Add a conformance check to the delivery pipeline that probes the agreed negative paths on every service, rather than relying on each team to remember. - Make the anonymous path list a reviewed artefact with named owners, since it is the one place the boundary is deliberately open. ## What you trade away - **Diagnosability.** A legitimate user who cannot tell "no access" from "wrong URL" opens a support ticket instead of fixing their own request. - **Client affordance.** An integration that could have prompted "request access" now sees only an absent resource. - **Support and incident cost.** Answering "why can this user not see this" needs an internal tool, because the API deliberately no longer answers it. - **Onboarding friction.** Every new engineer meets a surprising `401` on a typo URL, and the reasoning must be written down somewhere they will find it. ## Failure modes to watch for - The anonymous path list grows by prefix until it covers more than anyone intended. - The edge short-circuits with a different body than the services use, so the edge's refusals are distinguishable from theirs. - A new service is written from scratch instead of the template and quietly checks existence first. - Masking is applied to the credential failure as well, and clients lose the challenge that tells them to authenticate. - The true reason is dropped from the logs alongside the wire response, leaving nobody able to reconstruct what happened. The judgement to demonstrate here is that this is a standardisation problem with a security payoff, and that a policy which depends on every team remembering it is a policy that has already failed somewhere you have not looked yet.
- Why is enforcement at the edge alone not enough to guarantee the contract?Because not all traffic arrives through the edge. Internal calls, a mesh route, a misconfigured ingress or a test harness can reach a service directly, and a service that assumes it was already filtered will serve them. The edge makes the behaviour uniform for the common path; the service is what makes it true.
- Which resources should not hide their existence?Ones whose existence is already public: catalogue entries, published documents, status pages, anything discoverable elsewhere. Masking them costs client experience and hides nothing. Deciding this by resource class rather than endpoint also stops the mask itself from marking which resources you consider sensitive.
- How do you stop the anonymous path list from becoming the weak point?Keep it exact rather than prefix-based, give it named owners and a review whenever routes change, and require each entry to be safe to serve to anybody on its own merits. An entry that is only safe because 'nobody knows the path' is not an exception, it is an unprotected endpoint.
saying these in an interview costs you the question
- Leaves precedence to each team and assumes the shapes will converge
- Masks existence at one endpoint and treats the policy as delivered
- Puts the whole policy in the edge and skips service-side refusal
- Drops the true failure reason from logs as well as from the response
- Treats masking as free and never accounts for the support cost