skip to content

Every request for precise collar coordinates needs an authorization decision within a few milliseconds — do you embed the evaluator, run it beside the service, or call a shared decision service?

level: principalimportance: should knowfreq 42%

answer

  1. latency and blast radius, the same two currencies
  2. budget the tail, not the median
  3. a remote answer is a hard dependency
  4. undecided is not denied
  5. 503 for cannot decide, 403 for no

basics

~20 s

Embed for latency and blast radius, run a co-located process for language independence, call a shared service only where central control outweighs a network hop on every request. A remote evaluator makes its availability a hard dependency of every guarded request.

solid answer

~50 s

The three placements trade the same two currencies: **per-request latency** and **who fails when it fails**. An embedded evaluator — a library such as a Cedar policy evaluator linked into the service — answers in microseconds with no network and no shared fate, but every process must carry the rule set and every runtime you use needs a binding. A co-located process beside the service costs a loopback hop, is language-independent, and takes down one instance rather than all of them. A shared remote decision service is one place to change rules and observe them, but it adds a round trip to every guarded request and couples the availability of the whole product to it. Whatever you pick, fail closed when no answer arrives — and return a retryable `503`, not a `403`, because "I could not decide" and "I decided no" are different facts for the researcher and for whoever is on call.

code

pseudocode · 13 lines
pseudocode
function enforceReadPreciseFix(request, credential, input):
    try:
        decision = evaluator.evaluate(input, timeout = 5ms)
    catch Timeout, Unreachable:
        # UNDECIDED: fail closed, but report availability, not entitlement
        metrics.increment("authz.undecided")
        return Response(503, "authorization temporarily unavailable")

    if decision.allow:
        return Response(200, body = render(input.resource))

    # the evaluator answered: identity known, refused
    return Response(403, "not permitted for this animal")

go deeper

for a junior

Know the shape of the choice: the rules can be evaluated inside your process, next to it, or over the network, and the further away the answer comes from, the more a request depends on something else being up.

for a middle

Explain the two currencies — per-request latency and blast radius — and the three status outcomes: 401 when the caller is unknown, 403 when the evaluator refused, 503 when nothing could be decided.

for a senior

Bring the numbers: budget the decision's p99 against the containing request, count the fan-out per page, and show the enforcement code where a timeout takes a different branch from a deny. Say what the operator's dashboard shows in each case.

for a principal

Price the outage, not the millisecond. A remote decision service turns every guarded request into a dependency on one component, and for a product whose data is all guarded that is a full outage under another name. Say what you would accept and why.

## The three places evaluation can run The rule is the same in all three. What changes is where the expression is evaluated relative to the code that enforces the answer. | placement | per-request cost | who fails with it | rule updates reach it by | |---|---|---|---| | **embedded library** | microseconds, in-process | nobody else; the process answers from its own copy | shipping a new rule set to every process | | **co-located process** | a loopback round trip, tens to hundreds of microseconds | that one instance | updating the co-located process, no service redeploy | | **remote decision service** | a network round trip, and the tail is what matters | every caller at once | one update, everywhere, immediately | An embedded evaluator distributed as a library — a Cedar policy evaluator, for example, is shipped this way — is the cheapest per request by roughly an order of magnitude, and it removes the decision from the list of things that can be unreachable. Its price is distribution: every service process holds a copy of the rule set, and every language runtime in your estate needs a binding to the same evaluator or the rules diverge in interpretation, which is far worse than them diverging in version. ## The latency budget is a tail budget "A few milliseconds" is not a median. If a map view fans out to 40 tile requests, each carrying its own decision, then a remote evaluator with a 1 ms median and a 40 ms p99 will show up in the user-visible p50 of the *view*, because one slow decision in forty is likely on every view. Two rules of thumb survive contact with production: - Budget the **p99 of the decision** against the p50 of the request that contains it, not median against median. - Count the **fan-out**. A decision per request is a different system from a decision per object in a page, and the second is where a remote evaluator stops being viable first. ## Availability coupling, and what it actually costs A remote decision service is a synchronous dependency of every guarded request. If it is down, nothing guarded serves — and for a service whose whole point is guarded data, that is a full outage with a different name. That is the cost a principal has to price, not the millisecond. An embedded evaluator does not remove dependency altogether; it changes its shape. The path that distributes rule sets is still a dependency, but its failure mode is **staleness** rather than unavailability: the process keeps answering with the last rule set it loaded. That is usually the better failure, and it is only better if you can tell it is happening — a process serving last month's rules while reporting itself healthy is the classic version of this going wrong. ## Fail closed, and say which failure it was When no answer arrives — timeout, connection refused, an attribute read that failed — the request is **undecided**. Undecided must not serve the data. For poaching-sensitive coordinates the asymmetry is total: a wrongly refused read costs a researcher a retry, and a wrongly served one cannot be taken back. But denying is not the same as telling the caller they lack permission, and the distinction is what makes the outage diagnosable: 1. **The evaluator answered deny.** The endpoint returns `403`: the identity is known and refused. That is a statement about entitlement. 2. **No answer arrived.** The endpoint returns `503`: the server could not decide, and the caller may retry. The researcher sees "temporarily unavailable", not "you do not have access", and does not open a ticket asking for a grant they already hold. 3. **The caller is not authenticated at all.** That is a `401` with `WWW-Authenticate`, and it is a different question from either of the above — decided before any rule runs. Collapse 2 into 1 and an evaluator outage arrives on the operator's screen as a flood of permission denials, which reads like an attack and is investigated as one. That confusion has cost teams hours in the middle of an incident, and it is avoided by one branch in the enforcement code. ## The rollout nobody budgets for A new version of the rule set does not reach every evaluator at the same instant. During the rollout window, two identical requests one second apart can legitimately get different answers, because they were evaluated by processes holding different versions. Three things make that survivable: keep the window short and known; carry the rule-set version in the answer so a support question can be settled without guesswork; and when a change reverses an answer in the unsafe direction, roll the restrictive version out first and let the permissive one lag. None of this makes skew disappear — it makes it bounded and explainable.

  • A rollout leaves half your evaluators on the new rule set. What can two identical requests one second apart return?
    Different answers, legitimately — they were evaluated against different versions. Keep the rollout window short and known, make the rule-set version readable from the answer so a support question can be settled, and when a change reverses an answer in the unsafe direction, deploy the restrictive version first and let the permissive one lag behind it.
  • Does an embedded evaluator remove availability coupling entirely?
    It removes the per-request coupling: no hop, so no timeout on the hot path. The path that distributes rule sets is still a dependency, but it fails as staleness rather than as an outage — the process keeps answering from the last set it loaded. That is usually the better failure, and only if the age of the loaded set is visible.
  • When is a shared remote decision service the right answer despite the hop?
    When rules must change everywhere within seconds, when many services in several languages must agree exactly, or when the guarded traffic is low enough that a round trip is invisible. Administrative and back-office paths often qualify; a high-fan-out read path for map tiles rarely does.

Three ways to lock a door. A key in the lock works with the power out. A keypad wired to a panel in the same building fails only for that building. A reader that phones head office for every badge is trivial to reprogram centrally and useless the moment the line drops.

saying these in an interview costs you the question

  • Returns 403 when the evaluator timed out rather than denied.
  • Fails open on read endpoints because reads feel harmless.
  • Compares median decision latency against a tail-sensitive request.
  • Claims an embedded evaluator has no dependencies at all.
  • Assumes every evaluator holds the same rule-set version at all times.
  • Treats a co-located process as faster than an in-process call.