skip to content

Your authorization log records only denials to keep volume down — what does that cost you when an access is later disputed?

level: seniorimportance: should knowfreq 40%

answer

  1. a denial is the system working
  2. disputes are about allows
  3. the wrongly allowed read leaves no trace
  4. one record per request, not per row
  5. never sample cross-boundary or acting-as allows

basics

~20 s

Denials are the system working; the event a dispute is about is an access that was wrongly allowed, and a deny-only log has no record of it. The fix is not logging every allow, but recording every deny plus the allows that are not routine.

solid answer

~40 s

A denial is the control functioning, and it is the cheap half: rare, bounded, and interesting mostly in aggregate. The expensive half is the allow, and it is also the only half that answers the question you will actually be asked — *why was this read permitted?* Logging every allow is genuinely unaffordable on a read-heavy listing route, where one request can produce hundreds of object-level verdicts. The workable middle is: record **one decision per request, not per row**; record **every deny**; record **every allowed write**; record **every allow that is not routine** — granted across a site boundary, granted under an acting-as session, or granted by an administrative rule; and sample the remaining routine same-scope reads at a low rate with a stable key, so the class stays visible without the bill.

code

pseudocode · 7 lines
pseudocode
function shouldRecord(d):            # d is one decision, one request
    if d.verdict == "DENY":          return true   # always: rare, and the attack signal
    if d.action.isWrite:             return true   # allowed writes are never sampled
    if d.crossesSiteBoundary:        return true   # the shape a dispute is about
    if d.actor != d.subject:         return true   # someone acted on another's behalf
    if d.firedRule.isBreakGlass:     return true   # the exception is the point
    return sampledIn(d.requestId, rate = 0.01)     # routine same-site read: 1 in 100

go deeper

for a junior

Recall the core asymmetry: a refusal means the system worked, while the event people argue about later is something that was allowed. A log of refusals alone cannot describe it.

for a middle

Explain the two cuts that shrink the volume before any sampling — one record per request instead of per row, and separating reads from writes — and why they change the arithmetic more than sampling does.

for a senior

Name the cost and the classes: what an ingest bill looks like at page-size multiplication, which allow classes are exempt from sampling, and why sampling by route or principal is the dangerous kind.

for a principal

Own the policy: state which decision classes are recorded in full, defend it as a decision rather than an accident, and use the non-routine allow counts as the early signal that a rule was widened by mistake.

## The asymmetry nobody says out loud Deny-only logging is chosen for a good reason and defended with a bad one. The good reason is volume. The bad one is the implicit claim that denials are the security-relevant events. They are not. A denial is the system doing its job; a wall of denials from one principal is a useful signal, but it is a signal about an attempt that failed. The event that ends up in front of an auditor, a regulator or a customer is always the same shape: *this person saw data they should not have seen, and your system said yes.* A deny-only log has, by construction, nothing to say about it. So the honest framing is not allow-versus-deny. It is: which allows are worth their storage, and how do you cut the rest without losing the class? ## The arithmetic that makes people choose deny-only Take a trial platform where a monitor opens a site's participant list. One request, a page of two hundred participants, and an object-level verdict for each. At a modest fifty such requests a second that is ten thousand decisions a second, against perhaps a few hundred denials an hour. Two cuts change the shape of that number before any sampling: 1. **One record per request, not per object.** The interesting fact is *this principal listed participants at this site under this rule and received two hundred rows*, not two hundred near-identical rows. Record the query, the rule, the boundary and the count. 2. **Separate reads from writes.** Writes are orders of magnitude rarer and are the decisions with lasting effect. There is no volume argument for sampling an allowed write. ## Which allows are never routine | decision class | recorded | why | |---|---|---| | any deny | always | rare, cheap, and the signal for an attempt in progress | | allowed write | always | it changed something, and it is rare enough to afford | | allow crossing a site or tenant boundary | always | precisely the shape of the disputed access | | allow where the actor differs from the subject | always | somebody acted on another principal's behalf | | allow from an administrative or break-glass rule | always | the exception is the thing under review | | routine same-scope read | sampled | high volume, low information, still needs a visible baseline | The rows above are the ones a dispute is ever about. The sampled row is the one that is expensive and, ninety-nine times in a hundred, says only that the ordinary thing happened again. ## Sampling without losing the request Sample with a stable function of the request correlation identifier rather than a coin flip, so the choice is reproducible and so a trace and its decision record agree about whether they exist. And accept what sampling costs: for a sampled-out request you keep the fact that the request happened — the access log line, with principal, route and status — and lose the explanation of why it was permitted. That is a deliberate trade, and it is only safe because everything in the table above is exempt from it. The failure to avoid is sampling by principal or by route: a rule that stops recording a whole endpoint makes exactly that endpoint unreviewable, and the endpoints teams are tempted to exempt are the high-traffic read endpoints where cross-boundary reads hide. ## What you tell the auditor The defensible position is a stated policy about which decision classes are recorded in full and which are sampled, plus evidence that the classes in the table were never sampled. The indefensible position is a deny-only log and an explanation that begins *we only record failures.* The first says you decided; the second says you did not realise there was a decision to make. ## The second-order effect Once allows are recorded for the non-routine classes, their counts become an operational signal in their own right: cross-boundary allows per day, break-glass allows per week, acting-as reads per support agent. Those numbers move when a rule is widened by accident, and they move *before* anyone disputes an access — which is the only cheap way a wrongly-allowed class of read is ever caught. A deny-only log cannot produce any of them.

  • A routine read was sampled out and is now the request being asked about. What can you still say?
    That it happened, and little more. The access log line gives principal, route, object identifier, status and timestamp, which establishes the access and its time. What is gone is the explanation — which rule permitted it and what that rule read. That is the honest cost of sampling, and it is bounded by the fact that cross-boundary, acting-as, break-glass and write decisions are exempt.
  • Why record one decision per request rather than one per object returned?
    Because the per-object records are near-identical and the request-level fact is the one with meaning: this principal ran this query, under this rule, against this boundary, and received this many rows. Per-object records multiply volume by page size for almost no added explanatory power, and they are what makes people give up and switch to deny-only.
  • Is it ever right to sample denials?
    Rarely, and only for a single known-noisy source such as an unauthenticated scanner hammering one route, where the aggregate count is kept even when individual records are dropped. Denials are normally cheap enough to keep whole, and they are the input to rate-based alerting, which degrades badly when the underlying stream is thinned unevenly.

A shop whose camera only records when the alarm sounds. It has footage of everyone the door turned away, and none of the one person who walked in with a badge that should not have worked.

saying these in an interview costs you the question

  • Says denials are the security-relevant events and allows are noise.
  • Records an object-level decision for every row a listing returns.
  • Samples allows uniformly, including cross-boundary and break-glass ones.
  • Exempts the highest-traffic read route from recording entirely.
  • Believes the access log already explains why a read was permitted.
  • Samples with a random draw, so trace and decision record disagree.