You will rewrite the indexer's wildcard grant from thirty days of recorded calls — what can that window miss, and how do you cover it?
answer
- evidence runs one way only
- absent is not unneeded
- the window versus the business cycle
- quarterly job, recovery path, flagged branch
- watched rollout with cheap rollback
basics
~20 sA usage window shows what was called, never what is needed. It misses work that did not run inside it — periodic jobs, failure and recovery paths, rarely-taken branches. Cover it with a window longer than the longest business cycle, a declared list of rare operations, and a watched rollout.
solid answer
~50 sTightening from observed usage is the right method, but the evidence is one-directional: a permission that appears was definitely used; a permission that does not appear was merely not used **in that window**. Thirty days misses anything on a quarterly, half-yearly or annual rhythm, plus every error and recovery path the job did not hit, plus branches behind a feature flag nobody turned on. I would take a window longer than the longest business cycle the workload participates in, ask the owning team for the operations they know are rare, and then ship the narrow grant in a watched rollout — non-production first, alerting on refusals, with the previous grant restorable in minutes. Refusals after the change are the signal that the derivation was incomplete, and they should page someone rather than sit in a dashboard.
code
pseudocode · 19 linesobserved = empty set
for each call in usageRecord(principal = "workload/search-indexer",
from = today - 400 days, to = today):
if call.outcome == "allowed":
observed.add(call.action, call.resource)
else:
# refused calls are NOT evidence of a need, and are never
# silently dropped: each one goes to a human for a decision
refusals.add(call)
candidate = grant(effect = "allow",
action = distinct actions in observed,
resource = commonPrefix(resources in observed))
for each operation in ownerDeclaredRarities:
candidate.add(operation) # nothing in the record can prove these unused
return candidate, refusals # ship candidate in a watched mode, review refusalsgo deeper
Recall that platforms record the calls a principal makes, and that this record is where a narrow grant is derived from rather than guessed.
Explain why the inference is one-directional, and name concrete work a short window cannot have observed: periodic jobs, recovery paths, flag-gated branches.
Show the rollout you would run — window choice, declared rarities, comparison against live traffic, refusal alerting, and a rollback cheap enough that nobody pre-emptively widens the grant.
Decide what the organisation buys to make this repeatable: how long call records are kept, who owns a grant after derivation, and what review rhythm stops the snapshot from rotting.
## What the record actually proves Every platform keeps a record of the calls its principals make. Tightening a grant from that record is the only honest way to narrow a grant you did not write and do not fully understand — reading the code tells you what the happy path does, and the record tells you what actually happened. But the inference only runs one way: | Observation | Sound conclusion | Unsound conclusion | |---|---|---| | Permission appears in the window | it is needed | — | | Permission never appears | it was not used in this window | it is not needed | | No refusals in the window | nothing was blocked then | the grant is correctly sized | The second row is where tightening goes wrong, and it goes wrong weeks later, in the one execution nobody was watching. ## The shapes of work a thirty-day window misses - **Periodic work on a longer rhythm.** Quarter-end reconciliation, a half-yearly index rebuild, an annual archive. The job did not run, so no evidence of it can exist. - **Failure and recovery paths.** The retry that reads from a second location, the compensating write that undoes a partial run, the path that only executes after a crash. A healthy month produces no trace of any of them. - **Branches behind a switch.** A code path nobody enabled during the window — a fallback format, a re-index mode, a migration branch. - **Human-triggered operations.** An operator re-running the job by hand with a flag the schedule never sets. - **Onboarding-shaped work.** First-run behaviour that only occurs when a new store or a new tenant appears. None of these are exotic. They are the ordinary reason a grant derived from observation breaks in the second month rather than the first. ## Choosing the window The window should be **longer than the longest cycle the workload participates in**, which in practice means a year and a bit for anything financial or regulatory, and at least a full quarter for everything else. Two practical constraints push back: 1. The record is usually retained for a limited period, and long retention is a deliberate, paid-for decision that may not have been made before you needed it. 2. The workload may not be old enough to have a year of history. A six-month-old service has six months of evidence and no more, and honesty about that is part of the answer. When the window is shorter than the cycle, the gap is closed by people rather than data: ask the owning team to declare the operations they know are rare, and treat that list as first-class input alongside the record. ## Rolling the tightened grant out safely 1. **Derive the candidate** from the longest window available, plus the declared rarities. 2. **Compare, do not replace.** Evaluate the candidate against continuing traffic before it decides anything — every call the old grant allowed and the new one would not is a finding to explain, not a bug to tolerate. 3. **Ship to a non-production environment first**, where the same code runs on a schedule you can accelerate. 4. **Alert on refusals**, routed to the owning team with the principal and the attempted action in the alert. A refusal after tightening is the most valuable signal you will get, and it is worthless in a dashboard nobody opens. 5. **Keep rollback cheap.** The previous grant should be restorable in minutes by whoever is on call, without a review cycle. If restoring takes an hour, the team will pre-emptively widen the grant instead. ## What to do with refused calls in the window The record contains calls the platform refused as well as calls it allowed. These are excluded from the derived grant — a refused call is not evidence of a need — but they must never be silently dropped, because a refusal is exactly two things at once: a legitimate operation somebody has been failing to perform, or an attempt nobody should be making. Both deserve a human decision, and the second is a security finding that the tightening exercise happens to have surfaced. ## The trap underneath the trap Deriving the grant from usage produces a grant that fits **today's behaviour of today's code**. It is a snapshot, and the code keeps moving. This is why the output of the exercise should be a grant plus an owner plus a review rhythm, rather than a grant alone: without the second two, the first one is widened by the next person it inconveniences, and the widening is permanent.
- The record only goes back sixty days because that is how long it is kept. What do you do?Say so explicitly and close the gap with people rather than pretending the data covers it: ask the owning team for operations on a longer rhythm, read the schedule and the failure-handling code for calls the window cannot have seen, tighten anyway, and make refusals after the change page someone. Separately, raise the retention decision — the next tightening has the same ceiling until it changes.
- Why not simply tighten, wait for something to break, and widen again?Because the break is a production incident on somebody else's clock, and the widening that follows is done under pressure by whoever is paged — which reliably restores a wildcard rather than the one missing permission. A watched rollout converts the same discovery into an alert with a named principal and action, decided by the owning team while nothing is down.
- The derived grant is narrower than the one the code review approved six months ago. Which do you ship?The derived one, with the difference explained. A reviewed grant records what somebody expected the workload to need; the record shows what it did. Where they disagree, the gap is either dead capability to remove or a rare path to declare — and both are decisions worth making on purpose rather than inheriting a wider grant because it was once approved.
saying these in an interview costs you the question
- Treats a permission absent from the window as proven unnecessary
- Picks the window by what is convenient rather than by the business cycle
- Ships the tightened grant straight to production with no refusal alerting
- Silently discards refused calls instead of having someone decide on them
- Calls the job done at the grant, with no owner and no review rhythm