A partner's long-lived ingest key turns up in a public repository — when does the last request carrying it actually fail, and how does the replacement reach the station?
answer
- the write is instant, the caches are not
- lag equals worst cache lifetime plus propagation
- invalidation is best effort, not a floor
- two live keys make rotation routine
- cut over on telemetry, retire on a deadline
basics
~20 sNot instantly: the leaked key keeps working for as long as any cached verification answer lives, so the honest number is the worst-case cache lifetime plus propagation unless an invalidation is published. The replacement rides an overlap — issue a second key, let the station cut over, retire the first on a watched deadline.
solid answer
~50 sRevoking writes `revoked_at` on the key row, but the request path rarely reads that row directly: each ingest node caches key lookups, a shared cache tier may sit in front, and read replicas lag. **The revocation lag is the worst-case cache lifetime plus propagation**, and that is the number you quote — 'about a minute', not 'immediately' — unless you publish an explicit invalidation and can show every node consumed it. Shorten it by keeping the cache lifetime small; the key table is tiny, so a short lifetime is affordable. The replacement never needs a maintenance window: the model allows more than one live key per station, so you issue the new one, the technician deploys it, you watch the old key's `last_used_at` and request count fall to zero, and you delete it on a published deadline. Then do the part people skip — read the old key's request log for the exposure window.
code
pseudocode · 15 lines# verification path, with the lag made visible
row = local_cache.get(key_id) # answer may be up to TTL old
if row is null:
row = key_store.read(key_id) # authoritative
local_cache.put(key_id, row, ttl=10s)
if row is null: return 401
if not constant_time_equals(digest(secret), row.verifier): return 401
if row.revoked_at is not null: return 401 # only seen once the cache refreshes
return row
# revocation
key_store.write(key_id, revoked_at = now())
publish_invalidation(key_id) # best effort: a node that misses it
# still honours its cached row for <= ttlgo deeper
Recall that revoking a long-lived machine key is a write on the server, that it is not always instant, and that the replacement key is issued before the old one is switched off.
Explain where the lag lives — the per-node cache, a shared cache, replica lag, a proxy in front — and why the cache lifetime, not the invalidation message, is the number you can promise.
Own the budget. State the lag in seconds, say what would double it, run the cutover on per-key telemetry rather than a calendar, and read the exposed key's request log before calling the incident closed.
Decide what the organisation is buying: leaks are inevitable, so spend on making them findable, bounded and quickly killable rather than on the belief that a partner's configuration file will stay private.
## What revocation actually writes, and what it does not reach Revocation is one write: `revoked_at` on the key row, or the row's deletion. That write is instant. What is not instant is every place an answer about that key already lives. | Where an answer is cached | Typical lifetime | What shortens it | |---|---|---| | In-process cache in each ingest node | 30-60s | A shorter lifetime, or a published invalidation every node consumes | | A shared in-memory cache tier | Minutes | Delete the entry as part of the revocation write | | Read replicas of the key store | Replication lag | Route the verification read to the primary, or accept the lag knowingly | | A reverse proxy or API gateway authenticating ahead of the application | Its own configured lifetime | Usually the longest and least visible of the four | So the honest answer to 'when does it stop working' is **the worst of those, plus propagation**. If you revoke at 14:02 with a 60-second in-process cache and nothing else in front, the last node stops accepting it by about 14:03. If a proxy in front caches for five minutes, your answer is five minutes and you should know that before an incident asks you. There are two ways to make the number small, and you should be able to argue both: - **Keep the cached lifetime short.** A key table of four hundred rows is not a scaling problem; per-request reads against a small, well-indexed table are affordable for many services, and a 5-10 second cache lifetime is a reasonable compromise for the rest. - **Publish an invalidation.** Push the revoked key identifier to every verifying node so caches drop it. This makes the common case fast, but it is a *best-effort* improvement: a node that missed the message still honours its cached answer until the lifetime expires. Keep the short lifetime as the floor underneath it; the invalidation is an optimisation, not the guarantee. Write the number down as a budget — 'a revoked ingest key stops working within 60 seconds' — and put a test behind it, because it is the kind of number a cache change silently doubles. ## The replacement, overlapped The thing that makes a key swap an outage is a data model that allows one key per station. Allow several, and the swap becomes routine: 1. **Issue the new key** alongside the old one. Both are live; both verify. 2. **The station deploys it** on the technician's own schedule — which at a university department with one part-time technician may be next Tuesday. 3. **Watch the old key's traffic**, not the calendar. `last_used_at` and a per-key request count tell you whether the cutover actually happened. This is the step that replaces guessing. 4. **Retire on a published deadline.** The old key is deleted on a date the partner was told, and the telemetry says whether that will hurt. 5. **Delete it, do not disable-but-honour it.** A key that is 'retired' yet still accepted is a live credential with a misleading label. In a leak, steps 1-3 run *after* the revocation, not before: you kill the exposed key first and accept that the station is down until the technician deploys the new one. That is the trade a leak forces, and it is why the same overlap mechanism exists for the planned case — so that most rotations never cost anything. ## Revocation is not the whole incident The scanner found the key because it had a recognisable leading marker. That marker is also the reason you know roughly when it became public. So: - **Assume it was used.** Pull the per-key request log for the exposure window and look at what that key identifier actually did — which series it wrote, whether anything read outside the station's own data, whether the source addresses match the station's. - **Check the blast radius against the scope.** This is where narrow per-key scope pays off: the answer may honestly be 'it could only have written one channel'. - **Tell the partner what you know**, including the window, because their copy may be in more places than the one repository. ## What you cannot fix, and what you do instead You cannot un-publish the string, and you cannot know how many copies were taken before you revoked it. So the design goal is not to make leaks impossible but to make them cheap: a recognisable marker so leaks are *found*, narrow per-key scope so the blast radius is one station's own data, a short cache lifetime so the kill is fast, and per-key telemetry so the aftermath is answerable rather than guessed at.
- Could you avoid the lag entirely by reading the key store on every request?Yes, and for a four-hundred-row table it is often the right answer: one indexed read per request against a small hot table is cheap, and it makes the revocation lag zero by construction. The reasons not to are a store you do not want on the critical path of every ingest write, or a verification step that already sits behind a proxy with its own cache. Decide it deliberately, and write the resulting budget down.
- The partner cannot deploy the replacement for a week. Do you leave the leaked key live?No — revoke it and let them be down, or, if the data genuinely cannot be lost, narrow the leaked key hard rather than leaving it as it was: strip it to write-only on its own channel and cut its ceiling to the station's real rate. That converts an open credential into one whose worst case is bounded, and it buys the week without pretending the leak did not happen.
- How do you know the cutover finished, if the station will not answer email?By the telemetry, which is why per-key `last_used_at` and request counts exist. Zero requests on the old key identifier for several times the station's normal reporting interval is the signal; a station that reports continuously makes this unambiguous within minutes. Retire on that evidence plus the published deadline, not on a reply you may never get.
- Why delete the old key rather than mark it disabled and keep accepting it for a grace period?Because 'disabled but accepted' is a live credential that everyone believes is dead, which is the worst possible state to be in during the next incident. If you need a grace period, the correct shape is that both keys are genuinely live and both are labelled as such, with a deadline on the old one.
Cancelling a door code does not retract the copies people wrote down; it only stops the lock accepting it — and only once every lock in the building has heard. The building's honest answer is 'within a minute', and it is the slowest lock that decides.
saying these in an interview costs you the question
- Says a revoked key stops working immediately, with caches in front of the check
- Treats a published invalidation as the guarantee instead of the cache lifetime
- Plans a single cutover instant for every partner, then staffs a support queue
- Keeps a retired key accepted while labelling it disabled
- Considers the incident closed once the key is revoked and replaced
- Guesses that the partner has cut over instead of reading the old key's traffic