A revoked contractor's flow to an internal database was still moving bytes 40 minutes later: how do you close that gap without re-authenticating the whole estate hourly?
answer
- four causes: decision, delivery, execution, the flow
- pooled connections are admitted once and kept warm
- keepalives mean no idle timer ever fires
- bound session age, not the human
- measure to last byte, not console timestamp
basics
~20 sTrace the chain: where the decision changed, what learned about it, and what was actually torn down. Long-lived pooled flows with keepalives are never re-decided, so cap session age at the enforcement point and push revocations — and scope aggressive re-decision to high-value applications, not everything.
solid answer
~60 sFirst establish where it broke, because "revocation is slow" has four different causes. Was the decision only made at the directory and never delivered to the enforcement point? Did the enforcement point learn it on a poll interval? Did it act on its session table but on a session it did not own? Or was the flow simply never re-evaluated — a pooled database connection, established once, kept warm by keepalives so no idle timer ever fired? The 40-minute survival points at the last one. The fixes are: a **maximum session age** at the enforcement point, not just an idle timeout, so every flow is forced to re-ask eventually; **event-driven revocation push** to enforcement points with a poll as the floor when the push fails; and re-decision intervals **tiered by what the application is worth** — minutes for the finance and production-data brokers, an hour for the intranet — so you buy speed where it matters instead of re-authenticating nine thousand people every hour. Then measure it: fire a synthetic revocation on a schedule and time it to the last byte, not to the console timestamp.
code
text · 5 linesipv4 2 tcp 6 431997 ESTABLISHED
src=10.14.7.31 dst=10.60.2.9 sport=51344 dport=5432
src=10.60.2.9 dst=10.14.7.31 sport=5432 dport=51344
[ASSURED] mark=0 use=1
...go deeper
Know that a revoked account can still have a live connection, and that someone has to time how long that connection survives rather than assuming it dies at once.
Explain the four links in the chain — decision, delivery, execution, and a flow nothing re-evaluates — and why keepalives defeat an idle timeout.
Demonstrate you would bound session age on the paths that matter, move revocation to push with a poll floor, and produce a measured revocation-to-last-byte number instead of a claim.
Be ready to argue where tighter re-decision is worth its disruption and where it is not, and who decides that classification and keeps it current.
## Diagnose before you design "Revocation to disconnect took 40 minutes" is a symptom with four candidate causes, and they have different fixes. Walk the chain in order: 1. **Decision.** When did the authoritative source actually mark the principal as revoked? Sometimes the answer is that a ticket was closed at 14:00 and the directory change landed at 14:25. 2. **Delivery.** How does the enforcement point learn? If it polls, the poll interval is a hard floor on your best case. If it is pushed, does a failed push get retried, and does anyone notice when it silently does not arrive? 3. **Execution.** Did the enforcement point terminate everything matching, or only sessions it directly owns? A session it proxies it can close; a flow that only passed through it at setup it may not be able to reach at all. 4. **The flow itself.** This is the usual culprit for a survival measured in tens of minutes. An application connection pool opens a small number of long-lived connections and reuses them for days. Keepalives ride the flow often enough that no idle timer ever expires, and the connection-tracking entry for an established TCP connection carries a timeout measured in days, so nothing ages it out. Nothing in that flow ever asks a policy engine a second question after the handshake that created it. ## The fixes, and what each one costs | Fix | What it buys | What it costs | |---|---|---| | Maximum session age at the enforcement point | Every flow is forced to re-ask eventually, bounding worst case | Long jobs get cut mid-run unless they can resume | | Event-driven revocation push, poll as floor | Best case drops from an interval to seconds | A delivery path that must be monitored, or it fails silently | | Tiered re-decision by application value | Speed where it matters, cheap everywhere else | Someone must classify applications and keep that current | | Terminate on the broker rather than asking the client | Evidence you control, independent of endpoint cooperation | Only covers traffic the broker actually carries | The instinct to re-authenticate everyone every fifteen minutes is the wrong lever, and it is worth saying why out loud in an interview: it multiplies decision load across the whole population to fix a problem confined to a few long-lived flows, and every extra prompt trains people to click through prompts. Bound the **session**, not the human. ## Measure it, or you do not have a number The number worth quoting is *revocation issued to last byte observed*, and almost nobody measures it that way — they measure to the console timestamp, which is the moment the request was accepted, not the moment the adversary lost the channel. To get the real one: on a schedule, admit a synthetic principal, open a long-lived flow through the same path as real traffic, revoke it, and time the last byte at a point the endpoint does not control. Run it on the awkward cases too — a client on a mobile tether, a client that was asleep — because the distribution's tail is what an incident will actually hand you. That measurement is also the answer to "prove the boundary denies". A screenshot of a disabled account proves a control-plane change. A timed synthetic revocation proves the data plane closed. ## What good sounds like "I'd first check whether the enforcement point ever learned, then whether it had state for that flow. Given 40 minutes on a database connection, I'd bet on a pooled connection nobody re-evaluates — so I'd add a maximum session age on that path and move revocation from poll to push, but only tighten re-decision on the high-value brokers, because doing it everywhere means nine thousand people re-authenticating on a cycle to fix a handful of long flows. Then I'd put a scheduled synthetic revocation in place so the number is measured rather than assumed."
- Why is an idle timeout a poor way to bound how long a revoked flow survives?Because an idle timeout only fires when nothing is happening, and the flows you most want to bound are the busy ones. Application keepalives, or an adversary simply using the channel, reset the timer indefinitely. A maximum session age fires regardless of activity, which is why it is the control that actually bounds the worst case.
- What would you measure, and against what, to say your revocation is fast enough?Revocation issued to last byte observed, sampled by a scheduled synthetic revocation through the real path, reported as a distribution with its tail — not a mean. Compare it against what the business says the exposure window may be for that class of application, and report the offline-client tail separately, because you cannot disconnect a device you cannot reach.
- Why not simply force every user to re-authenticate every fifteen minutes?It applies a whole-population cost to fix a problem in a handful of long-lived flows, adds decision load and support load, and habituates people to approving prompts — which degrades the signal you rely on when a prompt is genuinely suspicious. Bound the session's lifetime at the enforcement point instead, and tier the interval by what the application is worth.
saying these in an interview costs you the question
- Measures revocation from the console timestamp, not last byte
- Expects an idle timeout to close a busy flow
- Proposes estate-wide re-authentication as the fix
- Overlooks a connection pool holding one flow open for days
- Assumes the push reached the enforcement point without checking