Your secret store's sessions back a fleet of unattended workloads — what containment time can you promise, and what does promising it cost?
answer
- containment is a number you sign
- shorter of revocation and expiry
- shorter windows, many more authentications
- the authentication path becomes critical
- measure it in a drill, per caller class
basics
~10 sPromise the worst case, not "immediately": containment time is the revocation path's own latency where sessions can be ended on demand, and the full remaining validity where they cannot. Shortening that window multiplies authentications.
solid answer
~50 sContainment time is a number you sign for, and it is the **shorter of two paths**: how long it takes a revocation to take effect, and how much validity a session has left if no revocation is possible. Where the store is consulted per call, the first path dominates and the promise is minutes. Where the proof is self-contained, the second path *is* the promise, so the only way to shorten it is to issue shorter windows — and that is the cost. Take four hundred unattended workloads on eight-hour sessions: roughly 1,200 authentications a day. Cut the window to fifteen minutes and it becomes 38,400, about 27 a minute, which makes the authentication path fleet-critical and its outage a fleet outage. The honest deliverable is a written worst case per class of caller, measured in a drill rather than asserted.
go deeper
Understand that shorter-lived access limits how long a stolen credential is useful, and that it means the systems using it must re-authenticate far more often.
Explain both paths to containment — a revocation taking effect, or the window simply running out — and work out how often a fleet re-authenticates when you halve the window.
Quote a number you have measured in a drill rather than assumed, per class of caller, and make sure the runbook's first step is an action that works for that class.
Decide the estate standard and own its bill: revocability costs a dependency on every request; independence costs a containment window that cannot be shortened after the fact. Write the resulting promise down before an incident needs it.
## What you are actually being asked to promise Someone — an incident lead, a customer's security review, your own runbook — wants a number: *if a workload is compromised, how long until its access to the secret store ends?* This is a lead's question because it is not answered by a feature. It is answered by a design choice with a recurring bill attached. The number is the **shorter of two paths**: 1. **The revocation path.** Someone notices, decides, and issues the action; the store applies it; the caller's next evaluated request is refused. This exists only where there is something to end. 2. **The expiry path.** The session simply stops being accepted when its window runs out. This always exists, and it is an upper bound on everything. Where no kill path exists, path two *is* the promise. That is why the length of the window you issue is a containment decision and not merely an operational one. ## The arithmetic of shortening the window Assume four hundred unattended workloads, each authenticating when its session runs out. | Session window | Authentications per workload per day | Fleet total per day | Roughly per minute | |---|---|---|---| | 8 hours | 3 | 1,200 | under 1 | | 1 hour | 24 | 9,600 | about 7 | | 15 minutes | 96 | 38,400 | about 27 | The worst-case exposure falls from eight hours to fifteen minutes, and the authentication path goes from background noise to a component whose outage stops the fleet within a quarter of an hour. That is the trade, stated plainly. Note also what does **not** scale with it: the number of secret reads is driven by what the workloads do, not by how often they re-authenticate. ## The costs people forget to count - **The authentication path becomes critical.** At one authentication every couple of seconds, its availability target is now the fleet's availability target. - **Whatever the workloads authenticate *with* gets exercised constantly.** If the proof comes from an identity the platform signs for each workload, that signer is now on the same critical path. - **Detection noise changes shape.** A fleet re-authenticating every fifteen minutes makes a genuinely anomalous authentication far harder to see than one that happens three times a day. - **Recovery gets worse before it gets better.** After any interruption, a short-window fleet re-authenticates all at once rather than trickling back. ## What a defensible promise looks like - **Per class of caller, not one number for everything.** An unattended consumer and a person at a terminal deserve different windows and different answers. - **Measured, not asserted.** In a drill: end a live session by reference and time the interval until reads are refused. That measured interval, plus the longest window you issue, is what you may quote. - **Stated as a worst case with its residual.** "Within X minutes of the decision, and anything read before that stands." - **Explicit about the classes with no kill path**, where the promise is simply the window length. Naming that is more valuable than a number nobody can meet. ## The deeper choice underneath The two shapes trade the same quantity in opposite directions. Consulting a record on every call buys revocability and pays for it with a lookup on every request and a dependency that must be up. Verifying a self-contained proof buys independence — callers keep working through interruptions — and pays for it with a containment window nobody can shorten after the fact. A lead's job is to decide which one the estate standardises on, to write the resulting number down before an incident rather than during one, and to make sure the runbook's first containment step is an action that actually works in the shape that was chosen. A runbook that says "revoke the session" for a class of caller that has nothing to revoke is worse than no runbook, because it consumes the first ten minutes of an incident and produces a false sense that containment happened. One boundary worth stating to whoever receives the promise: this number is about access **to the store**. Credentials already fetched and in use elsewhere are separate objects, withdrawn in the systems that accept them, and they need their own number.
- Which callers deserve a shorter promise than the fleet default?The ones whose rights are broad or whose hosts are most exposed: anything reading across several branches of the name space, anything running where a person can reach the host, and any human session, which is short-lived by nature and cheap to re-establish. Steady unattended workloads with one narrow branch are the cheapest place to accept a longer window, because the blast radius is already small.
- How do you measure the revocation path rather than assume it?Drill it. Issue a session for a test workload that reads on a known interval, have someone end it using only a reference, and record the interval from the action to the first refused read. Repeat while the store is under normal load and, if reads can be served from more than one place, from each of them. The slowest observed interval is what you quote.
saying these in an interview costs you the question
- Promises instant revocation without naming the mechanism behind it
- Shortens the session window without costing the extra authentications
- Ignores the revocation path's own latency when quoting a number
- Treats containment time as a policy statement rather than a measured one
- Gives one promise for unattended workloads and people alike
- Writes a runbook step to revoke where nothing can be revoked