How would you choose the maximum session lifetime a platform team sets for every workload across an estate?
answer
- one number cannot fit every caller
- classify by who can renew
- the cost of expiry differs per class
- default, ceiling, expiring exception
- returns flatten below a few hours
basics
~20 sNot as one number. Classify callers by what an expiry costs them and what a leaked credential would cost you, set a short default with a documented ceiling, and grant exceptions that carry an owner and a review date.
solid answer
~50 sThe estate is the wrong unit, because the two costs that decide the number vary by caller. For an automated consumer that can renew, expiry costs nothing and the lifetime should be short. For an interactive operator, expiry costs a sign-in and the credential sits on a laptop, so short is also right. For a long unit of work that cannot renew, expiry costs the run — which is an argument to restructure the work, not to raise the ceiling for everyone. So publish a short default, a ceiling nobody exceeds silently, and an exception path with a named owner and a review date. Then watch the returns: the large reduction comes from having an expiry at all rather than none; below a few hours you are buying much less while paying real operational cost.
go deeper
Recall that credentials are issued with a lifetime and that somebody has to choose it; the shortest possible value is not automatically the right one.
Explain the two costs that decide the number — what an expiry costs a given caller and what a leaked credential from it would cost — and why they differ by caller class.
Show the operational consequences: renewal failures concentrated in one class, break-glass access that must still work when the issuing path is degraded, and the exception that quietly became permanent.
Own the non-linearity and the governance. Get every caller onto credentials that expire first, then set a default, a ceiling and an exception path with owners and dates, and measure how many callers still hold something that never expires.
## Why the estate is the wrong unit A single maximum applied everywhere optimises for the caller it fits worst. Pick the number from the longest-running consumer and you hand every laptop and every routine service the same long-lived credential. Pick the shortest the platform allows and you generate failures in exactly the callers that cannot renew — often the ones running at night, unattended, doing the work that matters. The decision is not one number. It is a **default**, a **ceiling**, and a named exception path, chosen from two costs that differ by caller: what an expiry costs that caller, and what a leaked credential from that caller would cost you. ## Classify the callers | caller class | what expiry costs | what a leak costs | lifetime | |---|---|---|---| | service that renews automatically | nothing — a renewal | bounded by the window | as short as the renewal path tolerates | | interactive operator | one sign-in | high: broad access, a laptop, a terminal history | short | | long unit of work that cannot renew | the run | bounded by the window | the argument for restructuring, not for a longer ceiling | | break-glass access during an incident | the incident response itself | very high: it is the most powerful access you have | short, but never so short it fails while the issuing path is degraded | | credential handed to an outside party | their integration | highest: it is outside your control entirely | shortest possible, with the narrowest permissions | The table is the actual work. Once the classes are written down, most of the argument about the number evaporates, because the disagreements turn out to be about which class a caller is in. ## Diminishing returns are the point The lever is strongly non-linear, and a lead who cannot say where it flattens will over-tighten: - from **never expires** to **expires at all** is the enormous step: an exposure that used to persist until someone noticed now closes by itself. - from **hours** to **minutes** is a much smaller step, because damage inside the window is dominated by what the credential is permitted to do and by how fast you detect misuse — neither of which the lifetime touches. - the operational cost runs the other way: shorter lifetimes mean more renewals, more code paths that must handle them, more callers that fail when the issuing path is briefly unavailable, and more dependence on that path during exactly the incidents when you need access. So the honest advice is: get everything onto credentials that expire, then argue about minutes. ## What the standard contains 1. **A default** that most callers get without asking, deliberately short. 2. **A ceiling** that no caller exceeds without a decision, so nobody quietly configures their way past it. 3. **An exception path** with three fields that matter: who owns this exception, why it exists, and when it is reviewed. An exception without a date is a permanent change wearing a temporary label. 4. **A classification rule** for new callers, so the discussion happens once per class rather than once per team. ## What to measure - How many sessions actually run near the ceiling. The usual finding is very few, which turns a contested number into an uncontested one. - The renewal failure rate, split by caller class. A rising rate is the early signal that the default is below what some class can sustain. - The number of live exceptions and their age. A list that only grows means the exception path has become the standard. - How many callers still present a credential that never expires. This is the number that actually correlates with risk, and it should be heading to a small, owned set. ## What the lever does not buy Be explicit with the teams you set this for, because over-claiming is how a control loses credibility. A shorter lifetime bounds **how long** a captured credential is useful. It does not bound **what** that credential may do — a broadly permitted short-lived credential does broad damage inside its window — and it does not detect misuse. It does not protect the long-lived identity that issues the short-lived credentials, which remains something to defend in its own right. And it does not retire the response drill for a credential found somewhere it should not be; it changes the drill's urgency from hours to the time remaining on the window.
- Where does shortening the lifetime stop paying?Well above the shortest value the platform allows. The large reduction comes from having any expiry at all; below a few hours the damage inside the window is set by what the credential may do and by detection speed, while renewal load, failure paths and dependence on the issuing path all grow. Spend the remaining effort on permissions instead.
- What do you do about a caller that genuinely cannot renew?Make it an exception with an owner, a reason and a review date, and pay for it by narrowing that caller's permissions hard. In parallel, treat the inability to renew as a defect in the caller: restructuring the work into resumable units usually costs less than carrying a long-lived credential indefinitely.
A hire car with a full tank and no return date is a different risk from one booked for the afternoon. Shortening the booking from a day to an hour changes much less than going from open-ended to booked at all.
saying these in an interview costs you the question
- Picks the shortest lifetime the platform allows and calls it finished
- Sets the estate's lifetime from its longest-running job
- Treats a shorter lifetime as a substitute for narrowing permissions
- Grants a permanent exception with no owner and no review date
- Assumes shortening the lifetime costs nothing operationally