Your mass credential reset covers 4,000 accounts but the helpdesk can re-verify 400 a day. How do you scope it?
answer
- ten days means ten days of exposure
- order by exposure, not by seniority
- they can answer the security questions too
- shrink the queue by disabling dormant accounts
- a named owner accepts the residual risk
basics
~10 sRank the population by privilege and evidence of adversary use rather than resetting everyone equally, verify identity out-of-band, apply compensating controls to the un-reset tail, and have a named executive accept the residual exposure.
solid answer
~50 sA ten-day reset at that rate is a decision about who stays exposed, so make it explicitly. Tier the population: identities with evidence of adversary use, then privileged and service identities, then the business functions the intrusion touched, then everyone else. Fix the verification path first, because a helpdesk re-verifying callers with knowledge questions is answering an adversary who has read the mailbox; require an out-of-band check they cannot satisfy, and use self-service flows only where the factors involved are not ones they control. For the un-reset tail, buy time with compensating controls rather than pretending they are safe: block legacy authentication, tighten conditional access, force re-authentication, and detect use of pre-reset material. Then have a named executive accept the residual risk and the disruption window, because the real trade-off is business availability against days of continued adversary access, and that is not the incident lead's call alone.
go deeper
Understand that a mass reset is rate-limited by identity verification, and that verifying who is on the phone is the hard part, not changing the password.
Be able to describe compensating controls for accounts still waiting: blocking legacy authentication, tightening sign-in policy and disabling dormant accounts.
Demonstrate a defensible tiering derived from the investigation, an out-of-band verification path, and completion measured per account rather than by tickets closed.
Own the risk acceptance: price the alternatives, put the exposure window and compensating controls in front of a named executive, and hold the line against a seniority-ordered queue.
## Why this is a decision and not a schedule With four thousand accounts and four hundred verifications a day, the reset takes about ten working days. That is not an operational detail, it is a statement that some population remains reachable with credentials the adversary may hold for up to ten days. Treating it as a queue to be worked alphabetically converts a security decision into an accident of surname ordering. The job is to choose who is exposed for how long, put compensating controls under them, and have the right person accept that. ## Tiering the population Build the order from the investigation, not from the org chart: 1. **Identities with evidence of adversary use** in the audit trail, plus any identity whose credential material was in reach of what they accessed. 2. **Privileged identities**, including administrative accounts, break-glass accounts and anything holding directory or cloud control-plane rights. 3. **Service and application identities**, which need separate handling because rotating them breaks software, so app owners must be lined up first. 4. **The business functions the intrusion touched**, where lateral movement is most plausible. 5. **Everyone else**, on the assumption that the estate-wide reset is precautionary. A common political failure is executives first regardless of exposure. Seniority is not exposure, and a VIP-first queue delays the accounts that actually carry risk. ## Fix the verification path before the volume Re-verification is the step the adversary will attack, because it is the moment credentials are handed out over the phone. If the helpdesk verifies by knowledge (employee number, manager's name, recent activity), the adversary who has been reading a mailbox for weeks answers better than the real employee. During an active intrusion the verification must be out-of-band relative to anything they control: a manager confirming on a video call they initiate, an in-person check for the highest tier, or a code delivered through a channel that is not the compromised platform. Self-service reset is a throughput multiplier and worth using, but only where the factors it relies on are ones you have already audited for adversary-registered devices. Warn the helpdesk explicitly that a surge of urgent reset calls is an expected adversary behaviour during a mass reset, and give them an escalation path rather than a target handle time. ## Compensating controls for the tail The un-reset population is not safe, but it can be made less useful. Block legacy authentication protocols that bypass modern policy; tighten conditional access to compliant devices or known networks for the highest-risk applications; force sign-in frequency down so tokens have to be re-issued; disable dormant accounts outright rather than resetting them, which costs nothing and shrinks the queue; and run a detection specifically for authentication using pre-reset material or from the infrastructure attributed to the intrusion. Disabling the long tail of unused accounts is often the single biggest reduction available, and it needs no helpdesk time at all. ## The decision that is not yours Someone has to accept that the estate carries known residual adversary access for the duration. That is an executive risk acceptance, documented, with the numbers in it: how many accounts, for how long, with which compensating controls, and what would change the plan. The alternative options should be priced rather than assumed away: surging capacity with contractors or by deputising team leads as verifiers, accepting a partial outage by disabling rather than resetting a tranche, or extending the window. Present the trade-off as availability against dwell time, give a recommendation, and let the accountable owner choose. Where a regulator or a contractual notification clock is running, that constraint belongs in the same paper. ## Knowing when it is actually done Completion is not tickets closed. The measure is per-account: each identity has authenticated at least once with post-reset credentials, and nothing has authenticated with pre-reset material since. Accounts that never come back are a finding in their own right, because an account nobody claims is either dormant, which means disable it, or belongs to someone who is not the person you think. Report the two numbers separately and keep the detection running after the incident closes.
- The helpdesk verifies callers with knowledge questions. Why is that dangerous here?Because the adversary has been inside the mailbox and the directory and can answer employee number, manager and recent-activity questions more fluently than the genuine user. During an active intrusion, verification has to be out-of-band relative to anything they hold: a manager-initiated video confirmation, an in-person check for privileged accounts, or a code sent through a channel the intrusion did not touch.
- How would you shrink a 4,000-account reset queue without adding capacity?Disable rather than reset the dormant tail, which removes accounts from the queue at no helpdesk cost and eliminates the exposure entirely. Route low-risk users through a self-service flow whose factors you have already audited for adversary-registered devices. Handle service identities on a separate track with their application owners. What remains is a much smaller population that genuinely needs a human verifier.
- What tells you the mass reset actually completed?Per-account evidence, not queue metrics: each identity has authenticated with post-reset credentials, and no authentication has succeeded with pre-reset material since the cut. Accounts that never re-authenticate are themselves a finding to chase down. Keep the detection for pre-reset material running after the incident closes, since a patient adversary simply waits out the attention.
- Who signs off on leaving part of the estate un-reset for a week?A named accountable executive, in writing, with the numbers in front of them: how many accounts, for how long, under which compensating controls, and what would trigger a change of plan. The incident lead's job is to price the alternatives, recommend one, and make the residual risk legible. Taking that decision silently inside the security team is the failure mode.
saying these in an interview costs you the question
- Works the reset queue alphabetically or by seniority
- Lets the helpdesk verify identity with knowledge questions
- Calls the reset complete when tickets are closed
- Leaves the un-reset tail with no compensating controls
- Accepts multi-day residual exposure without an executive decision