A critical advisory hits 180 services you already know are affected. How do you order the fixes in the first hour?
answer
- the same score on every row
- flaw property versus system property
- who can reach it, with what input
- bands: public, tenant, internal, batch
- most of the list rides the normal train
basics
~20 sOrder by exposure, not by score. Services that are internet-reachable and feed attacker-controlled input into the vulnerable code path go first, then internally reachable ones, then batch and build-only workloads. The score is identical across all 180, so it cannot rank them.
solid answer
~50 sThe severity score is a property of the flaw, so it is the same number on all 180 rows and ranks nothing. What differs between the services is exposure, and in the first hour you can establish that with coarse, checkable facts rather than analysis. I use four bands: first, anything an unauthenticated internet client can reach that passes attacker-controlled data into the affected component; second, anything reachable by a low-privilege authenticated tenant, or holding credentials and data whose loss is the worst outcome; third, internal service-to-service traffic where an attacker would already need a foothold; fourth, batch jobs, dev-only and build-time dependencies. The first band gets emergency handling — mitigate now, patch out of band. The tail, which is usually most of the 180, gets the same version bump on the normal release train. Deeper analysis of whether the vulnerable code is actually called refines this later; it does not gate hour one.
go deeper
Know that a severity score describes the flaw, not your deployment, and be able to name the first question that ranks two affected services: can a stranger on the internet reach it with data they control?
Explain a concrete band ordering and the cheap yes/no facts that place a service in a band, and say what the majority of affected services should get instead of emergency handling.
Demonstrate that you can produce and communicate a defensible order in the first hour without waiting for analysis, including where a high-value internal service breaks the simple reachability ordering.
Own the cost side: emergency handling spends reviewer attention and team goodwill, so be ready to argue how much of an estate should ever be in the top band and what that budget buys.
## Why the score cannot order the queue When one advisory lands in one widely-used component, every affected service inherits the same severity rating. A severity score describes intrinsic characteristics of the flaw — how it is reached, what privileges it needs, what it costs you if it works — under assumptions about a generic deployment. It deliberately says nothing about *your* deployment: not whether the affected service is on the internet, not whether it holds anything worth stealing, not whether the vulnerable function is ever invoked. So a list sorted by score is a list in arbitrary order, and treating it as a work queue means the first thing you fix is whichever row the dashboard happened to render first. The useful ordering key is **exposure**: how much attacker effort stands between an attacker and this instance of the flaw, and what is behind it. ## The four bands Write this ordering rule down before an incident, on one page, so that during one you apply it instead of debating it. **Band 1 — internet-reachable and fed untrusted input.** An anonymous client on the internet can send data that reaches the affected component. Nothing else competes with this. These get mitigated within the hour, patched out of band, and watched. **Band 2 — authenticated but low-trust, or high-value.** Reachable by any signed-up tenant, or by a partner integration. Also in this band: services whose compromise is not about the service itself but what it holds — credential stores, signing material, the systems whose logs and records are the audit truth you would rely on after the incident. Attacker effort is higher; consequence is higher too. **Band 3 — internal service-to-service.** Reachable only from inside, so exploitation requires an existing foothold. Real, but it is the second move in an attack rather than the first. **Band 4 — batch, offline, build-time and dev-only.** Jobs that process data you produced, tooling that never sees a stranger's bytes, test-scope dependencies. These almost never justify an out-of-band release. ## The facts you can actually establish in hour one The bands are only useful if you can assign services to them fast. Use facts that are cheap and checkable: - Does this service have a public route at all? - Does the affected component sit on a request path, or on a startup/admin path? - Does it process a format an attacker controls — the request body, an uploaded file, a URL, a header? - Is authentication required before that code runs? - What credentials or data does the process hold if it is taken? These are yes/no questions the owning team can answer from memory. Precise analysis of whether the vulnerable function is genuinely invoked, and whether an attacker can steer input into it, is a real discipline and it will sharpen this ranking — but it takes longer than the first hour and it is not available for every language. Do not stall the response waiting for it. Coarse and immediate beats precise and late. ## What happens to the tail Most of the 180 will land in bands 3 and 4. The correct outcome for them is *not* an emergency: it is the same version bump riding the normal release train with normal tests. Emergency handling is expensive — it consumes reviewers, it skips soak time, and it burns the willingness of teams to respond next time. Spending it on a batch job is how an organisation trains itself to ignore the next real one. Equally, the tail must not be dropped. "Not urgent" has to mean scheduled, with the same bump, not silently deferred. ## The orderings that quietly go wrong - **By team responsiveness.** You patch the services whose owners answered the page, which correlates with nothing. - **By ease.** The dev-only dependency is a one-line change, so it gets done first and the queue looks like progress. - **By dashboard order or alphabetically.** The default sort becomes the plan. - **By whoever escalated loudest.** The executive's favourite service is not necessarily the exposed one. - **By count of findings per service.** A service with forty low-relevance rows outranks the one internet-facing service with one. ## Communicating the order Publish the ranking, not just the work. Teams in band 4 need to know they are in band 4 and why, or they will either panic-deploy on a Friday or conclude the security team cannot prioritise. A one-line rationale per band — "public endpoint, untrusted body, affected parser on the request path" — is what turns an arbitrary-looking list into an order people follow. ## The judgment being tested An interviewer asking this is checking whether you understand that severity is a property of the flaw and priority is a property of your system, and whether you can produce a defensible order under time pressure without a tool telling you the answer. The strongest answers state the rule in one sentence, name the cheap facts that assign services to bands, and say explicitly what the majority of the list *does not* get.
- Why does a CVSS base score rank nothing when one advisory affects many of your services?A base score rates intrinsic characteristics of the flaw under a generic deployment assumption, so every affected service inherits the identical number. It is not environment-aware by design, which is exactly why it cannot distinguish an internet-facing gateway from a nightly batch job. Ordering has to come from facts about your deployment, not from the advisory.
- One of the 180 is a service that holds signing keys but is internal-only. Where does it go?Above the internal band, near the top. Exposure is attacker effort times consequence, and this one has a low-probability path to a very high-value asset — compromising it undermines every artifact you have vouched for. Purely internal reachability lowers the odds; it does not lower the stakes, and the ordering rule has to have room for that.
- A team says their service is unaffected because the vulnerable function is never called. Do you accept it in hour one?Record it and move them down the queue, but do not close it on an assertion. In the first hour you are taking coarse facts on trust to move fast; the claim still needs evidence afterwards, and it needs to survive the next refactor that starts calling the function. Accepting unverified not-affected claims permanently is how a real exposure disappears.
saying these in an interview costs you the question
- Sorts the queue by severity score when every row has the same one
- Starts with the easiest fixes to show progress
- Waits for full reachability analysis before starting anything
- Treats all 180 as an emergency and freezes every release train
- Deprioritises the tail into silence rather than into the next release