In what order do you threat-model an estate of 140 services with no existing models?
answer
- Prioritisation problem, not a modeling one
- Consequence times reach
- Pick the right asset for this business
- Tier the depth, name the tail
- Close the intake or it regrows
basics
~20 sRank by consequence and reach, not alphabetically: services that cross an external or tenant boundary, hold the most damaging assets, or sit on the critical path go first. Accept that the long tail never gets a full model.
solid answer
~50 sI would not try to model all of them. First build a coarse picture — what is reachable from outside, what holds sensitive or safety-relevant data, what everything else depends on — even if the inventory is rough. Then rank on consequence multiplied by reach. For a logistics operator with around 140 undocumented dispatch and routing services, the asset that matters is the availability of dispatch, so ranking by 'holds customer data' would put the genuinely critical services last; the right top tier is whatever stops trucks moving. Deep models go to that tier, a short triage pass to the middle, and the tail is explicitly out of scope until something changes. In parallel I would close the intake: new services must arrive with a model, or the backlog regrows faster than I can burn it down.
go deeper
Know that you cannot model everything and that some systems matter more. Being able to say 'start with what is exposed and what would hurt most' is enough at this level.
Be able to list concrete ranking inputs — external reachability, sensitivity of what is held, how many things depend on it — and explain why alphabetical or volunteer order wastes effort.
Show how you would build a usable ranking from imperfect data: what you can read from ingress configuration, data stores and on-call knowledge instead of waiting for an inventory.
Own the programme call. Commit to tiered depth, name the tail you will not model, close the intake so the backlog stops regrowing, and be ready to defend that promise to leadership.
## The shape of the problem Inheriting a large estate with no threat models is not a modeling problem, it is a prioritisation problem. A logistics operator with roughly 140 undocumented dispatch and routing services has more systems than a small security function can analyse in a year, no reliable inventory, and no way to know in advance which analysis will pay. Anyone who answers "we model all of them" has failed the question: at a realistic rate that programme finishes after the architecture has turned over, and it spends most of its effort on services where nothing much was at stake. ## Rank on consequence and reach The ranking inputs that actually predict value are: - **Consequence if compromised.** What is lost, and to whom. Crucially, this must be the *right* asset for this business. In a dispatch estate the dominant asset is the availability of dispatch — trucks not moving, deliveries not routed — not a customer record. A generic ranking that sorts by "holds personal data" would push the services that stop the business to the bottom of the list. Getting the asset class right is the single highest-leverage judgement in the exercise. - **Reach and exposure.** What can talk to it, and from where. External reachability matters, but so does a service that any internal caller can invoke without further authorisation, and so does one an operator or vendor can drive directly. - **Blast radius and centrality.** A shared routing or scheduling component that dozens of services depend on carries the consequences of all of them. Central components deserve a model well before leaf services. - **Rate of change.** A service under heavy active development is both more likely to acquire new boundaries and cheaper to fix, because the design is still moving. - **Existing evidence.** Prior incidents, known operational pain and the services engineers are visibly nervous about are free signal. Ask the on-call rota which system frightens them. ## Tier the depth, not just the order Sequencing alone still implies you will eventually reach number 140. Better to commit to different depths: - **Top tier:** a real session with a diagram, boundaries, enumerated threats and owned actions. This is a small number of systems — think tens at most, often fewer than ten to start. - **Middle tier:** a short triage pass answering only whether the service reaches anything that matters or holds anything that matters. Most come back as "no", which both de-risks and shrinks the queue honestly. - **Tail:** explicitly out of scope until a trigger fires. Writing this down is what turns an unfinishable backlog into a defensible decision. ## Close the intake, or the backlog regrows The strategic half of the answer is that ranking is worthless if new unmodeled services keep arriving. Put the trigger criteria into whatever gate new services already pass — the platform onboarding step, the design review, the production-readiness list — so that anything new enters the estate with a model. A programme that only burns down the backlog is running up an escalator. ## What you can and cannot get without a good inventory You will not have a clean inventory, and waiting for one is a way of doing nothing for two quarters. Derive what you can from what already exists: ingress and load-balancer configuration tells you what is externally reachable; data stores and their schemas tell you where the sensitive classes live; deployment metadata and on-call ownership tell you what is central and who to ask. A ranking that is roughly right today beats a perfect one that arrives after the reorganisation. ## How you know it is working Measure the top tier's coverage rather than the estate's. Two signals are worth watching: whether any service crossing an external or tenant boundary is still unmodeled, and how many genuinely new threats each session finds. When sessions stop finding new threats in the top tier, the marginal model is no longer the best use of the hour, and the effort should shift to keeping those models current and to the intake gate. ## The conversation with leadership Be honest about what is being promised. "We will model the systems whose failure stops dispatch or exposes the estate, and we will not model the rest" is defensible and fundable. "We will model everything" is neither, and it converts the first missed service in the tail into a broken promise rather than a known, accepted position.
- Why not start with the easiest services to build momentum?Modeling one easy service as a training run is fine, but the ordering itself must follow consequence. A programme that spends its first quarter on low-stakes services produces a stack of models and no reduction in risk, and when something goes wrong in the untouched top tier it is very hard to defend how the time was spent.
- What if you cannot get a reliable inventory of the 140 services?Rank on the evidence you can observe rather than waiting. Ingress and load-balancer configuration reveals external reachability, data stores reveal where sensitive classes sit, and on-call engineers will tell you which systems are central and which frighten them. A roughly correct ranking now beats an exact one that arrives two quarters late.
- How do you decide when to stop working the backlog?Watch the yield. When several consecutive sessions in the top tier surface no genuinely new threats, the marginal model has stopped paying, and the effort should move to keeping the existing top-tier models current and to the gate that stops new unmodeled services arriving. Stopping is a decision to state explicitly, not to drift into.
It is triage in an emergency department: you do not treat patients in arrival order, you sort by how bad the outcome is and how fast it is moving, and you say plainly that the minor cases wait.
saying these in an interview costs you the question
- Model every service before shipping anything else
- Start alphabetically or with whoever volunteers
- Internet exposure is the only ranking input
- Wait for a complete inventory before ranking
- Rank by data sensitivity even when availability is the asset
- Burn down the backlog without closing the intake