skip to content

How does an exploit that crashes the service one attempt in five limit a worm's spread?

level: middleimportance: should knowfreq 45%

answer

  1. no operator is watching to retry
  2. the pool is shared between all infected nodes
  3. only the first attempt against a host counts
  4. candidates reached times success rate
  5. below one, each generation shrinks

basics

~20 s

Every failure subtracts a host from the pool that all infected nodes share. A host killed on the first attempt against it is never recruited by anyone, so per-attempt success caps total reach and cuts the number of hosts each infected node can pass the code to.

solid answer

~50 s

Reliability is a precondition of self-spread, not polish. Self-spreading code has no operator to retry, and on many services a failed attempt kills the listener - so the failure does not merely waste a try, it permanently removes a candidate from the pool every infected node is drawing from. A host survives to be infected only if the first attempt made against it succeeds, so with 80 percent per-attempt success roughly a fifth of the reachable set is destroyed rather than recruited, and those dead hosts never carry the code onward either. What decides whether spread continues is the branching factor: candidates each infected host can reach, times the success rate. On a dense flat segment that product stays comfortably above one and the crash rate only caps coverage; where each host finds only a handful of live candidates, it drops below one and propagation dies out.

go deeper

for a junior

Know that self-spreading code has no operator to retry a failed attempt, and that a failure which crashes the target service removes that machine from the set of hosts the code could ever reach.

for a middle

Explain the shared-pool argument and the first-attempt rule, and give the branching-factor product - candidates reachable per host times success rate - as the thing that decides whether spread continues or dies out.

for a senior

Bring the caveats unprompted: automatic restarts partly heal the pool, credential-based spreading has a non-destructive failure mode, and the crash rate matters most where each host reaches few live candidates.

for a principal

Be able to argue why a funded operator may reject self-spread precisely because the reliability term is uncontrollable across a heterogeneous estate, and what they buy instead with operator hours.

## Why a hands-on operator tolerates unreliability and self-spreading code cannot An operator working host by host treats a technique that works four times in five as perfectly usable. They watch the attempt, and when it fails they retry, pick a different route, or leave that host alone. The unreliability costs them a few minutes. Self-spreading code has none of that. There is no one to notice the failure, no fallback route unless it was engineered in advance, and - the part candidates usually miss - the failure is frequently **destructive**. Techniques that abuse a memory-handling defect in a listening service typically leave the service dead when the attempt lands wrong: the process crashes, and until something restarts it the host stops listening. The attempt did not just fail; it removed the target. ## Poisoning a shared pool The crucial structural point is that every infected host is drawing candidates from the **same** finite set - on a flat unmanaged segment, the couple of hundred devices in the broadcast domain. When one infected device crashes a printer's listening service, that printer is no longer available to any of the other infected devices either. The loss is not local to the node that made the attempt. That gives a simple and slightly counter-intuitive result. A host is recruited only if the **first** attempt made against it succeeds; a first attempt that fails kills the service, and later attempts by other nodes find nothing to talk to. Retrying does not help. So with per-attempt success `p`, roughly `p` of the reachable set ever runs the code and roughly `1 - p` is destroyed. At `p = 0.8`, about a fifth of the segment is turned into wreckage instead of into forwarders. ## Whether it merely caps reach or actually stalls What decides between "caps coverage" and "stops" is the **branching factor**: how many live candidates one infected host can reach, multiplied by the success rate. - **Dense range.** Each infected device can reach a couple of hundred live neighbours. Even at `p = 0.8` the branching factor is enormous, so the crash rate costs reach - the fifth that died - and some speed, but the wave still crosses the segment. - **Sparse range.** Each infected host finds only a handful of live candidates, and some of those have already been taken by another node. Now the product `p x k` is close to one, and a crash rate that removes a fifth of the pool can push it under one. Below one, each generation is smaller than the last and propagation dies out on its own. This is why unreliability is a governing term rather than a detail: it does not scale the spread down smoothly, it moves it across a threshold. ## Two nuances worth raising unprompted **Failure is not always permanent.** Embedded and appliance-class devices - the ones filling a flat unmanaged segment - often run their listeners under a supervisor or a hardware watchdog that restarts them, sometimes within seconds. There the crash costs a window rather than a host, and the pool partially heals. Conversely a device that needs a physical power cycle, on a segment nothing reaches to restart it remotely, is gone for the entire campaign. **A crash may not be the failure mode at all.** Where the technique is a valid login with a carried credential, a wrong guess costs nothing but an authentication failure and possibly a lockout counter on that account. The target keeps listening, so the pool is not poisoned - one reason credential-based spreaders can be far more relentless than exploit-based ones despite being technically unremarkable. ## How to answer this in the room The strong answer names three things in order: failures subtract from a **shared** pool rather than from one node's list; recruitment depends on the **first** attempt against a host, so retries do not recover it; and whether that ends the spread depends on the **branching factor**, not on the crash rate alone. Then add the honest caveat about restarts. What weak answers do instead is treat reliability as a quality issue - "they would tidy that up before release" - which misses that the failure rate is arithmetic acting directly on the propagation, not an engineering blemish.

  • Why do retries not recover a host that a failed attempt crashed?
    Because the crash removes the listener, not the attacker's chance. A later attempt from a different infected node finds nothing accepting connections on that port, so it fails for a new reason. The host is recruited only if the first attempt made against it - by anyone - succeeded, which is why per-attempt success rate caps total reach.
  • How does the picture change when the spreader authenticates instead of exploiting?
    A failed authentication normally leaves the service running, so the pool is not poisoned and repeated attempts stay possible. The cost moves elsewhere: account lockout thresholds, and the fact that the credential either works estate-wide or not at all. Reliability stops being the limiting term and reachability of the service becomes the binding one.
  • Where does a device watchdog change the calculation?
    It converts a permanent loss into a temporary one. Appliance-class devices often restart a crashed listener automatically within seconds, so the candidate returns to the pool and a later attempt can succeed. The spread is slowed rather than capped. A device that only comes back after a manual power cycle is a genuine permanent subtraction.

saying these in an interview costs you the question

  • Treats reliability as polish rather than a precondition
  • Assumes a later retry can recover a crashed target
  • Thinks each infected host has its own private target pool
  • Says a 20 percent failure rate simply makes spread 20 percent slower
  • Ignores that crashed hosts also stop being forwarders

context