skip to content

In a Denial of Service review, why does unauthenticated expensive work dominate the findings?

level: middleimportance: must knowfreq 58%

answer

  1. compare what each side spends
  2. who can trigger it without an account
  3. the order of the checks matters
  4. declared size is not real size
  5. cheap checks first, hard caps after

basics

~20 s

One cheap request forces the server to spend seconds of CPU or gigabytes of memory, and any anonymous caller can send it. Fix the ordering - identity and cheap checks first - and cap what the expensive step may consume.

solid answer

~50 s

The classic finding is an upload endpoint that decodes and thumbnails an image before any authentication check. The attacker pays for one HTTP request; the server pays to decode a highly compressed image whose pixel buffer is orders of magnitude larger than the bytes on the wire. Anyone on the internet can trigger it in parallel, so the threat sits at the edge of the trust boundary and needs no account and no vulnerability. Two things fix it, and you need both. First, **ordering**: authenticate, authorise, then run the cheap checks (declared size, declared type, quota, format sniffing on a bounded prefix) and only then the expensive step. Second, **bounding**: cap decoded dimensions and memory, set a timeout, run the work in a bounded pool that rejects fast, and queue it rather than holding a request thread. Ordering shrinks who can trigger it; bounding limits what each trigger costs.

go deeper

for a junior

Be ready to say why an endpoint that does heavy work before checking who is calling is risky, and to name the ordering fix: authenticate and validate cheaply first, do the expensive step after.

for a middle

Explain the mechanics of cost asymmetry — compressed bytes versus decoded pixels, input-driven cost, holding a worker thread — and give both halves of the fix, ordering the checks and hard-capping the work.

for a senior

Demonstrate you enforce limits at the point of consumption rather than trusting declared metadata, and that you know which paths cannot be reordered, such as password hashing, and must be bounded and shed instead.

for a principal

Own the standard: entry points carry a stated cost budget and a documented cap before they ship, so this class of finding is caught by the design template rather than rediscovered in every review.

## The pattern Walk any design and mark, for every entry point, two numbers: what the **caller** spends to make the request, and what the **system** spends to serve it. Where the second is far larger than the first, and the request needs no identity, you have the highest-yield denial-of-service finding available at design time. It is the one worth writing first because it requires nothing of the attacker — no credentials, no privileged position, no flaw in the code. The feature working exactly as designed is the attack. The worked example: an upload path that generates a thumbnail before any authentication check. The request is a single POST of a small compressed file. The server decodes it into a raw pixel buffer whose size is width times height times bytes per pixel — a highly compressed image of a few megabytes can expand into gigabytes of resident memory. A handful of concurrent requests exhaust the heap; the process dies or the host starts swapping, and every other request on that instance dies with it. The cost ratio between the two sides is the whole finding. ## Why cost asymmetry, not request volume, is the metric Counting requests is the wrong lens. A design that can absorb ten thousand cheap requests per second may fall over on twenty expensive ones. During review, classify each operation by **input-driven cost** — cost that grows with something the caller controls: - Decompression and decoding (images, archives, video, documents) — output size is caller-chosen and unbounded relative to input. - Parsing that recurses or expands (deeply nested documents, entity expansion, huge collections). - Pattern matching whose runtime blows up on crafted input. - Deliberately slow cryptography — key derivation and password hashing are *designed* to be expensive. - Anything that fans out: one request that becomes many downstream calls, database rows scanned, or third-party API calls you are billed for. - Anything that holds a scarce handle for a long time: a connection, a lock, a worker thread, a temp file. Cloud spend belongs in this list. An unauthenticated endpoint that fans out to a metered third party denies service to your budget before it denies service to your users. ## The two independent fixes **Ordering.** Push cost behind checks, cheapest first. A defensible pipeline: terminate the connection with a size limit already applied; authenticate; authorise; validate declared metadata (content length, declared type, declared dimensions) and reject on mismatch; check the caller's quota; only then perform the expensive step. Each step earlier in the chain is cheaper than the one after it, so the further an attacker gets, the more they have had to prove. Ordering does not make the work cheap — it shrinks the population that can trigger it from *everyone* to *an authenticated principal you can name, throttle and revoke*. Be honest about where ordering cannot apply. Login is the standard counterexample: a password hash is expensive **by design** and must run before you know who the caller is. There the answer is bounding alone — keep the cost parameter at a defensible level, cap concurrent hashing with a bounded pool, reject immediately when that pool is full rather than queueing, and throttle per source and per targeted account so the rest of the API survives while the login path is under pressure. **Bounding.** Never trust declared metadata; enforce the real limit at the point of consumption. Cap decoded dimensions and total pixels, not just the byte length on the wire. Cap decompressed output and abort mid-stream when it is exceeded. Set a wall-clock timeout on the operation and on every downstream call it makes. Run it in a pool with a fixed size and a fast rejection path, so overload becomes a clean error rather than an out-of-memory kill. Where the work is not needed synchronously, accept, enqueue and return — the queue is where you apply per-tenant fairness and where you can shed load without dropping the interactive path. ## How to write the threat and the mitigation A useful entry names the element, the actor, the resource and the effect: *an anonymous internet caller posts a crafted image to the upload process, whose decode step allocates memory proportional to caller-chosen dimensions, exhausting the heap and denying uploads and all other traffic on that instance.* The mitigation then reads as an ordering claim plus a bound: *authenticate and check declared dimensions against a hard maximum before decoding; decode with an enforced pixel and memory cap inside a bounded worker pool with a timeout.* That pairing is what a reviewer is looking for. A candidate who answers only *add rate limiting* has bounded the count of requests, not the cost of any one of them — the single worst request still lands. ## Common wrong turns - Trusting `Content-Length` or a declared type as the size bound. The compressed size tells you nothing about the resident cost of the decoded form. - Treating it as a performance bug for later. Performance work targets the average request; this threat is about the worst request a stranger can choose. - Assuming autoscaling absorbs it. Autoscaling converts an availability problem into a bill, and scales too slowly against a burst anyway. - Adding a queue with no bound, so the memory pressure simply moves.

  • A login handler must run a slow password hash before it knows who the caller is. How do you keep that D risk down?
    You cannot move the cost behind authentication, so bound and meter it. Keep the hash parameters defensible rather than maximal, cap concurrent hashing with a fixed pool and reject immediately when it is full, throttle per source address and per targeted account, and shed load on that path so the rest of the API keeps serving. Fail fast beats queueing.
  • How do you decide during a design review that an operation counts as expensive?
    Compare attacker cost per request with server cost per request across CPU seconds, peak memory, disk bytes, downstream and third-party calls, and how long it holds a scarce handle. Treat anything with input-driven cost — decompression, decode, recursive parsing, pattern matching, fan-out — as expensive by default until someone states a hard cap.
  • Why is per-IP rate limiting an incomplete answer here?
    It caps how many requests arrive, not what each one costs, so the single most expensive request still executes in full. It is also easy to spread across many sources. Use it as one layer, on top of ordering the checks and putting hard caps on the work itself.

It is the postal equivalent of a reply-paid envelope: the sender spends a stamp, you pay for whatever they put inside. You either check who is writing before you open the parcel, or you cap how big a parcel you will accept.

saying these in an interview costs you the question

  • Says rate limiting alone closes the threat
  • Assumes a small Content-Length means a small decoded size
  • Counts only bandwidth, never memory or CPU
  • Believes only authenticated endpoints need quotas
  • Calls it a performance bug rather than a threat
  • Claims autoscaling makes exhaustion impossible

context