skip to content

How would a web crawler detect and contain a spider trap, such as a calendar page whose next-month links never end?

level: seniorimportance: should knowfreq 45%

answer

  1. every generated URL looks new
  2. calendars, session IDs, loops, facets
  3. length, depth, repeated segments
  4. many URLs, one fingerprint
  5. per-host budget as backstop

basics

~20 s

Assume no host deserves unlimited pages: cap each host's page budget and each URL's length and depth, flag repeating path segments and session-ID parameters, and watch for hosts whose many URLs share content fingerprints. Suspect hosts get lower priority or are cut off.

solid answer

~40 s

A spider trap is a host that generates an effectively unbounded URL space: calendars with endless next links, session IDs added to every link, relative-link loops like `/a/b/a/b/`, faceted filter combinations, or soft 404s that return 200 for any path. The seen-URL set does not help, because every generated URL is new. Defences are layered: **structural limits** (maximum URL length, path depth, repeated segments, parameter count), **content signals** (many URLs with the same or near-duplicate fingerprint, or a probe showing the host answers nonsense paths with 200), **learned normalisation** that strips parameters which never change content, and a **per-host page budget** scaled by the host's importance. The budget is the backstop, because it bounds the damage from trap shapes nobody anticipated. The trade-off is that tight budgets under-crawl large legitimate sites.

go deeper

for a junior

Recall the classic trap shapes, such as endless calendars and session IDs in URLs, and why they keep a crawler busy forever.

for a middle

Explain why URL dedup cannot stop a trap and name the structural limits that reject runaway URLs at admission.

for a senior

Show how you would combine content-fingerprint yield, soft-404 probes, learned parameter rules and per-host budgets, and what you would alert on.

for a principal

Own the budget policy: how importance and observed yield set each host's allotment, and the coverage you accept losing to keep traps bounded.

## What a spider trap is A **spider trap** (or crawler trap) is any part of a site that produces an unbounded, or practically unbounded, number of distinct URLs. A crawler that follows links naively can spend its whole budget inside one such site. Most traps are accidental rather than malicious. Common shapes: - **Infinite calendars.** Every month page links to the next month, forever. - **Session IDs in URLs.** The server appends a fresh session token to every link, so each visit mints new URLs for the same pages. - **Relative-link loops.** A misconfigured relative link produces `/a/b/`, `/a/b/a/b/`, `/a/b/a/b/a/b/` and so on. - **Faceted navigation.** Filters for colour, size, price and sort order multiply into millions of combinations of the same catalogue. - **Soft 404s.** The server returns HTTP 200 with a generic page for any path, so broken or generated links never fail. ## Why the usual dedup does not help The **seen-URL set** only stops the crawler from fetching the same *address* twice. A trap produces a new address every time, so every URL passes the check. Politeness does not help either: it spreads the fetches out in time but still performs them all. Traps need their own defences. ## Layered defences | Layer | Examples | Catches | Weakness | |---|---|---|---| | structural limits | max URL length, max path depth, repeated path segments, max query parameters | loops, runaway paths | pattern-specific; misses traps with short URLs | | URL pattern rules | strip known session and tracking parameters, learned per host | session IDs, reordered filters | must be learned or configured | | content signals | many URLs with the same or near-duplicate fingerprint; soft-404 probe | generated variants, soft 404s | requires fetching some pages first | | per-host budget | cap on pages fetched per host per crawl cycle | anything, including unknown shapes | can under-crawl big legitimate sites | In practice they combine: 1. **Reject obvious runaways at admission.** URLs above a length or depth limit, or with the same segment repeated several times, never enter the frontier. 2. **Learn parameter rules.** When URLs on a host that differ only in one parameter keep returning identical content fingerprints, that parameter is stripped during normalisation for that host. 3. **Probe for soft 404s.** Fetch a random nonsense path on the host. If it returns 200 with a page whose fingerprint matches many other pages, the host's 200 responses cannot be trusted as proof of existence, and pages matching the probe are dropped. 4. **Watch the yield.** Track, per host, how many fetched pages are new content versus duplicates. A falling yield lowers the host's priority. 5. **Enforce a budget.** Every host gets a maximum number of pages per crawl cycle, scaled by an importance signal such as links from other hosts. ## Why the budget is the backstop Heuristics recognise trap *shapes*, and there are always shapes nobody has seen. A per-host budget does not care about shape: whatever a host does, it can consume at most its allotment. That bounds the damage from unknown traps to a known number of fetches. A depth limit from the seed is a weaker backstop. It cuts off legitimate deep sites, and a wide trap such as faceted navigation can generate millions of URLs only two or three links from a seed. ## Tuning the trade-off - **Too tight** a budget under-crawls large legitimate sites such as big forums, encyclopaedias or shops. - **Too loose** a budget lets a trap eat a large share of daily capacity. The usual answer is an **adaptive budget**: start from an importance-based allotment, raise it while the host keeps yielding unique content, and cut it when the duplicate rate climbs. Importance signals that are hard for a single site to inflate, such as links from many unrelated hosts, carry the most weight. ## Operating it - Alert on hosts whose queued URL count grows much faster than their fetched unique content. - Keep a reviewable list of per-host rules, since a learned rule that strips the wrong parameter silently loses pages. - Honour `robots.txt` disallow lines: many sites already disallow their calendar or filter paths, which avoids some traps entirely. The interview answer to lead with: *every generated URL is new, so dedup cannot help; bound each host with a budget, and use structure and content signals to spend that budget well.*

  • How can a web crawler detect that a host serves soft 404 pages?
    Request a random path that almost certainly does not exist. A well-behaved server answers 404; a soft-404 server answers 200 with a generic page. The crawler stores that probe page's content fingerprint and treats other pages on the host that match it as missing, so generated or broken links stop inflating the crawl.
  • Why is a per-host page budget a safer backstop than a maximum link depth from the seed?
    A depth limit cuts off legitimate deep content and still lets a wide trap, such as faceted filters, generate millions of URLs within two or three hops. A per-host budget bounds total fetches regardless of the trap's shape, so the worst case from any unknown trap is a known number of wasted requests.
  • How do you keep per-host budgets from starving a large legitimate site?
    Scale the budget by an importance signal that one site cannot easily inflate, such as links from many unrelated hosts, and adapt it: raise the allotment while fetched pages keep yielding unique content, and cut it when the duplicate rate rises.

It is like exploring a hall of mirrors: every reflection looks like a new room, so instead of trying to recognise each illusion you decide up front how many rooms you will enter before leaving.

saying these in an interview costs you the question

  • The seen-URL set already protects the crawler from spider traps.
  • A maximum link depth from the seed is enough to stop every trap.
  • Spider traps are always malicious; honest sites never generate them.
  • Every host should get exactly the same page budget.
  • An HTTP 200 response proves the page exists and has unique content.