What is the browser's preload scanner (also called the speculative or lookahead parser), and which resources on a page can it not discover?
answer
- a second, dumber parser
- keeps the network busy while blocked
- only sees literal URLs in markup
- CSS and JS references are invisible
- preload puts the URL back in the head
basics
~20 sThe preload scanner is a secondary, lightweight parser that scans the raw HTML ahead of the main parser and starts downloading resources it finds in markup attributes. It cannot see anything that only exists after CSS or JavaScript runs.
solid answer
~50 sWhen the main HTML parser stalls — typically on a synchronous script — the browser does not leave the network idle. A second, much simpler parser scans the not-yet-parsed HTML bytes looking for URLs in attributes such as `src` and `href` on `<img>`, `<script>`, and `<link>`, and kicks off those fetches speculatively. It builds no DOM and executes nothing; it only warms the network. The consequence for authors is that anything not visible as a plain URL in the markup is invisible to it: images referenced from CSS such as `background-image`, resources injected by JavaScript, URLs assembled at runtime, and lazy-loading patterns that hide the real address in `data-src`. Those cannot start downloading until CSS or JS has run. When a genuinely critical resource is hidden that way, `<link rel="preload">` in the head puts a real URL back in the markup for the scanner to find.
code
html · 9 lines<!-- Found by the preload scanner: a literal URL in a src attribute. -->
<img src="/hero.jpg" width="1200" height="600" alt="Product hero">
<!-- Invisible to the scanner: the real address is hidden in data-src. -->
<img data-src="/hero.jpg" class="lazy" alt="Product hero">
<!-- Invisible to the scanner: the URL only exists once the CSSOM is built. -->
<style>.hero { background-image: url("/hero.jpg"); }</style>
<div class="hero"></div>go deeper
Know that the browser looks ahead in the HTML to start downloads early, and that only URLs written directly in the markup get that head start.
Explain that it is a separate lightweight parser that runs while the main parser is blocked, builds no DOM, and misses anything reachable only through CSS or JavaScript. Name preload as the fix.
Read a waterfall and attribute a late-starting critical resource to scanner blindness rather than to server latency. Justify which few resources deserve a preload hint and what over-preloading costs.
Be ready to argue how a rendering architecture affects discoverability at all — markup that ships real URLs versus a shell that builds every reference from script — and who owns the guardrail that stops preload hints proliferating until they compete with each other.
## The problem it solves HTML parsing is not a smooth process. The main parser stops whenever it hits a classic synchronous script, because that script can modify the document it is being parsed into. In the early days of the web that stall meant the network went quiet too — the browser knew nothing about the twelve images and three stylesheets sitting further down the document, because it had not read that far. The preload scanner exists to remove that dead air. It is a second parser, deliberately dumb and fast, that runs over the raw bytes ahead of the main parser's position. ## What it actually does The scanner tokenizes markup looking only for attributes that name a resource — chiefly `src` on `<img>` and `<script>`, `href` on `<link>`, plus related ones like `srcset` and `poster`. When it finds one it starts the fetch immediately, so by the time the main parser reaches that element the bytes are already in flight or in the cache. What it explicitly does **not** do: - It does not construct DOM nodes. - It does not execute scripts or evaluate CSS. - It does not run the cascade, so it has no idea which elements will be visible. It is a speculative optimisation. If the main parse later diverges — a script rewrites part of the document — some of those fetches were simply wasted bandwidth, which is a trade the browser accepts. ## What it cannot see This is the interview-relevant half, because it explains a whole class of "why did my hero image start downloading so late" bugs. **Resources referenced from CSS.** A `background-image` URL lives inside a stylesheet, which the scanner does not parse. That image cannot even be discovered until the stylesheet has been downloaded and the CSSOM built, and the element that uses it has been matched. Two serialised round trips before the request starts. **Anything injected by JavaScript.** An image element created with `document.createElement('img')` and given a `src` in a bundle does not exist in the markup at all. The download begins after the bundle has downloaded, parsed and executed. For a client-rendered app, this is why the largest image on the screen is often the last request to start. **URLs assembled at runtime.** `img.src = base + '/' + id + '.jpg'` is not a URL in the source; it is string concatenation. There is nothing for a byte scanner to match. **Lazy-loading patterns that hide the address.** The classic library idiom writes `data-src="real.jpg"` and swaps it into `src` from script. That deliberately defeats the scanner, which is the point for below-the-fold images and a serious own-goal for above-the-fold ones. Note that the native `loading="lazy"` attribute is different — the real URL stays in `src`, and the browser itself decides when to fetch. **Imports inside CSS.** An `@import` in a stylesheet is invisible for the same reason `background-image` is: it lives in a file the scanner never reads. ## Putting the URL back where it can be seen The standard repair is a preload hint in the head: ```html <link rel="preload" as="image" href="/hero.jpg" fetchpriority="high"> <link rel="preload" as="font" href="/inter.woff2" type="font/woff2" crossorigin> ``` This is a real URL in the markup, so the scanner finds it in the first bytes of the document and starts the fetch immediately; the response is then waiting in the memory cache when CSS or JS finally asks for it. The `as` attribute is not optional cosmetics — it sets the request's destination, and therefore its priority and its CORS handling. Fonts always need `crossorigin` because font fetches are made in CORS mode, and a mismatch causes the resource to be downloaded twice. Preload hints are a scarce resource: every hint you add competes for bandwidth with the things the scanner would have found anyway. A handful of genuinely critical, otherwise-invisible resources is the right dose. ## How to observe it In a network waterfall, scanner-discovered requests all start very near the beginning of the document download, in a cluster, even when the elements they belong to are far down the page. A resource that starts only after a script or stylesheet has finished is one the scanner missed — and that gap is the cost you are being asked to explain.
- Does the preload scanner build DOM nodes for what it finds?No. It only tokenizes enough to spot resource URLs and start fetches; the DOM is built exclusively by the main parser. That is why its work is safely discardable if a script later rewrites the document — at worst the browser wasted a download.
- Why does an image set as a CSS background-image start downloading so much later than the same image in an img tag?The scanner reads markup, not stylesheets. The URL is only discovered after the stylesheet is downloaded and parsed into the CSSOM and the rule is matched against an element, so two dependencies serialise ahead of the request. An img src is a literal URL in the first bytes of the document.
- When is a preload hint the wrong tool?When the resource is already discoverable in markup — you gain nothing and add priority contention — or when it is not actually needed for the first screen, in which case you are stealing bandwidth from things that are. Over-preloading measurably slows pages down.
saying these in an interview costs you the question
- Thinks the preload scanner executes scripts
- Believes it parses stylesheets for image URLs
- Says data-src lazy loading is free with no download cost
- Confuses rel=preload with rel=prefetch for the current page
- Assumes it builds part of the DOM ahead of the parser