A team marks every <script> on a page as deferred. Lab First Contentful Paint improves, but real users report the page is unresponsive for about a second after it appears, and Total Blocking Time barely moved. Why did deferring not fix that, and what would?
answer
- moved, not removed
- one burst right after parsing
- fast paint, frozen page
- execution drives blocking time
basics
~20 sDeferring changes when JavaScript runs, not how much runs. All that execution now lands in one burst right after the page paints, occupying the main thread exactly when the user starts interacting. Only shipping or running less code reduces blocking time.
solid answer
~50 sDeferring is a scheduling change, not a reduction. The same bytes are still downloaded and the same functions still execute on the main thread — they now execute together, just after parsing finishes, which is precisely the window in which the page looks ready and the user tries to click. So the paint metric improves and the blocking metrics do not, and the field responsiveness metric INP can even get worse because input now arrives during the burst instead of before it. The fixes are all about volume and necessity rather than attributes: delete tags nobody uses, cut the amount of JavaScript that runs at load, move code the first screen does not need behind the interaction that needs it, hold heavy vendor tags until the page is idle or the user engages, and break any remaining long initialisation into pieces that yield back to the browser between them.
code
javascript · 8 linesdocument.getElementById('open-editor').addEventListener(
'click',
async () => {
const { mountEditor } = await import('/editor.js');
mountEditor(document.getElementById('editor-root'));
},
{ once: true }
);go deeper
Remember the core fact: changing how a script is loaded does not change how much work it does. If a page feels stuck after it appears, the reason is JavaScript running, not JavaScript downloading.
Explain the two separate costs — download and execution — and match each to the metric it moves. Be able to say why all-deferred scripts land in one burst and why that timing is worse for the user than it looks in a lab report.
Demonstrate the diagnosis path: notice that the blocking metric did not move, profile the load, attribute main-thread time to individual scripts, then propose removal, conditional loading and yielding with expected effects. Interviewers want the reasoning, not a list of tips.
Frame it as a limit the team keeps rediscovering: attribute tuning has a ceiling, and past it the only lever is how much code the page is allowed to run. Be ready to describe the guardrails and the ownership conversations that keep startup work from creeping back.
## The confusion at the heart of it There are two separate costs in a script, and people who have only ever tuned attributes tend to collapse them into one: - **Discovery and download cost.** The bytes have to be found and fetched, and while they are fetching they compete with everything else the page needs. This is what loading attributes and hints affect. - **Execution cost.** Once the bytes arrive, the engine parses, compiles and runs them on the main thread. Nothing about how you asked for the file changes this number. Marking scripts as deferred addresses the first cost and leaves the second untouched. First Contentful Paint measures when something appears, so it improves. Total Blocking Time measures main-thread work that keeps the page from responding, so it does not. ## Why the burst is the worst possible timing When every script on the page is deferred, they all become eligible to run at roughly the same moment — after the document has been parsed. The result is a dense block of initialisation: frameworks booting, tag managers evaluating rules, analytics building payloads, widgets constructing DOM. That block lands immediately after the page has visibly rendered. From the user's perspective this is the worst arrangement available. The page looks finished, so they click, scroll or type — and the main thread is busy, so nothing happens for several hundred milliseconds. This is exactly what the field responsiveness metric, Interaction to Next Paint, captures, and it is why a page can hold a fast paint score and a poor real-user experience simultaneously. (INP became the responsiveness Core Web Vital in 2024, replacing First Input Delay, which only measured the delay before the *first* interaction's handler started and largely missed this pattern.) ## Diagnosing it honestly The symptom description already contains the diagnosis: a good paint number next to an unchanged blocking number is the signature of moved-not-removed work. To confirm and to assign blame, record a load profile and look at the main-thread work grouped by which script it came from. Two numbers per script matter: how long it occupies the thread, and whether anything on the first screen depended on it. The scripts that are expensive and unnecessary are your list. ## What actually reduces the cost **Delete.** The largest reliable wins on mature pages come from tags with no consumer. Nothing else improves bytes, execution, connections and failure modes simultaneously. **Ship less that runs at load.** Not all bytes cost the same: what matters is how much code actually executes during startup. A dependency that registers dozens of handlers and builds objects at import time is more expensive than a larger one that only defines functions. Attack initialisation work, not just the file size in a bundle report. **Move code behind the need.** Code for a feature that a minority of visitors use should be fetched when they use it. This converts a guaranteed cost for everybody into an occasional cost for the people who benefit. **Hold third-party tags.** Vendor code is the classic contributor to this burst because several tags each do their own initialisation and none of them are needed for the first screen. Loading them once the page is idle, or on first meaningful engagement, takes their work out of the sensitive window entirely. **Break up what remains.** Some initialisation genuinely has to happen at load. If it runs as one long stretch, split it so the browser gets control back between pieces and can respond to input in the gaps. This does not reduce total work, but it does reduce the longest uninterrupted block, which is what the user feels. ```javascript // Deferred but unconditional: everyone pays for the editor. import { mountEditor } from '/editor.js'; // Conditional: only users who open it pay. document.getElementById('open-editor') .addEventListener('click', async () => { const { mountEditor } = await import('/editor.js'); mountEditor(document.getElementById('editor-root')); }, { once: true }); ``` ## How to say this in an interview State the principle first — deferring relocates work, it does not remove it — then show that you know which metric each action moves. Candidates who answer this question by proposing a different loading attribute have not understood the symptom; the fact that the blocking number did not move is the evidence that the attribute layer has already been exhausted. The remaining levers are all about quantity: less code, later code, or no code. ## The lab-versus-field lesson underneath This scenario is also a good illustration of why a green local score is not a report card. A lab run on a fast machine compresses the burst enough that it may not read as a problem, and lab runs do not include a user clicking during it. The complaint came from real users because the pattern is only fully visible when real people interact with a page while it is still booting.
- The burst comes mostly from third-party tags rather than your own bundle. What changes about the fix?You lose the ability to make the code smaller, so the levers become fewer and later: run fewer tags, hold the survivors until the page is idle or the user engages, and put anything that does not need the page's DOM somewhere it cannot occupy the main thread. It also becomes a negotiation rather than a refactor, because each tag has an owner who wants it, so bring per-tag execution numbers to that conversation.
- Does making the bundle smaller always reduce Total Blocking Time?No. Blocking time tracks main-thread work, and bytes are only a proxy for it. Removing a large dependency that mostly defines functions may barely move the number, while removing a small one that walks the DOM and registers handlers at import time can move it a lot. Measure execution time per script and target that, rather than assuming the byte chart ranks your problems.
- If the work genuinely must happen at load, what do you do?Split it so the browser regains control between pieces, and order the pieces by what the user needs first — the interactive parts of the first screen before background bookkeeping. Total work stays the same, but the longest uninterrupted block shrinks, so an interaction arriving mid-startup waits for one short piece rather than the whole sequence.
Deferring is like pushing dishes off the counter into the sink: the washing still has to happen, it now happens all at once, and it happens exactly when someone wants to use the kitchen.
saying these in an interview costs you the question
- Believes deferring removes JavaScript cost
- Treats a green lab paint score as proof of speed
- Proposes switching defer to async as the fix
- Confuses a fast first paint with a responsive page
- Adds preload hints to fix a main-thread problem