Core Web Vitals replaced First Input Delay (FID) with Interaction to Next Paint (INP) in 2024. What did FID measure, and why could a page with a good FID score still have poor INP?
answer
- FID watched the first input only
- delay before the handler, nothing after
- almost every site passed it
- INP is end-to-end, every interaction
- 100/300 became 200/500
basics
~20 sFID timed only the queueing delay before the first interaction's handler started. INP measures every qualifying interaction all the way to the next paint, so slow handlers, heavy rendering, and interactions after the first one — all invisible to FID — now count.
solid answer
~50 sFID measured one narrow thing: for the **first** interaction on a page, how long the browser waited before the first event handler could start. It stopped there — it never measured the handler itself, never measured the rendering that followed, and never looked at any later interaction. That made it easy to score well on: a page could pass FID while every click after the first one took half a second to visibly respond. INP closes all three gaps. It considers **all** qualifying interactions in the visit and reports essentially the worst one, and it measures each interaction end to end — input delay, handler processing, and the presentation delay before the next frame. So a single-page app whose first click is trivial but whose filter controls re-render a huge list will look fine under FID and poor under INP. Since 2024 the bands are 200 ms good, 500 ms poor, at the 75th percentile.
go deeper
Remember that FID looked only at the delay before the first interaction's handler started, while INP measures every qualifying interaction all the way to the next paint.
Explain the three blind spots FID had — handler cost, rendering cost, and later interactions — and why the thresholds moved to 200 ms and 500 ms when the measured window widened.
Show what the swap changed operationally: FID-era tactics only address the input-delay phase, so responsiveness work now has to follow the user through the whole session, especially in single-page apps.
Own the reporting story — a metric almost every origin passed had no power to direct investment, and replacing it moved responsibility for responsiveness from a load-time checklist into ongoing feature work.
## What FID actually measured First Input Delay took the first click, tap or key press on a page and measured the gap between the input arriving and the first corresponding event handler beginning to execute. Nothing else. Concretely, it captured only the first of the three phases INP now reports, and only once per page visit. The reasoning at the time was defensible. FID was designed as a load-responsiveness signal: it answered "when the page first looked ready, was the main thread actually free?" That is a real failure mode — a page that paints quickly and then blocks on hydration or startup scripts feels broken when the user reaches for it. ## The three blind spots **1. It ignored handler cost.** The clock stopped the instant your listener started. A handler that then spent 400 ms sorting an array contributed nothing to the score. **2. It ignored rendering.** Everything after the handler — style recalculation, layout, paint, and the wait for the next frame — was outside the measurement, even though that is the part the user is staring at. **3. It ignored everything after the first interaction.** A visit could contain two hundred interactions; FID reported one. Single-page apps are the worst case here: the first click is often a benign navigation or a cookie-banner dismissal, while the genuinely expensive interactions — opening a filter panel, typing in a search box, expanding a table row — happen minutes later and were never sampled. The practical consequence was well documented before the change: the overwhelming majority of origins passed FID, which meant the metric had almost no diagnostic power. A metric nearly everyone passes cannot tell you where to invest. ## What INP changed INP fixes each blind spot directly: - **End-to-end timing.** Each interaction is measured from the input to the **next paint** after the handlers finish, so all three phases count. - **All interactions, not the first.** Every qualifying click, tap and key press in the visit is a candidate. - **Worst-case reporting.** The reported value is drawn from the slowest interactions rather than an average, because users remember the click that hung, not the mean. The thresholds also moved to match the wider measurement: FID used 100 ms and 300 ms, INP uses **200 ms (good)** and **500 ms (poor)**, both evaluated at the 75th percentile of visits. The INP numbers are larger because the window is larger — it is not a loosening of the bar. ## Why the swap changed team behaviour Under FID, the winning tactic was to keep the main thread free around first input — defer scripts, break up hydration — and that was the end of the responsiveness conversation. Under INP those tactics still help the input-delay phase, but they are no longer sufficient. Teams now have to care about what handlers do and about how much of the page a handler asks the browser to re-render, for the whole life of the session. For long-lived single-page applications that was a substantial shift: the metric now follows the user through the app instead of stopping at the front door. ```js // FID's territory was only this gap: // entry.processingStart - entry.startTime // INP measures the whole span: // entry.duration (input -> next paint after handling) ``` ## What carried over The input-delay phase of INP is conceptually the thing FID measured, so old FID work was not wasted — long tasks that blocked first input still show up, now as one contributor among three. Field data collection also carried over: both are field metrics, gathered from real users, not something a synthetic run can produce faithfully, because a lab run rarely reproduces the interactions a real user performs. ## Answering the question crisply A good short answer names the three widenings — first interaction to all interactions, delay-only to end-to-end, and the threshold change that follows from measuring more — and gives one concrete example, such as a search field whose keystroke handler re-renders a thousand rows: invisible to FID, and a 600 ms INP.
- Given how narrow FID was, why was it introduced in the first place?It was a load-responsiveness signal, cheap to collect and hard to game: it asked whether the main thread was genuinely free when the page first looked usable. That is a real failure mode. Its weakness was that nearly every origin passed it, so it stopped distinguishing good pages from bad ones.
- Why did the good threshold rise from 100 ms under FID to 200 ms under INP?Because the measured window grew. FID timed only the queueing delay, while INP adds handler execution and all the rendering up to the next paint. Keeping 100 ms would have been a much harsher bar measuring three phases instead of one, so the bands were recalibrated to 200 ms good and 500 ms poor.
- Which kind of application was most exposed by the switch from FID to INP?Long-lived single-page apps. Their first interaction is often trivial while the expensive ones — filtering, searching, expanding rows — happen deep in the session and re-render large regions. FID never sampled those; INP reports essentially the worst of them, so scores that had been comfortably green often turned amber or red.
- Does work done to improve FID still help INP?Yes, partly. Keeping the main thread free — deferring third parties, breaking up startup and hydration tasks — reduces the input-delay phase, which INP still counts. It is simply no longer enough on its own, because handler cost and rendering cost now sit inside the same measurement.
saying these in an interview costs you the question
- Says INP is just FID with a bigger threshold
- Claims FID measured how long the handler ran
- Thinks FID covered every interaction on the page
- States that INP is a lab metric measurable in a synthetic run
- Believes improving first-input delay alone guarantees good INP