A team splits a heavy click handler so it paints feedback and then yields before the expensive work, and local traces show much shorter tasks — yet real-user Interaction to Next Paint for that control has not improved, and a few interactions look worse than before. What would you investigate, and in what order?
answer
- the number did not move, so measure differently
- which phase grew, not which task shrank
- work relocated, not removed
- input delay is the next interaction's problem
- you may have fixed the wrong control
basics
~20 sUsually the deferred work did not disappear — it now runs while the user's next input arrives and shows up as input delay on the following interaction. Check the interaction's phase breakdown before concluding the split failed.
solid answer
~50 sFirst I'd look at the phase breakdown of the real interactions rather than the task lengths: input delay, processing time, and presentation delay. If input delay grew, the split relocated the cost — the deferred continuation is on the thread when the next click lands — and the fix is cancelling superseded work or slicing the continuation smaller. If presentation delay dominates, scheduling was never the lever; the frame itself is expensive and needs a smaller DOM update. If processing is still long, the yield may not be a real handback, so the paint never happened where the code implies. Only then would I question the target: the reported score reflects the worst interactions on a visit, so fixing one control moves nothing if a different control is worse. And a fast local machine flatters every split, so lab evidence alone was never going to settle it.
go deeper
Know that a change looking faster on your own machine is not evidence that users experienced it, and that responsiveness has to be checked with real-user data.
Be able to name the three phases of an interaction — input delay, processing, presentation — and say which phase splitting a handler is supposed to shorten.
Demonstrate the diagnosis path: read the phase breakdown, recognise relocated work appearing as input delay on the next interaction, confirm a paint really lands between the chunks, and know when the frame itself is the cost.
Own the verification discipline — define the field success criterion before the change ships, target the interactions that actually set the score, and stop the team from declaring victory on lab traces recorded on developer hardware.
## Start from the phases, not the tasks Shorter tasks in a trace is the wrong success criterion. An interaction is measured from input to the next paint, and that span decomposes into three phases: - **Input delay** — the event waited because the main thread was busy with something else. - **Processing** — the event's own handlers ran. - **Presentation delay** — style, layout and paint of the resulting frame. A split handler attacks the *processing* phase. If the number did not move, the honest first question is which phase actually dominates the slow interactions in the field. Real-user data can be collected per interaction with the phase timings attached, and that breakdown is what turns "it didn't work" into a specific diagnosis. ## Diagnosis 1 — the work was relocated, not removed This is the most common outcome and it explains the "a few got worse" part of the symptom. Before the split, one click produced one long handler. After the split, one click produces a short handler *plus* a continuation that runs a moment later. If the user is doing anything — clicking again, typing, scrolling — that continuation is now sitting in front of the *next* input, and it appears as input delay on that interaction. The accounting looks better on a per-handler basis and identical (or worse) to the user. The tells: - Input delay on interactions that follow the fixed one has risen. - The regression shows up mostly in rapid interaction sequences. Fixes: abort superseded continuations so only the latest survives; slice the continuation so the thread is reclaimable within a few milliseconds; or debounce, so a burst of clicks produces one piece of deferred work rather than five. ## Diagnosis 2 — the yield was not a real yield If processing time is still long in the field, the continuation may be resuming before the browser rendered. Not every way of "awaiting" hands the frame back; some resume in the same turn, so the code reads as split while the paint still sits behind everything. Confirm in a trace that a paint appears *between* the two script chunks. If it does not, the split is cosmetic. ## Diagnosis 3 — the frame itself is expensive If presentation delay dominates, no scheduling change will help. The acknowledgement is doing more than it looks: replacing a large subtree, toggling a class that invalidates a lot of layout, revealing a big list. The browser must still style, lay out and paint all of it before the interaction ends. The lever here is rendering less — narrower updates, fewer nodes touched — not moving script around. ## Diagnosis 4 — you fixed a control that was not the bottleneck The page-level score is driven by the worst interactions a visit contains, not the average and not the one you chose. If a different control — a filter that re-sorts ten thousand rows, a menu that mounts a heavy panel — is worse, improving the save button changes nothing visible in the aggregate even though the fix was real. Segment field data by interaction target before optimising, so the first question you answer is *which* interaction sets the number. ## Diagnosis 5 — lab evidence was never going to settle it The local trace was recorded on a developer machine: fast CPU, warm caches, no contention, few extensions, small data. Splitting a handler helps most exactly where it is easiest to measure and least where it matters. Real users bring slow devices, thermal throttling, third-party scripts competing for the thread, and datasets an order of magnitude larger. A green local trace and flat field data is not a contradiction — it is the normal case, and it is why the field number is the one that decides. ## The order I would work in 1. Pull the phase breakdown for the affected interactions from real-user data. 2. If input delay grew → the continuation is colliding with the next input; cancel, coalesce, or slice it. 3. If processing is unchanged → verify a paint really happens between the chunks. 4. If presentation dominates → stop scheduling and shrink the frame. 5. If the fixed interaction is not the worst one → re-target. ## The lesson to state out loud Yielding never removes work; it only changes *when* the work competes with the user. That is a genuine win when it moves work out of a frame the user is waiting for, and a wash — or a regression — when it moves work into the moment the user does something else. Verifying the change means looking at the metric it was supposed to move, in the environment it was supposed to move it, broken down by the phase the change was supposed to touch.
- Which single signal most cleanly distinguishes "relocated the work" from "the frame is expensive"?The phase that grew. If input delay on subsequent interactions rose, the deferred continuation is colliding with the next input — that is relocation. If presentation delay dominates while processing is already short, the frame itself is the cost and no amount of scheduling touches it. Task length in a trace does not distinguish the two.
- How would you verify the fix now, before shipping the next attempt?Decide the success criterion up front and make it a field one: the phase breakdown for that interaction target, at a high percentile, over a comparable population. Ship behind a flag if possible so the comparison is concurrent rather than week-over-week. A local trace is only useful for confirming a paint lands between the chunks — it cannot tell you the metric moved.
- The team asks whether they should just revert the split. What do you say?Not yet — the split is probably correct and incomplete. Reverting brings back a long handler in front of the feedback paint. The missing piece is cancellation and sizing of the continuation, so stale work doesn't queue behind the user. Revert only if the evidence says presentation delay dominated all along, in which case the split addressed the wrong phase entirely.
saying these in an interview costs you the question
- Treats shorter tasks in a trace as proof the metric improved
- Never checks which phase of the interaction is dominant
- Assumes deferred work is free because it left the handler
- Concludes the field data must be wrong when lab looks green
- Optimises a control that was never setting the page's score