A stakeholder says a page in your web app "feels slow". Walk through the profiling workflow you would follow in the browser, from that complaint to a verified fix.
answer
- start before you touch code
- one scenario, one number
- dominant cost, not familiar cost
- one change per measurement
- re-record the same scenario
basics
~20 sTurn the complaint into one reproducible scenario and one number, record a browser performance trace under representative conditions, attack the single dominant cost the trace shows, then re-record the same scenario and compare. One change per measurement.
solid answer
~50 sFirst I make the complaint concrete: which page, which moment, on what device and network, and what "slow" means as a number — a load that takes six seconds, or a filter click that takes 800 ms to repaint. Then I reproduce it against a **production build** with realistic data, CPU and network throttling, and a cold cache, because an unthrottled dev build hides the very cost I am hunting. I record a trace of exactly that scenario and look for the *dominant* cost first: is the main thread solidly busy, or idle waiting on the network? I turn that into one falsifiable hypothesis, make one change, and re-record the same scenario under the same conditions. If the number did not move, I revert — an unverified optimization is pure risk. Then I repeat, because the second-largest cost is usually only visible once the first one is gone.
go deeper
Be ready to say you would reproduce the slowness and record a trace before touching code, and to name the conditions you would set: production build, throttling, cold cache.
Explain how you read a trace to find the dominant cost rather than the familiar one, and why one change per measurement is what makes any result attributable.
Show production judgment: how you choose which cost to attack given effort and risk, how you handle a symptom that only reproduces on a real device, and the honesty to say when a number moved but the experience did not.
Own the meta-decision: when a profiling investigation is worth the engineering time at all, how you size expected gain against cost, and how you stop an investigation from quietly becoming a rewrite.
## Why the workflow matters more than the tricks Performance work fails in a predictable way. Someone hears "the app is slow", remembers an optimization they read about, applies it, and ships. Sometimes the page gets faster; nobody can say why, and the next complaint restarts the cycle. A profiling workflow replaces guessing with a loop that produces evidence at every step: reproduce, quantify, record, isolate the dominant cost, change one thing, re-measure. Interviewers ask for this walkthrough because it separates people who have actually made a page faster from people who have collected tips. ## Step 1 — Turn the complaint into a scenario and a number "Feels slow" is a symptom report, not a bug. Extract four things from whoever reported it: - **Which page or flow**, exactly, and with what data (an empty account and a ten-thousand-row account are different products). - **Which moment feels slow** — the page taking too long to appear at all, or the page appearing and then not responding to a click. Load problems and interaction problems have entirely different causes and entirely different fixes; conflating them is the single most common way an investigation goes nowhere. - **On what device and network**, and whether it happens on the first visit or every visit. "Only the first time" points at download and parse cost; "every time" points at work the page redoes. - **What the target is.** Write it down: "the product list should show its content in under 2.5 s on a mid-range phone over a slow connection". You cannot verify a fix against a feeling. ## Step 2 — Make the environment representative A trace only tells you the truth about the machine and build you recorded. Before recording: - Profile a **production build**. Development builds carry unminified code, development-only warnings, and a hot-reload client — costs no user ever pays, and they can easily dominate the trace. - Use **realistic data volumes**, not a seeded fixture with five rows. - Apply **CPU and network throttling** and start from a **cold cache**, so the recording resembles the device and connection the complaint came from. - Record in a **clean browser profile** with extensions disabled; extensions inject scripts and their cost lands in your trace. ## Step 3 — Record the scenario, and read the trace largest-first Record only the scenario in question, and keep it short — a sixty-second recording is hard to read and the profiler's own overhead grows with it. Then read the trace from the top down, not from the first thing you recognize: 1. Over the span you care about, where did the time actually go? 2. Was the main thread solidly busy, or idle waiting on bytes? 3. What is the single largest contributor, and roughly what share of the total is it? That last question is the whole game. If a page takes four seconds and you find a 40 ms function you know how to make ten times faster, fixing it perfectly buys you 36 ms — invisible. The dominant cost sets the ceiling on any improvement, so it is where the effort goes even when it is the less interesting problem. ## Step 4 — One hypothesis, one change State the hypothesis so it can be wrong: "the hero image is discovered late because it is referenced from a stylesheet, which is why the largest paint lands at 3.4 s", or "parsing the 900 KB analytics bundle is the 700 ms block before first paint". Then make **one** change. Batching five changes and observing a faster page tells you the batch helped; it does not tell you which part helped, which part did nothing, and which part quietly made things worse. ## Step 5 — Re-record the same scenario Same build type, same throttling, same cache state, same data. Compare the same span you measured before. Two outcomes are useful: the number moved and the trace shows the specific cost you targeted shrinking, or it did not move and you revert the change. A change that costs complexity and buys nothing measurable is a net loss, and reverting it is a result, not a failure. ## Step 6 — Repeat, and know when to stop Removing the biggest cost reshapes the trace: work that was hidden behind it becomes the new bottleneck, and it is frequently something you would never have guessed from the first recording. So loop. Stop when you have hit the target you wrote down in step 1, or when the next fix costs more engineering and risk than the milliseconds are worth — that judgment is part of the job, not a cop-out. ## Failure modes worth naming out loud Optimizing before measuring; measuring an environment no user has; fixing the *recognizable* cost instead of the *dominant* one; changing several things at once; declaring victory from a single lucky run; and improving a number while the moment the user actually cares about lands at exactly the same time as before.
- The trace shows five separate costs that each look worth fixing. How do you order them?By share of the number I am trying to move, adjusted for effort and risk. A cost that is 60% of the span sets the ceiling on any improvement, so it goes first even if it is the dull one. Among similar-sized costs I take the one with the cheapest, most reversible fix, and I re-record after each so the ordering stays honest as the trace reshapes.
- Why insist on profiling a production build rather than the dev server?Development builds ship unminified code, dev-only warnings and assertions, and a hot-reload client, and they often skip minification and bundling entirely. Those costs can dominate a trace and none of them reach users, so the profile tells you about your tooling rather than your product. The numbers are also not comparable to anything you will measure after deploying.
- What do you do when the slowness only reproduces for the reporter and not on your machine?Treat the difference as the clue. Narrow the variables one at a time — device class, network, account data size, extensions, locale — and try to reproduce with those conditions applied. If it still will not reproduce locally, get a trace from the actual device through remote debugging, or add timing instrumentation around the suspect span and ship it to gather evidence before guessing.
saying these in an interview costs you the question
- Starts changing code before recording any trace
- Fixes the first familiar thing, not the biggest cost
- Profiles the dev server build and quotes those numbers
- Changes several things at once, then cannot attribute the win
- Never re-records after applying the fix