You are designing a rollout where React's `<Profiler>` feeds render timings into your telemetry pipeline. What decisions would you settle before turning it on for real traffic, and how do you keep the instrumentation from distorting what it measures?
answer
- what decision does the number drive
- few boundaries, stable ids
- the callback runs on the hot path
- buffer and flush, never send per commit
- sample sessions, and set an end date
basics
~20 sSettle what decision the data will drive, then place a handful of Profilers at meaningful boundaries, buffer and aggregate in the callback instead of sending per commit, sample traffic rather than instrumenting everyone, and agree in advance when the instrumentation gets removed.
solid answer
~50 sStart from the decision: if no plausible number changes what the team does next, do not ship the instrumentation. Then keep the surface small — a few Profilers with stable, low-cardinality ids at boundaries that map to real product surfaces, not one per component. The callback runs after every commit of each profiled subtree, so it must be trivially cheap: push a compact record into a buffer, aggregate in memory, and flush on an idle callback or page hide, never a request per commit. Group by `commitTime` to reconstruct per-commit totals across boundaries, and record `phase` so mount cost, update cost and nested updates stay separable. Sample a slice of traffic rather than everyone, since the profiling build costs every user who receives it. Report distributions rather than averages, and set a date for turning it off.
code
jsx · 28 linesimport { Profiler } from 'react';
const buffer = new Map();
function record(id, phase, actualDuration, baseDuration) {
const key = id + ':' + phase;
const entry = buffer.get(key) ?? { count: 0, actual: 0, base: 0, max: 0 };
entry.count += 1;
entry.actual += actualDuration;
entry.base += baseDuration;
entry.max = Math.max(entry.max, actualDuration);
buffer.set(key, entry);
}
document.addEventListener('visibilitychange', () => {
if (document.visibilityState !== 'hidden' || buffer.size === 0) return;
const rows = [...buffer].map(([key, e]) => ({ key, ...e }));
buffer.clear();
navigator.sendBeacon('/telemetry/render', JSON.stringify(rows));
});
export function InstrumentedRoute({ name, children }) {
return (
<Profiler id={name} onRender={record}>
{children}
</Profiler>
);
}go deeper
Understand that the onRender callback runs very often, so anything slow inside it — logging, formatting, network calls — makes the page worse rather than measuring it.
Be able to describe a workable shape: a handful of named boundaries, an in-memory buffer keyed by id and phase, and one flush on idle or page hide instead of a send per commit.
Show operational judgment on sampling and segmentation: whole sessions rather than commits, percentiles rather than means, device class as a dimension, and grouping by commitTime so nested boundaries are not double-counted.
Own the whole tradeoff: what decision the data unlocks, who pays for the instrumented build, the observer effect on the numbers you will quote, metric cardinality and retention cost, and the exit condition that keeps a temporary experiment from becoming permanent overhead.
## Start with the decision, not the metric The first question is not where to put the boundaries; it is what changes if the number is bad. "If the dashboard subtree's update cost at the 95th percentile exceeds 50 ms on mid-range Android devices, we prioritise splitting it this quarter" is a decision. "We would like render timings on a dashboard" is not, and it will not survive the first argument about the cost of the profiling build. Render telemetry is expensive in a way most telemetry is not, because collecting it requires shipping an instrumented React to the people you are measuring. ## Where the boundaries go A few Profilers beat many. Good boundaries are the ones you can name in a product conversation: a route, a data-heavy table, an editor pane, a feed. They give ids that stay meaningful in a dashboard six months later. Avoid instrumenting per component. It multiplies the measurement overhead, and it explodes the cardinality of your metric labels — thousands of ids that nobody queries, at real storage cost. Ids must be **stable and low-cardinality**: never bake a record id or a user identifier into them, both because cardinality is unbounded and because ids leave the browser in telemetry payloads. Boundaries can nest. An outer route Profiler plus one inner Profiler around the component you suspect gives you "the route cost 60 ms, of which the table was 48 ms" — which is exactly the shape of a useful conversation. ## Make the callback disappear into the noise `onRender` fires after every commit in which the profiled subtree rendered. On an interactive page that is a stream, not an event. Two rules follow: 1. **No I/O in the callback.** No `fetch`, no `sendBeacon` per commit, no JSON serialisation, no string formatting. Append a small record to a preallocated array or fold it straight into running aggregates. 2. **Flush elsewhere.** Send on an idle callback, on a timer, or on the `visibilitychange` transition to hidden — the moment where a beacon is most likely to actually be delivered. The reason is not only throughput. Your callback runs on the same path whose duration you are reporting, so expensive work inside it inflates the surrounding interaction and, indirectly, the very numbers you are collecting. ```jsx const buffer = new Map(); function handleRender(id, phase, actualDuration, baseDuration, startTime, commitTime) { const key = `${id}:${phase}`; const entry = buffer.get(key) ?? { count: 0, actual: 0, base: 0, max: 0 }; entry.count += 1; entry.actual += actualDuration; entry.base += baseDuration; entry.max = Math.max(entry.max, actualDuration); buffer.set(key, entry); } ``` ## What to keep from each call Keep `phase`, and keep it separate in the aggregate. Mount cost and update cost answer different questions, and folding them together produces a mean that describes neither. `'nested-update'` deserves its own counter: it means an update was scheduled from inside a lifecycle or effect during a commit, so one interaction produced two render passes — a structural finding, not a slow-render finding. Keep both durations. `actualDuration` is what ran; `baseDuration` is the no-bail-out estimate. Their ratio is the part that survives hardware differences, because a slow device inflates both. Use `commitTime` as the grouping key when several boundaries report for the same commit — it is identical across those callbacks, so it reconstructs a per-commit total instead of double-reporting nested subtrees as if they were independent. ## Sampling and rollout The profiling build costs every user who receives it, so the population you instrument is a budget decision. A single-digit percentage of sessions, or an internal/canary channel, answers most questions. Sample **by session**, not by commit: a session that reports only some of its commits gives you a distorted distribution, while a fully-instrumented minority of sessions gives you a clean one. Segment what you keep. Device class, connection type and route matter far more than a global number, because render cost is dominated by the slowest hardware in your user base. ## Reporting and honesty about the numbers Report percentiles, not means: render cost is long-tailed, and the mean hides exactly the sessions you care about. Compare against the same instrumented build over time, not against uninstrumented production — the instrumentation is part of what you measured. And state the observer effect explicitly when the numbers are quoted to anyone making a decision. ## An exit Agree up front on what ends the experiment: the question is answered, or a threshold has held for a defined period. Instrumentation without an owner and an end date becomes permanent overhead that nobody can justify removing because nobody remembers why it is there. If a metric earns permanence, that is a deliberate second decision with its own budget — not the default outcome of forgetting to turn the first one off.
- Why sample whole sessions rather than a fraction of commits?Because the interesting quantity is a distribution per interaction, and dropping random commits inside a session biases it — you lose exactly the bursts where several commits fired for one interaction. Fully instrumenting a small, randomly chosen set of sessions gives you an unbiased sample of real behaviour, and it also bounds who pays the profiling build's cost.
- How would you keep mount cost from polluting your update-cost metric?Aggregate by `phase` as part of the key. Mounts are inherently expensive because nothing can be skipped, so mixing them into an update metric raises the tail on every freshly-navigated route. Track 'mount', 'update' and 'nested-update' as separate series, and treat a rising nested-update count as a structural problem rather than a slow render.
- Several nested Profilers report for one commit. How do you avoid double-counting their durations?Group the callbacks by `commitTime`, which is identical for every Profiler in a commit, and treat the outermost boundary's actualDuration as the commit total for that region. Inner boundaries are a breakdown of the outer one, not additional cost — summing them produces a number larger than the work React actually did.
- What would make you argue against shipping this instrumentation at all?If the suspected problem reproduces locally, a devtools profiling session answers it at zero cost to users. If no threshold would change the roadmap, there is no decision to inform. And if the team cannot name an owner and an end date, the likely outcome is a permanently instrumented React that nobody dares remove — a cost with no corresponding benefit.
saying these in an interview costs you the question
- Wrapping every component in a Profiler for full coverage
- Sending an analytics request from inside onRender per commit
- Averaging mount and update timings into one metric
- Instrumenting all traffic instead of a sampled slice
- Putting record or user identifiers into the Profiler id
- Leaving the profiling build enabled with no end date