You are virtualizing a React chat log where each message's height depends on its content, so offsets cannot be computed from a single row height. How do virtualizers handle unknown row heights, and what goes wrong while the measurements settle?
answer
- you cannot divide by one row height
- start from a guess per item
- measure after mount, cache by item id
- prefix sums plus binary search
- correcting a row above you shifts the view
basics
~20 sThe virtualizer starts from an estimated height per row, renders the window, then measures each mounted row and caches its real height, rebuilding the offset table and total size. Until a row is measured its position is a guess, so correcting rows above the viewport shifts content unless the scroll is compensated.
solid answer
~50 sWith variable heights you cannot divide `scrollTop` by a row height, so the virtualizer keeps a per-item height cache and a prefix-sum offset table, and finds the first visible row by binary search. Items that have never rendered contribute an estimate — TanStack Virtual's `estimateSize`, react-window's `VariableSizeList` item-size function. When an item mounts it is measured, typically through a ref plus `ResizeObserver`, and the real height replaces the estimate, which changes the total size and every offset after it. That is where the pain lives: if a corrected item sits *above* the viewport, everything below it moves and the user sees a jump unless the virtualizer anchors and adjusts `scrollTop` by the delta. You also have to invalidate the cache when data or container width changes — react-window exposes `resetAfterIndex` for exactly that — and reserve space for images so late-loading media does not re-measure the list under the user.
code
javascript · 22 linesfunction buildOffsets(heights) {
const offsets = [0];
for (let i = 0; i < heights.length; i++) {
offsets.push(offsets[i] + heights[i]);
}
return offsets;
}
function indexAtOffset(offsets, scrollTop) {
let lo = 0;
let hi = offsets.length - 2;
while (lo < hi) {
const mid = (lo + hi + 1) >> 1;
if (offsets[mid] <= scrollTop) lo = mid;
else hi = mid - 1;
}
return lo;
}
const offsets = buildOffsets([40, 120, 60, 200, 80]);
console.log(offsets); // [0, 40, 160, 220, 420, 500]
console.log(indexAtOffset(offsets, 250)); // 3go deeper
Know that variable-height rows need an estimated size up front and a real measurement after the row renders, and that virtualizer libraries ask you for that estimate.
Explain the data structures — a per-item height cache, prefix-sum offsets, binary search for the first visible index — and why total list height changes as more rows get measured.
Demonstrate the scroll-anchoring judgment: which corrections are visible, why scrolling up is the bad direction, when to invalidate the cache, and how to reserve space for late-loading media.
Weigh whether variable heights are worth their complexity at all — clamping rows to a fixed height or moving overflow into a detail view removes an entire class of bugs, and that tradeoff is yours to call.
## Why fixed-height math stops working Uniform rows give you a closed form: index from `scrollTop / rowHeight`, offset from `index * rowHeight`, total from `count * rowHeight`. All three are O(1) and touch no DOM. Variable heights destroy every one of them. You cannot know a message's height until it has been laid out with its real text, at the real container width, in the real font. So the virtualizer maintains state it did not previously need: - a **height cache**, one entry per item, holding either the estimate or the measured value - a **prefix-sum offset table**, where entry `i` is the summed height of everything before item `i` - **binary search** over that table to answer "which item is at this scroll position" ```js function indexAtOffset(offsets, scrollTop) { let lo = 0, hi = offsets.length - 2; while (lo < hi) { const mid = (lo + hi + 1) >> 1; if (offsets[mid] <= scrollTop) lo = mid; else hi = mid - 1; } return lo; } ``` ## Estimate, then measure, then correct The standard algorithm runs three phases continuously: 1. **Estimate.** Every unmeasured item gets a provisional height from an estimator. The total size is the sum of estimates and measurements, so the scrollbar is approximately right from the first frame. 2. **Measure.** When an item mounts, the virtualizer reads its real box — a ref callback plus `ResizeObserver`, or an explicit measure call the library provides — and writes the value into the cache. 3. **Correct.** The offset table is rebuilt from the changed index onward, the spacer's total height changes, and the window is recomputed. The consequence is that total height is *unstable*: it converges as more of the list has been seen, and a badly wrong estimator makes the scrollbar thumb visibly grow or shrink while the user scrolls. A good estimator — the median of a sample of real rows, not a round number you liked — is the cheapest fix for most complaints about jumpy scrollbars. ## The scroll-anchoring problem This is the failure an interviewer is really probing for. Correcting an item *below* the viewport is invisible: content further down moves, but nothing on screen shifts. Correcting an item *above* the viewport moves everything below it, including the pixels the user is currently reading. Scrolling upward is therefore the bad direction: previously unmeasured items enter, get measured, and their delta pushes the content you are looking at. Virtualizers handle this by anchoring — remembering an item and its offset within the viewport, and after a measurement pass adjusting `scrollTop` by the accumulated delta so the anchor stays visually still. If you hand-roll a virtualizer, this is the part you will get wrong. If you use a library, this is the behaviour to test explicitly: scroll fast into the middle, then scroll up, and watch whether text stays put. ## Cache invalidation A measured height is valid only for the content and the width it was measured at. Four events invalidate it: - **Container width changes** (window resize, sidebar toggle). Every wrapped-text row can change height, so the whole cache is suspect. - **The item's data changes** — an edited message, an expanded "show more". - **Reordering, insertion or deletion.** Key the cache by a stable item id rather than by index, or heights follow positions instead of records. - **Late-arriving content**, most often images and embeds that load after first paint. react-window's `VariableSizeList` makes this explicit with `resetAfterIndex(index)`, which discards cached sizes from that index on. Headless virtualizers expose an equivalent remeasure call. Forgetting to invalidate produces the signature bug: rows overlapping or leaving gaps after a resize. ## Practical mitigations - **Reserve space for media.** Give images intrinsic width and height attributes or a CSS `aspect-ratio` so the row's height is known before the bytes arrive. This removes the largest source of post-measurement jumps. - **Measure at the right width.** Do not measure inside a container that is still animating open. - **Prefer content-derived estimates.** For chat, estimating from character count and known line height beats a constant by a wide margin. - **Consider making rows uniform.** Clamping to a fixed number of lines, or moving overflow behind a control that opens a detail panel, converts the whole problem back into the fixed-height case. That is a legitimate product answer and a strong one to offer. - **Accept that scroll-to-index is approximate.** Targeting an unmeasured item lands near it, not on it; libraries scroll, measure and correct, which can take a frame or two.
- Why does scrolling upward through a measured list feel worse than scrolling downward?Because upward scrolling brings unmeasured items into view above the content you are reading. Each measurement replaces an estimate with a real height, and that delta pushes everything below it — including the current viewport — unless the virtualizer anchors on a visible item and compensates `scrollTop` by the accumulated difference. Scrolling down corrects items already behind you, so nothing visible moves.
- The list is stable until images finish loading, then rows jump. What is the fix?Reserve the space before the bytes arrive: set intrinsic width and height attributes or a CSS aspect-ratio so the row lays out at its final height on first paint. Otherwise the row is measured short, then grows, invalidating every offset after it. As a backstop, remeasure the row on the image's load event so the cache does not stay wrong.
- Your height cache is an array indexed by row position. What breaks when a message is deleted from the middle?Every item after the deletion shifts down one index and inherits the previous occupant's cached height, so offsets are silently wrong until each row is remeasured. Key the cache by the record's stable id instead, so a deletion removes exactly one entry, leaves the rest correct, and remeasures only what actually changed.
saying these in an interview costs you the question
- Assumes every row can be measured before it is rendered
- Keys the height cache by index rather than item id
- Never invalidates measurements after a container resize
- Thinks a bigger overscan fixes measurement jumps
- Ignores scroll compensation when rows above the viewport change