skip to content

How would you summarize a 3-hour video into chapters when it exceeds the context window?

level: principalimportance: should knowfreq 40%

answer

  1. one impossible request becomes many affordable ones
  2. split, summarize, then summarize the summaries
  3. boundaries need overlap
  4. the second pass sees only text
  5. go back to the footage for the chosen moments

basics

~20 s

Process it hierarchically: split into overlapping windows, summarize each with timestamps, then summarize the summaries into chapters. The reduce pass sees only the intermediate text, so detail and cross-window continuity have to be deliberately carried through it.

solid answer

~50 s

The workable shape is map-reduce over time. Cut the three hours into windows of several minutes, overlapping by enough to keep an event that straddles a boundary intact, summarize each window into timestamped observations, and then run a reduce pass over those summaries to produce chapters. Two design decisions carry most of the quality. First, what the map pass emits: free prose loses the detail the reduce pass needs, so a fixed structure with timestamps, entities and salience scores travels much better. Second, whether chapters are final. For a highlights product they usually are not, so a third pass re-visits the raw footage for the selected intervals at a higher frame rate to produce the actual clips. The costs are error compounding, since the reduce pass cannot check anything the map pass got wrong, and boundary blindness for slow arcs that span many windows. Windows are parallelizable, so wall-clock latency is mostly the reduce pass.

go deeper

for a junior

Know that content too long for the context window is processed in chunks and then combined, and that each chunk's summary needs timestamps so the combined result can point back into the footage.

for a middle

Be able to describe the map-reduce shape concretely: window length, overlap, structured per-window output, and a reduce pass, and explain why the reducer cannot fix what the mapper got wrong.

for a senior

Show the coarse-then-fine discipline, with a refine pass at a higher frame rate over selected intervals, and give real cost and latency reasoning including the parallelism of the map stage.

for a principal

Own the architecture choice and its uncertainty: why hierarchical map-reduce over memory-style or compression alternatives, what evaluation would justify switching, and what grounding checks keep a fluent but fabricated chapter list from shipping.

## Why hierarchy rather than a bigger window Even with a million-token context, three hours of footage at any useful frame rate does not fit, and the fraction that does fit is expensive on every call. Hierarchical processing trades one impossible request for many affordable ones, and the structure it imposes turns out to be useful in its own right, because chapters are themselves a hierarchy. ## The map pass Split the footage into windows. Window length is a real parameter: too short and each summary lacks the context to know what matters, too long and you are back at the original problem. Several minutes is a common landing zone for continuous footage, and for structured content the natural units, a half of a match, a scene, an agenda item, beat the clock. Overlap the windows. An event that begins at 41:58 and ends at 42:06 is mangled by a hard cut at 42:00, and the reduce pass has no way to know it was mangled. Thirty seconds of overlap on each side costs a few per cent of the budget and removes an entire failure class, at the price of duplicate observations that the reduce pass must merge. What each window emits matters more than most teams expect. If the map pass returns a paragraph of prose, the reduce pass receives three hours compressed into fluent text that has already discarded the specifics. A structured emission survives much better: timestamped observations, the entities involved, a salience or confidence score, and a flag for anything that appears unresolved at the window boundary. The reduce pass can then rank and group rather than re-write. ## The reduce pass The reduce pass sees only text. This is the defining constraint and the source of the main failure mode: it cannot verify anything, cannot recover detail the map pass dropped, and will happily produce a confident chapter list built on a map-pass hallucination. Error compounds one way only. It is also where cross-window structure is created. Arcs that develop slowly, a mood shifting across an hour, a score changing across a match, a machine degrading across a shift, exist nowhere in any single window summary and must be assembled here. That only works if the map pass emitted comparable, structured signals rather than independent prose. When the summaries themselves do not fit, reduce in more than one level: summarize groups of windows, then summarize those. Each additional level is another lossy compression, so stop at the shallowest hierarchy that fits. ## The refine pass For most real products chapters are an index, not the deliverable. A highlights reel needs the actual moments, so a third pass takes the intervals the reduce pass chose and re-analyses only those at a much higher frame rate, where precise boundaries and fine detail matter and the token cost is affordable because the intervals are short. This coarse-then-fine structure is what makes the whole design economic: cheap sampling everywhere, expensive sampling only where the answer lives. ## Cost, latency and evaluation Cost is roughly the whole recording sampled once at the coarse rate, plus a small reduce, plus the refined intervals. That is dramatically less than a single dense pass and is the number to quote when someone asks why not just use a longer context. Latency is not the sum of the windows, because the map pass is embarrassingly parallel; wall-clock time is one window plus the reduce, subject to rate limits, which is usually the binding constraint in practice. Evaluation is the part teams skip. Chapter quality has no single metric, so pick two: agreement with human chapter boundaries within a tolerance, and a check that every claim in a chapter summary is grounded in a real timestamp. The second catches compounded map-pass errors, which are otherwise invisible because the output reads beautifully. ## The honest uncertainty There is no settled best architecture for long-video understanding as of 2026. Hierarchical map-reduce is the reliable, boring, well-understood option. Research alternatives, memory-style architectures that carry state forward across windows and token-compression schemes that squeeze many frames into few tokens, are actively moving and can beat it on benchmarks, but they are harder to reason about when they fail. A principal-level answer says which one it is choosing and why, names the failure it is most worried about, and describes what would change its mind.

  • Why does emitting structured output from each window beat emitting a prose summary?
    Because the reduce pass can only work with what it receives. Prose has already thrown away timestamps, entity identity and confidence, so the reducer cannot rank, deduplicate overlapping observations, or link an arc across windows. Structured emissions with times, entities and salience let the reducer do set operations rather than creative rewriting, which is both cheaper and far less prone to inventing continuity that was not there.
  • What happens to an arc that develops slowly across the whole three hours?
    No single window sees it, so it exists only if the reduce pass can assemble it from comparable signals across summaries. That is an argument for emitting consistent structured fields rather than free text, and sometimes for a dedicated pass that looks only at one tracked quantity across all windows. If neither is in the design, expect the system to report every local event correctly and miss the story entirely.
  • How would you decide between this and simply waiting for longer context windows?
    Cost and control, not capacity. Even when three hours fits, a single dense pass pays full price for every idle minute on every query, and you cannot re-analyse a chosen moment at higher fidelity without redoing everything. Hierarchy gives you a cheap index plus expensive detail on demand. Longer windows raise the ceiling on window size, which simplifies the hierarchy rather than removing the reason for it.

saying these in an interview costs you the question

  • Cuts windows with no overlap, splitting boundary events
  • Assumes the reduce pass can catch map-pass errors
  • Emits free prose per window and expects detail to survive
  • Treats chapters as the deliverable with no refine pass
  • Adds hierarchy levels without accounting for compounding loss

context