How would you restructure a video-on-demand transcode pipeline where a two-hour upload takes twelve hours to produce six renditions?
answer
- serial work, summed time
- cut where decoding can restart
- one task per chunk per rung
- leases plus deterministic output keys
- cap retries, park the poison
basics
~20 sTreat transcoding as a queue of small independent tasks: split the source at keyframes into chunks, transcode every chunk of every rendition in parallel on a worker fleet, then assemble, package and publish, with idempotent per-task retries.
solid answer
~50 sTwelve hours for six two-hour renditions means they are encoded one after another at roughly real time. The fix is **chunked parallel transcoding** driven by a **job queue**. A coordinator probes the source, splits it at keyframes into chunks, say 60 seconds each (120 for two hours), and enqueues one task per chunk per rendition: 720 tasks. Stateless workers lease a task, encode the chunk with closed groups of pictures and forced keyframe timing, write the output under a **deterministic key**, and acknowledge; an expired lease returns the task to the queue, and the deterministic key makes a rerun idempotent. When a rendition's chunks are all done, an assembly step joins them, packages segments and writes manifests. Assuming each worker encodes a minute of video in about a minute, 200 workers finish the encode phase in about 3.6 minutes plus split and assembly time. Retry caps, a dead-letter state and priority lanes keep one bad file from stalling the fleet.
code
json · 12 lines{
"taskId": "tx-7f3a",
"assetId": "vod-81234",
"rendition": "1080p-5000k",
"chunkIndex": 57,
"sourceRange": { "startSec": 3420, "endSec": 3480 },
"state": "LEASED",
"attempt": 2,
"maxAttempts": 4,
"leaseExpiresAt": "2026-09-17T10:04:30Z",
"outputKey": "vod-81234/1080p-5000k/chunk-00057"
}go deeper
Recall that transcoding turns one upload into several renditions, and that splitting the work into many small independent tasks lets many machines run it at once.
Explain the job graph, probe, split at keyframes, fan out per chunk per rendition, assemble, package and publish, and how leases and deterministic output keys keep retries correct.
Show operational judgment: retry caps and dead-lettering for poison files, priority lanes, scaling on queue depth, straggler handling, and the quality artefacts chunking can introduce.
Weigh chunk size, fleet size and priorities against cost and time to publish, including when to ship a partial ladder and how much reclaimable capacity the retry model can absorb.
## Why the serial pipeline is slow A **transcode pipeline** turns one uploaded source into the renditions of an adaptive bitrate ladder. If a single process encodes each rendition from start to end, one after another, the wall-clock time is the sum of all encodes. Six renditions of a two-hour source at about real-time encoding speed is roughly 6 × 2 hours = 12 hours. A bigger machine helps a little, but it cannot keep up as uploads grow, and one crash near the end wastes hours of work. ## The pipeline as a job graph The standard redesign treats the work as a **directed graph of small tasks** fed through a **job queue**: 1. **Probe**: read the source's duration, frame rate, resolution and keyframe positions. 2. **Split**: cut the source at keyframes into chunks, for example about 60 seconds each, so each chunk decodes on its own. 3. **Fan out**: enqueue one task per chunk per rendition — 120 chunks × 6 renditions = 720 tasks. 4. **Transcode**: workers encode each chunk with a fixed keyframe interval and closed groups of pictures, so chunks join cleanly and segments align across renditions. 5. **Assemble and package**: when every chunk of a rendition is done, join the chunks, cut segments and write the rendition's manifest entries. 6. **Publish**: write the top-level manifest and mark the asset playable. Audio is commonly encoded once as a separate track rather than per chunk, which avoids small gaps or clicks at chunk joins. ## The queue and the workers - **Stateless workers** pull tasks, so the fleet scales horizontally with queue depth. - **Leases**: a worker holds a task for a limited time and extends it while working; if it dies, the lease expires and the task becomes available again. - **Idempotent output**: each task writes to a key derived from asset, rendition and chunk index, and commits it as a whole, so a rerun overwrites rather than duplicates. - **Completion records**: the job store records which tasks finished; assembly reads only recorded outputs, never partial files a crashed worker left behind. - **Retry caps and dead-lettering**: a source that crashes workers repeatedly is parked with its error after a few attempts instead of cycling forever. - **Priority lanes**: fresh uploads can jump ahead of bulk re-encodes of the back catalogue. - **Interruptible capacity**: because tasks are short and retryable, cheaper capacity that may be reclaimed at short notice becomes usable. ## The arithmetic All figures assume a worker encodes one minute of video in about one minute. | Plan | Work | Parallelism | Encode wall-clock | |---|---|---|---| | Serial renditions | 6 × 120 min | 1 | about 12 h | | One worker per rendition | 6 × 120 min | 6 | about 2 h | | Chunked fan-out | 720 × 1 min | 200 | about 3.6 min | 720 task-minutes spread across 200 workers is 3.6 minutes. Split, assembly, packaging and queue wait add to that, and the slowest chunk sets the finish time for its rendition. ## Quality and correctness pitfalls - **Rate control at boundaries**: an encoder working on one chunk cannot borrow bits from its neighbours, so quality can step at a join. Constant-quality modes or per-chunk multi-pass encoding reduce the effect. - **Keyframe alignment**: every rendition must force keyframes at the same timestamps, or segment boundaries drift between rungs. - **Chunk size**: very short chunks multiply queue and assembly overhead and give the encoder less context; very long chunks reduce parallelism and make retries expensive. - **Stragglers**: one slow worker holds up a whole rendition; re-issuing tasks that run far longer than their peers bounds that delay. ## Publishing strategy Waiting for the full ladder delays every upload by its slowest rung. Many pipelines publish once a **minimum playable set** exists — say the lower and middle rungs — and rewrite the manifest as higher rungs finish. The trade-off is extra manifest versions and a window in which top quality is missing, and players that loaded the early manifest may not see the new rungs until they reload it.
- How do you stop a retried transcode task from producing duplicated or corrupted output?Derive each task's output key from asset, rendition and chunk index, and commit the output as a whole, for example by writing under a temporary name and switching to the final key only when complete. A rerun then replaces the same key with equivalent output, so duplicates collapse. Assembly reads only chunks whose completion is recorded in the job store, never partial files a crashed worker left behind.
- Why might the pipeline publish lower renditions before the whole ladder finishes?Viewers can start watching as soon as a playable subset exists, which matters most for fresh uploads. The manifest first lists only the finished rungs and is rewritten as more complete. The cost is extra manifest versions and a window without top quality, and players that loaded the early manifest may not see new rungs until they reload it.
- What should the pipeline do with a source file that makes workers crash repeatedly?Cap attempts per task, then move the task and its job to a dead-letter state with the error captured instead of letting it cycle through the queue. Surface it to operators or the uploader, and keep a way to re-drive it after a fix. Without the cap, one poison file keeps killing workers and delays every other job.
saying these in an interview costs you the question
- A bigger single machine will keep a serial pipeline fast at any catalogue size.
- Chunks can be cut at arbitrary timestamps as long as they are equal length.
- Retries are safe without idempotent outputs because workers rarely crash.
- No rendition can be published until every rung has finished encoding.
- Chunked encoding has no quality cost at chunk boundaries.