A job's only input is one 40 GB file compressed as a single stream — how many worker threads can read it?
answer
- the stored form decides, not the cluster
- can a reader start mid-file?
- entry points: block headers or sync markers
- one continuous stream means one piece
- redistribution afterwards restores width
basics
~20 sOne. A stored form that cannot be decoded from an arbitrary offset yields exactly one piece of the input — one slice a single worker thread reads end to end — so every other thread in the cluster stays idle during that read.
solid answer
~50 sA finite job derives its pieces from the stored bytes, and a piece of the input is one slice a single worker thread reads from start to finish. To make two pieces out of one file, a reader has to be able to begin decoding at some offset and find a record boundary from there. A file compressed as one continuous stream offers neither: the decoder's state at any byte depends on every byte before it. So the file is one piece, one thread works and the rest idle, and the wall clock is the single-thread decode rate. Note this binds the read pass and the steps fused to it, not the whole job — once records are redistributed between workers, later steps run at the new division's count. The durable fixes are a block-framed stored form, or many moderate files instead of one.
go deeper
Recall the rule and the reason: one piece of the input is worked by one worker thread, and a file that must be decoded from byte zero is one piece however large it is or however big the cluster is.
Explain the mechanics: dividing a file needs an offset a decoder can start at plus a findable record boundary, which block headers or sync markers supply and a single compression stream does not.
Show the production judgment: recognise the symptom (one lane busy, the rest idle), decide whether to redistribute straight after the read or to fix the stored form, and price the extra pass either way.
Frame it as a storage standard: what file size and framing every producer into a shared dataset must use, and what it costs to migrate existing data against the parallelism every downstream reader would gain.
## What "a piece of the input" means Before a **finite job** — a run over an input that ends, so the engine can measure the stored bytes before it starts — executes anything, the input is cut into **pieces of the input**: one slice of the stored input that a single **worker thread** reads and processes from start to finish. A worker thread is one lane inside a **worker process**, which is one operating-system process on one machine holding its own memory and several such lanes. The **piece count** is how many slices the run is divided into, and it is the ceiling on how many lanes can be busy. The machine count only decides how many lanes exist. That cut is not a free choice. It is derived from what is on disk, and the first thing the reader must establish about each file is whether it can be entered anywhere other than at the beginning. ## Why a stored form can or cannot be divided To make two pieces out of one file, a reader must be able to (1) begin decoding at some offset and (2) find the first whole record boundary at or after it. Forms that allow this build in entry points: - **Block framing** — the file is a sequence of independently decodable blocks, each with its own header and its own compressed payload, so a reader that lands mid-file scans forward to the next header and starts there. - **Sync markers** — a short, statistically unlikely byte pattern written between record groups, which a reader can scan for from any offset. - **Fixed-width uncompressed records**, where the boundary is arithmetic. A file compressed as **one continuous stream** has none of these. The decoder's state at byte N — its dictionary of earlier byte sequences — depends on every byte before N, so there is no offset a second lane could start from. The same holds for a file encrypted as one stream, for a single serialized document (one enormous array or one document tree) whose record boundaries cannot be located without parsing from the root, and for a text form whose records span lines with no framing to tell a mid-file reader whether it landed inside a record or between two. | stored form | pieces from one file | what decides it | |---|---|---| | whole-file stream compression | 1 | no offset the decoder can start at | | block framing, compressed per block | many | block headers are entry points | | records separated by sync markers | many | scan forward from any offset | | one serialized document | 1 | boundaries exist only relative to the root | ## What the single piece costs - **Concurrency** — one lane works and the rest idle; wall clock for the read is the single-lane decode rate, not the cluster's aggregate rate. - **Retry granularity** — the unit that fails is the whole piece, so a failure at 39 GB re-reads from zero. - **Memory** — whatever that pass accumulates accumulates inside one process instead of being spread across many. - **Machine time billed** — you pay for lanes that are doing nothing. What it does *not* cost is the whole job. The single piece constrains the read pass and the **narrow steps** fused into it — steps each worker finishes alone out of the piece already in its hands. As soon as the job performs a **wide step** — one a worker cannot finish from the records it already holds, because it needs records currently sitting on every other worker — the records are redistributed, and the steps after it run at whatever piece count that redistribution creates. A common deliberate shape is therefore: read the undividable file on one lane, redistribute immediately, then run everything expensive at full width. That pays one pass over the data to buy the parallelism back. ## What varies between engines - Deriving the pieces from stored bytes is a **finite job's** behaviour. In a **continuous job** — a run over an input with no end, where nothing about the input can be measured in advance — the author states a **declared operator width** instead, and it stands until the job is restarted. An undividable file read by a continuous job is read by one operator instance, and no width declaration changes that. - Some readers **pack several whole files into one piece** up to a byte target. That helps a population of small files and does nothing here, because one large undividable file is already one whole file. - Some engines plan one piece per file up front; others plan from measured sizes and discover at run time that only one piece produced records. The observable symptom — one lane busy, the rest idle — is the same either way. ## What to do about it 1. Store the data in a **block-framed** form. This is the only change that makes the file itself divisible. 2. Failing that, write **many moderate files** rather than one, so the piece count comes from the file count instead. What one job leaves behind is the layout the next job inherits. 3. Failing both, accept the single-lane read, redistribute immediately afterwards, and shape that first worker process for one heavy lane rather than for many idle ones.
- The same 40 GB arrives as 400 files of 100 MB, each compressed the same way. What changes?Each file is still one piece of the input, but now there are 400 of them, so up to 400 worker threads read at once. Divisibility inside a file never appeared; the piece count came from the file count instead. This is why writers choose a moderate file size rather than one giant object.
- Does adding worker processes shorten a job whose first step is this single undividable read?Not for that read. The piece count is one, so exactly one lane can be busy no matter how many exist. Extra capacity only helps the steps after a redistribution, which rebuilds the division at its own count. Until then you are paying for idle lanes.
- The reader reports a piece count of 320 for this file but only one lane produces records. What happened?The planner sized pieces from the file's byte length before checking whether the form admits an entry point, so the count is arithmetic rather than achievable. At run time each candidate piece except the first finds no boundary it can start from and finishes empty. Treat the reported count as a plan, not as evidence of parallelism.
A reel of film spliced end to end: to watch the scene an hour in, you still have to wind through everything before it, so a second projectionist cannot help. A book with numbered chapters is shared out immediately, because anyone can open it at chapter twelve. Divisibility comes from the entry points, not from the size of the thing or the number of readers you have.
saying these in an interview costs you the question
- Believes a big file is automatically divided because it is big
- Says more machines will speed up a single undividable read
- Thinks each machine can open its own handle and take a share
- Confuses compressing the file with being unable to read it
- Assumes the whole job is stuck at one lane after the read
- Treats a reported piece count as proof of actual parallelism