skip to content

A continuous job runs as repeated small finite jobs and writes changed state entries to a versioned file set each cycle - what dominates its state cost?

level: seniorimportance: nice to knowfreq 32%

answer

  1. paid per cycle, not per access
  2. partitions times cycles sets the file count
  3. fixed cost per file dominates small writes
  4. full copies bound the chain to load
  5. shorter cycle multiplies the overhead

basics

~10 s

Cost is paid per cycle, not per access: each cycle writes a file per state partition to durable shared storage, plus periodic full copies, so partitions multiplied by cycle rate sets the bill.

solid answer

~50 s

In this model the runtime cuts an endless input into short finite chunks and runs a complete job over each one. State is not a long-lived worker-local store; each cycle writes the entries it changed into a new versioned file for each state partition on durable shared storage, and periodically rewrites a full copy so the chain of changes does not grow without end. The consequences follow from that shape: cost scales with **partitions times cycles**, so many partitions and a short cycle produce a flood of tiny files and metadata operations whose fixed per-file cost can exceed the data itself. Reading state at the start of a cycle may mean applying the changes written since the last full copy, so infrequent full copies make loading slower. And an entry read a thousand times within one cycle costs no more than one read - the opposite of a per-access store.

go deeper

for a junior

Recall that in this model state is written out at the end of each cycle as files on shared storage, rather than kept on one worker for the life of the job.

for a middle

Explain why cost scales with partitions multiplied by cycles rather than with the number of state accesses, and what periodic full copies are for.

for a senior

Show the trade you would run: cycle length against state overhead and load time, with the file count computed rather than guessed, and say what you would change first when the overhead dominates.

for a principal

Decide the defaults nobody revisits: the partition count new jobs start with, the cycle length the platform recommends, and who pays when a thousand short-cycle jobs each create their own flood of tiny files.

## The model, stated by mechanism Some runtimes handle a continuous input by cutting it into short finite chunks and running a complete job over each chunk. In that model the retained set - everything the job remembers between records - is typically not a long-lived store sitting on a worker for the life of the job. Instead: - Each state partition has a versioned file set on **durable shared storage**: storage that outlives any worker and that every worker can read. - At the end of a cycle, the entries changed during that cycle are written as a new version. - Periodically the runtime writes a **full copy** of that partition's entries, so that loading does not require applying an unbounded chain of changes. - At the start of a cycle, the worker assigned a partition loads the last full copy plus the changes written after it, and works against that in memory for the duration of the cycle. The structural difference from a worker-local store is where the cost is paid. A worker-local store charges **per access**: an encode, a decode and possibly a disk read for each read and write. This model charges **per cycle**: an entry touched once and an entry touched a thousand times within the same cycle cost the same, and an entry untouched in a cycle costs nothing at all in that cycle. ## What actually dominates the bill | driver | why it costs | what makes it worse | |---|---|---| | files written per cycle | one or more per state partition per stateful step | many partitions, many stateful steps | | cycle rate | the per-cycle cost is paid again every cycle | a short cycle chosen for latency | | metadata operations on shared storage | creating, listing and committing files has a fixed cost per file | thousands of tiny files per minute | | periodic full copies | rewrites entries that did not change | a large retained set with few changes | | loading at cycle start | last full copy plus the changes after it | infrequent full copies, long chains | The entry that catches teams out is the third. A few kilobytes of changed entries still costs a file creation and a commit, and on shared storage those fixed costs do not shrink with the payload. Partitions multiplied by stateful steps multiplied by cycles per hour is the number to compute before anyone argues about bytes. ## Shortening the cycle Shortening the cycle is the obvious lever for latency, and it is the expensive one here: 1. **Per-cycle overhead is multiplied, not divided.** Halving the cycle doubles the number of file and metadata operations per hour while the changed bytes per hour stay roughly the same. 2. **Each file gets smaller.** The fixed cost per file becomes a larger share of the total, so the effective cost per useful byte rises. 3. **The chain since the last full copy grows faster in file count**, so either full copies must be written more often - more rewriting - or loading gets slower. 4. **Small files accumulate.** Old versions still have to be tracked and eventually removed, and the removal is itself work against shared storage. Lengthening the cycle inverts every line: cheaper state, higher latency, more work lost when a cycle fails and is redone. ## How to reason about it against the other placements - **Per-access cost is irrelevant here.** If your step reads the same entries many times per record, this model absorbs that almost for free, while a worker-local disk-backed store charges for each one. - **Per-cycle cost is irrelevant to the other placements.** A record-at-a-time runtime with a worker-local store has no cycle boundary to pay at; it writes a durable copy periodically instead. - **Memory still matters.** The working entries are held in the worker during the cycle, so a partition whose working entries do not fit is a problem here too. The model changes where durability and capacity come from, not the fact that a cycle's work happens in memory. - **Partition count is nearly irreversible.** The number of state partitions is fixed early, and it is simultaneously the parallelism ceiling and the multiplier on file count, so it deserves a decision rather than a default. ## What varies between engines This is one model of three, not the model. Some engines in this class hold a long-lived worker-local store and never write per-cycle state files; some offer a choice; the two-phase disk-to-disk batch model keeps no state between runs and none of this applies to it. A claim like *state is written to shared storage every cycle* is precisely right for one design and simply false for another, so name the model before you make the claim - in an interview that framing is most of the answer.

  • Why does the number of state partitions matter as much as the retained size?
    Because files per cycle scale with partitions multiplied by stateful steps, independently of how many bytes changed. Two hundred partitions on a ten-second cycle produce tens of thousands of small objects an hour, each paying a fixed creation and commit cost on shared storage that a few kilobytes of payload cannot amortise.
  • What does writing full copies less often save, and what does it cost?
    It saves rewriting entries that did not change, which is the bulk of the write volume when the retained set is large and churn is low. It costs load time: a worker starting a cycle, or a replacement worker after a failure, has to apply a longer chain of changes on top of an older full copy before it can process anything.
  • How does this compare with a worker-local disk-backed store on the same workload?
    They charge on different axes. The worker-local store charges per access, so it suffers when each record performs several state reads and writes, and is indifferent to cycle length. The per-cycle model charges at cycle boundaries, so it is indifferent to access count within a cycle and suffers when the cycle is short and partitions are many.

saying these in an interview costs you the question

  • Assumes per-cycle state cost behaves like per-access state cost
  • Thinks a shorter cycle only affects latency and not state overhead
  • Believes every cycle rewrites the whole retained set to shared storage
  • Treats small file and metadata operations on shared storage as free
  • Assumes every engine in this class writes state files per cycle
  • Forgets that the cycle's working entries still occupy worker memory