Why is a processor's generated output fed through another processing round, and what makes the rounds stop?
answer
- output becomes input
- a fixed point, not a pass
- monotonic: only adds
- stops when a round emits nothing
- a final pass for diagnostics only
basics
~20 sGenerated files are ordinary source: they can carry markers of their own, so the build re-runs processing over them. Each round's output becomes the next round's input, and the loop ends when a round emits no new file.
solid answer
~40 sA processor emits source, and source can itself carry markers — a generated companion mapper may be marked for injection, validation or serialization by some other tool. If the build processed only hand-written files, that second tool would silently never run. So processing is a **loop**: round one sees the hand-written declarations, round two sees what round one emitted, and so on. The loop ends when a round produces no new file, after which the build usually runs one **final pass** in which processors can report diagnostics but generation is no longer useful, because nothing would process the result. Termination is therefore a property of what the processors emit, not a limit the build imposes: a processor that emits a file carrying its own trigger marker on every round will not terminate.
code
pseudocode · 12 linesinputs = hand_written_sources
loop:
matched = declarations in inputs carrying a supported marker
generated = run_processors(matched) // may emit new source files
if generated is empty:
break // a round emitted no new file
inputs = generated // this round's output is next round's input
run_processors(nothing_new) // final pass: diagnostics, not generationgo deeper
Recall that generated source is ordinary source: it can carry markers, and the build looks at it again. Generation is not a single pass that happens once before compiling.
Explain the loop with its inputs named: hand-written sources, then each round's emitted files, ending when a round emits nothing, plus a final pass for diagnostics. Say why the loop only ever adds.
Diagnose the two real failures: a build that spins because a processor emits its own trigger marker, and a duplicate-output error from a processor written as if called once. Explain how you make emission deterministic.
Consider what chained processors cost across many teams' builds: hidden depth, build time that no source file explains, and errors that point at generated files. Decide what you require of a processor before it is allowed on the shared path.
## Why one pass is not enough A build-time processor emits **source**, and source is exactly the thing processors consume. Suppose one tool emits a companion mapper type for every marked declaration, and the emitted mapper is itself marked so that a second tool can register it. If the build ran processing once, over hand-written files only, the second tool would never see the mapper: the file exists, it compiles, and the registration silently never happens. That failure — a thing that exists but was never processed — is what the rounds model is designed to prevent. So processing is a fixed-point loop rather than a pass. ## The loop, step by step 1. **Round one.** The build gathers the hand-written declarations, dispatches the marked ones to the processors that advertised those markers, and collects the files they emit. 2. **Round two.** The inputs are the files emitted in round one. Their declarations are parsed, their markers dispatched, and whatever those processors emit is collected. 3. **Round n.** The same, over round n-1's output. 4. **Termination.** A round that emits no new file ends the loop. 5. **The final pass.** Most models then call the processors once more with nothing new, to let them report accumulated diagnostics or emit a summary — but a file created at this point has no round left to process it, so models either warn about it or reject it. ## What is and is not true of a round | true of a round | not true of a round | |---|---| | it is a phase inside one compilation | it is a separate invocation of the compiler | | its input is the previous round's output | it re-scans the hand-written files from scratch | | it may dispatch to several processors | each processor gets its own private round | | it can add files | it can amend a file an earlier round wrote | The last row matters most. The loop is **monotonic**: each round adds to the set of declarations and nothing is ever taken back. That is what makes the model tractable — a declaration a processor saw in round one is still there, unchanged, in round four — and it is the reason generation is an add-only operation. ## Termination is the processors' property, not the build's The honest statement is that the loop ends when a round emits nothing new, and that this is a property of what the processors do. Two failure shapes follow: - **Runaway generation.** A processor that emits a file carrying the very marker that triggers it regenerates on every round. Because each round's output is a *new* file name, the loop keeps finding work. The build does not rescue you here by capping rounds in the general model; the symptom is a build that spins, with the generated directory filling up. - **The one-round assumption.** A processor written as if it will be called exactly once — accumulating state in a field and writing it on the first call — misbehaves once its own output re-enters. It may try to write the same output path twice, which processing models reject rather than silently overwriting, and that surfaces as a confusing mid-build error about a duplicate file. The discipline that avoids both: make emission a pure function of the declaration it is emitted for, name the output deterministically, and never emit a file carrying a marker your own processor advertises. ## What this buys, and what it costs What it buys is **composition**. Tools that know nothing about each other chain automatically: one emits, the next processes, a third processes that. No build configuration expresses the ordering, and none needs to, because the loop discovers it. What it costs: - build time grows with the depth of the chain, and a deep chain is invisible in the source; - an error in a late round points at a generated file, so the author must trace back to the marked declaration that caused it; - a processor cannot ask about a declaration that a later round will generate, because it does not exist yet; code that needs the whole picture must therefore defer its work to a later round or collect across rounds and emit once, at the end. That last constraint is the practical one. "Generate the registry of everything marked" is a legitimate requirement and it cannot be satisfied in round one, because round one has not seen what rounds two and three will produce. The idiom is to accumulate across rounds and emit the registry in the last round that still has a round after it. ## The interview form The question is usually asked as *"a file was generated but the second tool did not run on it — why, and when would it?"*. Naming the loop, naming what each round's input is, and naming the emit-nothing termination condition is the full answer.
- A processor must emit one registry listing every marked declaration in the build. Why can it not do that in the first round?Because round one has only seen the hand-written declarations; later rounds may add marked declarations that belong in the registry. The idiom is to accumulate across rounds and emit once, in the last round where the emitted file still has a round left to be processed — emitting it in the final pass is too late for anything to consume it.
- What happens if a processor tries to write the same output path in two different rounds?Processing models reject the second write rather than overwriting the first, so the build fails with a duplicate-output error. It usually means the processor kept per-build state as if it would be called once, or its file naming is not a deterministic function of the declaration it is generating for.
- Is a round the same thing as compiling twice?No. Rounds are phases inside a single compilation: the declarations of earlier rounds stay loaded and resolved, and the compiler emits compiled output once, at the end, for everything. Compiling twice would restart resolution and lose the accumulated state that lets a processor reason across rounds.
saying these in an interview costs you the question
- Thinks generated files are never themselves offered to processors
- Believes the build caps rounds, so runaway generation is impossible
- Says a later round can rewrite a file an earlier round emitted
- Thinks each round starts over from the hand-written sources
- Confuses a round with a second invocation of the compiler
- Expects a first-round processor to see declarations generated later