skip to content

Should source generated from your telemetry description be committed to the repository, or produced by every build?

level: middleimportance: should knowfreq 50%

answer

  1. drift versus dependency
  2. committed output can go stale
  3. regeneration puts the generator in the build
  4. diff of consequences aids review
  5. regenerate, compare, fail on difference

basics

~20 s

Regenerating on every build keeps the description the only source of truth but makes the generator a build dependency. Committing the output removes that dependency and makes changes reviewable, at the price of drift nothing detects. A common middle path regenerates and fails the build on any difference.

solid answer

~50 s

The two options trade **drift** against **dependency**. Regenerating on every build means the emitted source cannot be stale: it is a function of the description, and no reviewer can approve a change to it. The price is that the generator now gates every build — it must be available, fast enough, and stable enough to produce the same output from the same input. Committing the output inverts that: anyone can build without the generator, the diff of a description change is visible in review, and stack traces point at files that are in the tree — but nothing forces a regeneration, so the committed files can quietly describe an older description. The middle path most teams land on is to regenerate in the build **and** keep the output committed, with a check that fails when the two differ.

go deeper

for a junior

Know that source produced from a description is either rebuilt each time or stored alongside it, and that stored copies can fall behind the description they came from.

for a middle

Be able to argue both sides: staleness and hand edits on one, a generator in the critical path of every build on the other, and describe the check that compares a fresh generation against what is stored.

for a senior

Show that you have operated it: stabilise output so the comparison is meaningful, make the failure message name the fix, and watch for the team that re-runs the fix reflexively without reading the change.

for a principal

Frame it as who pays. Regeneration taxes every build and every contributor's machine; committed output taxes review and invites drift. Pick according to how many consumers share the description and how reliable the generator really is.

## What is actually being decided A device-telemetry pipeline defines its message shapes once, in an external description, and a generator turns that description into source for every consumer. The question is what the repository holds: only the description, with source produced on the way to each compile, or the description **and** its emitted source, checked in like any other file. The decision is not about taste. It decides which of two failure modes you can suffer. ## Regenerating on every build What it buys: - The emitted source is a **function of the description**, so it cannot be stale. There is exactly one artifact to review and one to version. - Nobody can approve a change to generated code, because generated code never appears in a change set. - A description change reaches every consumer in the same build, so mismatches surface at compile time rather than later. What it costs: - The generator becomes a **build dependency**: it must be present, runnable on every machine and pipeline that builds, and reasonably fast, because its cost is paid on every clean build. - The output must be stable for a given input, or unrelated builds start differing for reasons nobody changed. - A new contributor cannot read the type they are working against without running the build first, and tooling that reads the tree sees nothing until generation has run. ## Committing the emitted source What it buys: - Anyone can build, read and navigate without the generator installed, which matters most when the generator is awkward to run or slow. - A description change produces a **visible diff of its consequences**, which is a genuinely good review artifact: you see that renaming one member removed four accessors. - Failures point at files that exist in the tree exactly as they were built. What it costs: - Nothing forces a regeneration. The committed files can describe last month's description, and the description stops being the source of truth in fact even though it still is on paper. - Generated files appear in review as ordinary source, which invites both review noise and — far worse — hand edits, because the files look editable. - Large emitted trees dominate the history of the repository and every diff that touches the description. | Axis | Regenerate every build | Commit the output | |---|---|---| | Staleness | Impossible by construction | Possible and silent | | Generator needed to build | Yes | No | | Review shows consequences | No | Yes | | Invites hand edits | Rarely — files are transient | Often — files look like source | | Clean-build cost | Generation on every build | None | ## The middle path Most teams that have lived with both end up doing this: 1. Commit the emitted source, so the repository is self-contained and description changes are reviewable through their consequences. 2. Regenerate in the build (or in a dedicated check) into a scratch location. 3. Fail the build when the regenerated output differs from what is committed, with a message naming the command that fixes it. That turns the committed files into a **cached, verified projection**: staleness becomes a build failure instead of a silent condition, and a hand edit is caught by the same check, because a hand edit is just a difference the description does not explain. The cost is that the check is only as good as the stability of the generator's output — a generator that emits a timestamp or an unordered member list will fail the check for no reason, and a team that learns to re-run the fix command reflexively stops reading what changed. ## How to choose when you cannot have the middle path - **Is the generator painful to run** — an unusual toolchain, a long run, a network fetch? Lean toward committing, because otherwise the pain is paid by everyone on every clean build. - **Is the emitted output large or noisy?** Lean toward regenerating; a diff nobody reads is not a review benefit. - **How many consumers regenerate from the same description?** The more there are, the more valuable it is that regeneration is automatic, because a stale copy in one consumer is the drift the single description existed to prevent. - **Who is likely to edit the files?** If the team is new to generated code, transient files teach the lesson faster than a committed tree ever will. The answer an interviewer is listening for is not a side. It is that you can name the failure mode each side owns — silent staleness versus a generator in the critical path — and that you know the check that collapses the difference.

  • What breaks the regenerate-and-compare check in practice?
    Output that is not a pure function of the description: an embedded timestamp or build host, members emitted in an unordered traversal, or a generator whose version changes the formatting. Each makes the check fail for reasons nobody changed, and the usual reaction is to switch it off. Stabilise the output first, then make the comparison a gate.
  • A consumer cannot run the generator at all. Does that settle the question?
    It settles it for that consumer, not for the description. Commit the output that consumer needs, but keep a regeneration check somewhere in the pipeline; otherwise the committed copy becomes an independent artifact that drifts. The rule to protect is that the description explains every line, not that every machine runs the generator.

saying these in an interview costs you the question

  • Says committed generated files are fine because someone will remember to regenerate
  • Treats generated files in review as ordinary code to be tidied
  • Assumes regeneration is free and ignores clean-build cost
  • Believes checking in the output makes it the source of truth
  • Cannot name a single cost of putting the generator in every build