skip to content

Some code generators regenerate an entire application from a model on every run, while others generate only specific artifacts, like a base class or a data layer, and leave the rest to be written by hand. What drives the choice to generate only part of a system rather than the whole thing, and what does a team give up by choosing partial generation?

level: seniorimportance: should knowfreq 45%

answer

  1. full generation needs a total, formal model
  2. partial = generate repetitive/mechanical parts only
  3. generation-gap or generated/ folder convention
  4. hand-written code is a second source of truth
  5. compile errors are the safety net at the seam

basics

~20 s

Full generation regenerates everything from the model every time, which only works if the model truly captures 100% of the system's logic. Partial generation, generating only the repetitive, well-understood parts (like data access boilerplate) and hand-writing the rest, is far more common because most real systems have logic that's awkward or impossible to fully model.

solid answer

~60 s

Full generation demands that the model be a complete, formal specification of system behavior, which is realistic for narrow, well-understood domains (e.g., generating a full CRUD data-access layer, or a state machine's dispatch code) but breaks down for business logic full of exceptions, external integrations, and evolving requirements that are cheaper to write directly in code than to formally model. Partial generation, following patterns like the generation gap (generated base classes, hand-written subclasses) or generating only specific layers (data access, DTOs, API client stubs, validation boilerplate) while leaving orchestration and business rules hand-written, accepts that trade-off deliberately: it targets generation at the parts of the system that are (a) highly repetitive, (b) mechanically derivable from structural information already captured elsewhere, and (c) low-risk if regenerated wrong, because they are easy to review or covered by tests. What a team gives up is the strong consistency guarantee full generation would provide, if the model changes, only the generated slice updates automatically, and hand-written code that depends on it can silently drift out of sync until a compiler error, a runtime failure, or a careful code review catches it.

go deeper

for a junior

Should recognize that not all code needs to be generated and that hand-written code and generated code can coexist in the same project.

for a middle

Should describe at least one concrete pattern (generation gap, or a generated-folder convention) for keeping the generated and hand-written portions from interfering with each other.

for a senior

Should articulate the selection criteria for what to generate (repetitive, mechanically derivable, low-risk) and clearly name the consistency guarantee that gets given up compared to full generation.

for a principal

Should reason about how to make the generated/hand-written seam safe at an organizational scale, choosing target languages, typing strictness, and code review practices specifically to make drift at that seam loud and early rather than silent and late.

## Why full generation is such a strong demand Partial generation exists because full generation makes a very strong demand: the model has to be a complete, unambiguous specification of everything the system does, not just its structure. That demand is achievable in narrow, well-bounded domains, a state machine model can fully determine dispatch code, a database schema model can fully determine a CRUD data-access layer, a REST API description can fully determine client and server stub code, because in each of these cases the mapping from model to code really is mechanical and total: there is no meaningful behavior left over that the model failed to capture. Most real production systems, however, are not that narrow. Business logic is full of conditional exceptions, integration quirks with third-party systems, performance-driven special cases, and requirements that change faster than a formal model of them could be kept accurate. Trying to fully model that and generate 100% of the system from it either produces a modeling language so complex it is effectively a second programming language (defeating the purpose), or forces developers to bypass the model constantly, which is worse than not having full generation at all, because now there are two sources of truth silently diverging. ## What earns a place in the generated slice Partial generation resolves this by being selective about what gets generated, and the selection criteria are fairly consistent across real projects — generate the parts that are: 1. **highly repetitive** (the same shape repeated across dozens of entities or endpoints, where hand-writing each instance is pure toil and a prime source of copy-paste bugs), 2. **mechanically derivable** from information that already exists in a structural, machine-readable form elsewhere (a database schema, an API description, a set of DTOs), 3. **low-risk to get wrong** because the generated slice is either simple enough to review at a glance or exercised thoroughly by generic tests. Data-access layers, DTO/mapper classes, API client and server stubs from an OpenAPI or Protobuf definition, and validation boilerplate from annotated fields are the textbook cases. Business orchestration, workflow-specific rules, and anything involving genuine judgment calls about edge cases are deliberately left to hand-written code, because trying to squeeze that logic into a model would either be impossible or would produce a model no more readable than the code it's replacing. ## The two structural patterns Mechanically, partial generation is usually implemented via one of two structural patterns, both aimed at making the boundary between generated and hand-written code explicit and safe. - **The generation-gap pattern** generates a base artifact (an abstract class, an interface, one half of a partial class) wholesale on every run, and hand-written code lives in a separate artifact that extends or implements it, so the generator never touches, and cannot corrupt, the hand-written portion. - **Alternatively**, some pipelines generate entire files but restrict generation to specific directories or file patterns by convention (a `generated/` folder that is entirely regenerated and never manually edited, versus everything outside it that the generator never writes), which is a coarser-grained version of the same ownership-splitting idea applied at the file-tree level rather than the class-hierarchy level. ## The cost: a forfeited consistency guarantee The cost of partial generation is the consistency guarantee it forfeits. With full generation, by construction, the generated system is always consistent with the model, because there is nothing else contributing to the system's behavior; regenerate, and everything is up to date. With partial generation, the model only guarantees consistency for the generated slice; the hand-written portion is a second, independent source of truth that can drift out of sync with the model or with the generated code it depends on. If a model change removes a field, the generated DTO correctly loses that field on the next regeneration, but hand-written business logic that still references the old field will only be caught if the target language's type system flags the reference at compile time; in a dynamically typed language, or if the hand-written code accesses the field through a reflection-based or loosely typed path, the drift can survive into production silently. This is the central risk teams accept: partial generation trades total consistency for practicality, and the price is paid in vigilance, code review discipline, and test coverage around the generated/hand-written seam, rather than in the generator's own guarantees. ## Where it shows up A concrete, common real-world instance: an OpenAPI-driven pipeline (OpenAPI Generator or similar) that regenerates DTO classes and a typed API client from a service's OpenAPI specification on every build, while all request-handling logic, business validation beyond simple field constraints, and orchestration across multiple service calls remain hand-written in files the generator never touches. When a field is renamed in the specification, the DTO regenerates correctly and (in a statically typed target language) the build breaks everywhere the old field name was used in hand-written code, which is the type system doing the job that full generation's inherent consistency would have done automatically; in practice, this compiler-enforced seam is exactly why teams accept partial generation's weaker guarantee, the failure mode is loud and caught at build time rather than silent, as long as the language and the generated artifact's typing are strict enough to make the seam visible.

  • Why doesn't a team just try to expand the model until it captures 100% of the business logic, so full generation becomes possible?
    Pushing a modeling language to capture every conditional exception and integration quirk of real business logic tends to make the model as complex, and as hard to read and change, as the code it was meant to replace, at which point the modeling layer has stopped saving effort and has just become a second, harder-to-tool programming language. Requirements also change faster in practice than most teams can keep a fully formal model synchronized, so the model itself becomes the stale artifact instead of the code.
  • How does the generation-gap pattern specifically protect against the consistency risk that partial generation introduces?
    It does not eliminate the risk, since hand-written code can still reference something the generated base class no longer provides, but it makes the seam explicit and structurally enforced: the generated artifact is always fully and safely regenerated because nothing hand-written lives inside it, and any drift between the two surfaces as a normal compile error at the extension point in a statically typed language, rather than as silent corruption of merged text.
  • What kind of target language or typing setup makes partial generation noticeably safer in practice, and why?
    A statically typed target language with strict compilation makes partial generation safer because any hand-written code referencing a field, method, or type that the generator has just removed or renamed will fail to compile immediately, surfacing the drift at build time. In a dynamically typed or loosely typed target, the same drift can pass compilation and only fail at runtime, or not fail at all if the code path is untested, which is a materially worse failure mode.

Full generation is like 3D-printing an entire finished product from a digital file: nothing is added by hand, so what you get always matches the file exactly. Partial generation is like 3D-printing a custom bracket to bolt onto an otherwise hand-assembled machine: you save effort on the repetitive part, but the hand-assembled portion can be wired up wrong in ways the printer has no way to catch.

saying these in an interview costs you the question

  • Assumes generation should always aim to cover 100% of the codebase eventually
  • Cannot name any concrete criterion for what makes a part of a system a good generation candidate
  • Believes partial generation gives the same consistency guarantee as full generation
  • Has no answer for what happens when hand-written code depends on something the model removes
  • Confuses partial generation with simply generating less code overall for no particular reason

context