skip to content

Model-Driven Design

Treating explicit models as primary artifacts and generating or deriving implementation from them: MDA and metamodels, transformations, DSLs, and keeping model and code aligned.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

29

In Model-Driven Design, what does it mean to treat the model as the 'primary artifact' instead of treating diagrams as throwaway documentation?

level: juniorimportance: must knowfreq 70%

answer

  1. model = blueprint not poster
  2. translation gap
  3. ubiquitous language enforcement
  4. generated vs disciplined sync
  5. round-trip engineering

basics

~20 s

The model isn't a picture you draw once and forget — it's the actual blueprint that the code, database, and APIs are built from and kept in sync with, so it stays accurate and useful as the system evolves.

solid answer

~40 s

Treating the model as the primary artifact means the model — the set of concepts, relationships, and rules describing the domain or system — is the authoritative description everyone, including the code, must agree with, not a one-off diagram made for a kickoff meeting and never touched again. Documentation-only models drift from reality within weeks because nothing enforces alignment; a primary-artifact model instead drives implementation, either literally (code generation, executable models) or disciplinarily (the team refuses to let code diverge from the model's vocabulary and structure, as in Domain-Driven Design's ubiquitous language). The practical test: if you change the model, does something concrete have to change too, and if the code changes, does someone update the model? If the answer is no on either side, it's decoration, not a primary artifact.

go deeper

for a junior

Should be able to state that the model isn't a throwaway diagram and give one concrete example of code reflecting model vocabulary; doesn't need to know MDA or generation tooling.

for a middle

Should distinguish the generative (tool-driven) approach from the disciplinary (code-review-enforced) approach and know at least one has to be actively maintained.

for a senior

Should be able to name the translation-gap problem this solves, describe how their own team keeps code and model aligned in practice, and identify where it has slipped before.

for a principal

Should be able to weigh model-primacy discipline against delivery speed at an org level and design the process or tooling — review gates, codegen, or lightweight linting — that keeps a whole organization's models honest over years, not just one team's sprint.

## The artifact of record In most software projects, 'the model' shows up as a UML diagram in a wiki page, drawn during initial design and never opened again. **Model-Driven Design rejects that pattern.** It insists the model — a structured representation of the important concepts, their relationships, invariants, and behavior — is the artifact of record: the thing developers, testers, and business stakeholders treat as authoritative when they disagree about what the system does. Code, database schemas, and API contracts are expected to be derived from or kept faithful to the model, not the other way around. ## Two flavors Mechanically this plays out in two flavors. 1. **The heavyweight version is generative.** Tooling (in the OMG's Model Driven Architecture tradition, using UML plus the Meta Object Facility) reads a formal model and produces skeleton code, database DDL, or even a fully executable system, so the model is quite literally compiled. 2. **The lighter, more common version** — closer to Eric Evans' Domain-Driven Design usage of the term 'model-driven design' — **is disciplinary rather than mechanical.** There is no generator, but the team enforces that every class, method, and table name traces back to a concept in the model, and every model change forces a corresponding code change in the same commit. | Flavor | What keeps the code faithful | |---|---| | Generative, heavyweight | Tooling reads the formal model and produces the code | | Disciplinary, lighter | The team enforces that every name traces back to a concept in the model | Both flavors share the same goal: **eliminate the gap between what was designed and what was built.** ## The translation gap The problem this solves is an old one — the **translation gap**. A business analyst writes a spec, an architect draws a diagram interpreting the spec, and a developer writes code interpreting the diagram; at each hop, information is lost or silently reinterpreted, and six months later nobody can say with confidence what the system is actually supposed to do, because the diagram — if anyone still has it — no longer matches the code. Making the model the primary artifact collapses those hops: there's one representation that must stay true, and everything else answers to it. ## The trade-off: discipline versus speed The trade-off is discipline versus speed. Keeping a model authoritative costs continuous maintenance effort — every schema tweak, every new business rule, has to be reflected in the model as well as the code, and code reviews have to check for drift, not just correctness. Teams under deadline pressure routinely let this slip: - a developer patches the generated code directly to hit a deadline, or - adds a field to a table without updating the corresponding model element. In generative setups this is sharper still, because hand-edited generated code gets silently overwritten the next time someone regenerates from the model — the classic **round-trip engineering** failure mode; teams work around it with generator 'protected regions' or by abandoning generation for hand-written code that merely mirrors the model's vocabulary. ## What it cashes out to day-to-day A concrete real-world instance: shops using Domain-Driven Design tactically keep an aggregate concept — say, an `Order` with its `OrderLine` children — as the model, and insist that - the `Order` class in code, - the `orders` and `order_lines` tables, and - the API's `Order` DTO all use exactly that vocabulary and exactly that consistency boundary, with no helper classes carrying logic the model doesn't describe and no direct SQL updates to `order_lines` that bypass the `Order` aggregate's invariants. That discipline is what 'model as primary artifact' cashes out to day-to-day: not a diagram, but a standing rule that code must justify itself against the model, and the model must justify itself against the domain. ## How it fails Failure shows up gradually rather than as an outage: - new hires read the stale model and build features that contradict how the code actually behaves; - two teams implement the 'same' concept differently because each trusted a different stale copy of the model; - and eventually the model is quietly deleted from onboarding material because 'it's wrong anyway' — at which point the team has reverted to code-only design with all the translation-gap risk that model-driven design was meant to remove.

  • If a team has no code-generation tooling at all, can they still practice 'model as primary artifact'?
    Yes — this is the Domain-Driven Design flavor of model-driven design: there's no generator, but the team enforces by convention and code review that class names, method names, and structure trace directly back to the model's ubiquitous language. The discipline is social and process-based rather than tool-enforced, which is more fragile but far cheaper to adopt than a full generative toolchain.
  • What's the first sign a team has stopped treating the model as primary and started treating it as documentation?
    Code review comments stop referencing the model, and a change to business rules gets implemented in code without anyone updating the model artifact in the same change. Once that happens once without pushback, the model starts drifting and rarely recovers without a deliberate re-sync effort.

It's the difference between an architect's blueprint that the construction crew is legally required to follow, and that gets updated the moment a wall moves, versus a lobby poster showing an 'artist's impression' of the building that nobody checks against the actual construction.

saying these in an interview costs you the question

  • calls the model 'just a diagram we made at the start'
  • can't explain what forces the model and code to stay in sync
  • conflates having a UML diagram with practicing model-driven design
  • no answer for what happens when code and model disagree

context

open as a page

In the context of Domain-Specific Languages (DSLs), what is the difference between an internal (embedded) DSL and an external (standalone) DSL, and can you give a concrete example of each?

level: juniorimportance: must knowfreq 55%

basics

~20 s

An internal DSL is written using the normal syntax of an existing programming language, just arranged to read like a mini-language for one job (e.g. a Kotlin build script). An external DSL has its own custom syntax and needs its own parser to be understood, like SQL or a regular expression.

open as a page

What is OMG's Model Driven Architecture (MDA), and what is the difference between a Platform-Independent Model (PIM) and a Platform-Specific Model (PSM)?

level: juniorimportance: must knowfreq 55%

basics

~20 s

MDA is a way of building software by first drawing a model of what the system does, ignoring the technology, then turning that model into one tied to a specific technology (like Java or a database), and finally generating code from it.

open as a page

When keeping a UML or domain model in sync with running code, what is the difference between 'forward engineering' (generating code from a model) and 'reverse engineering' (deriving a model from existing code)?

level: juniorimportance: must knowfreq 65%

basics

~20 s

Forward engineering writes the model first, then a tool generates code from it. Reverse engineering starts from existing code and builds or updates a model to describe it. Same goal—keep model and code matching—different starting point.

open as a page

In model-driven development, one kind of transformation turns a model into working code or documentation, while another kind turns one model into a different model. What is the practical difference between these two kinds of transformations, and why does the choice matter for a real pipeline?

level: juniorimportance: must knowfreq 55%

basics

~10 s

Model-to-text turns a model into plain text, like source code, mostly a one-way trip. Model-to-model turns one structured model into another structured model, so it stays machine-readable and can be transformed again.

open as a page

What do the CIM, PIM, and PSM layers represent in the Model Driven Architecture (MDA) approach to Model-Driven Design, and how does a change flow between them?

level: middleimportance: must knowfreq 55%

basics

~20 s

They're three levels of the same design, from business-language description (CIM), to a tech-neutral system design (PIM), to a version tied to one specific technology (PSM) — each layer adds detail the previous one deliberately left out.

open as a page

When designing the grammar for an external textual DSL - for instance using a tool like Xtext with EBNF-style grammar rules - what does the grammar actually define, and how does that feed into the rest of the language's tooling such as parsing, the abstract model, and editor support?

level: middleimportance: must knowfreq 45%

basics

~20 s

The grammar is a set of rules describing what a valid sentence in your mini-language looks like, similar to a sentence-structure rule like 'an order has a customer name and a list of items.' Tools like Xtext turn that grammar into both a parser and an in-memory data model, and then automatically generate editor features like syntax highlighting and autocomplete from the same rules.

open as a page

OMG's Meta Object Facility (MOF) defines a four-layer metamodeling stack, M0 through M3. What does each layer represent, and how does UML fit into this stack?

level: middleimportance: must knowfreq 45%

basics

~20 s

MOF is a stack of four levels, like Russian nesting dolls: M0 is the real data at runtime, M1 is a model of it (like a UML diagram), M2 is the language used to draw that diagram (UML's rules), and M3 is the top rule-set (MOF) that even defines UML's own rules and defines itself.

open as a page

How does annotation-driven synchronization (e.g., embedding metadata like JPA's @Entity/@Column or OpenAPI annotations directly in source code) keep a model and code aligned differently from maintaining a separate external model file?

level: middleimportance: must knowfreq 60%

basics

~20 s

Instead of a separate diagram file, you put small tags (annotations) right on the code, like @Entity on a class. A tool reads those tags to generate the model (like a database schema or API spec), so the model and code can never drift apart—they're the same file.

open as a page

A team generates Java classes from a domain model, but developers need to hand-write business logic inside those same generated files. If the model changes and the generator runs again, how do you avoid destroying the hand-written code, and what does this problem reveal about round-trip engineering?

level: middleimportance: must knowfreq 65%

basics

~20 s

Mark off areas in the generated file as 'protected' or split hand-written code into a separate file the generator never touches. Round-trip engineering is about letting generated output and manual edits coexist without one wiping out the other.

open as a page

Two common ways to build a model-to-text generator are template-based generation, filling text templates with model values, and rule-based generation, applying declarative mapping rules to produce output. What is the mechanical difference between the two approaches, and when would you choose one over the other?

level: middleimportance: must knowfreq 60%

basics

~20 s

Template-based generation fills in text templates that look like the output, with holes for model data. Rule-based generation runs a set of if-this-then-produce-that rules that decide what to build. Templates are easier to read; rules scale better for complex mappings.

open as a page

Model-Driven Design treats the model as the 'single source of truth' for a system. What actually happens in production teams when the code and the model start to drift apart, and how do teams recover from it?

level: seniorimportance: must knowfreq 60%

basics

~20 s

When developers change code without updating the model, or vice versa, the model stops being trustworthy — new features get built on wrong assumptions, and fixing it usually means a deliberate re-sync project rather than something that happens automatically.

open as a page

In a textual DSL editor - for example one built with Xtext - what are 'scope rules' (a scoping provider) responsible for, and why does getting them right matter so much once a DSL model spans many files?

level: seniorimportance: must knowfreq 40%

basics

~20 s

Scope rules decide which named things - like another element defined elsewhere - are visible and can be referenced from a given point in the file, similar to how you can only use a variable after it's declared. Getting this wrong makes autocomplete suggest the wrong things or 'go to definition' silently jump to the wrong place.

open as a page

Where do PIM and PSM sit in the M0–M3 MOF stack, and what does a QVT-based PIM-to-PSM transformation actually operate on?

level: seniorimportance: must knowfreq 35%

basics

~20 s

PIM and PSM are both just models sitting at the same level in the stack, built using rules like UML's. A transformation tool reads the PIM according to those rules and writes out a PSM according to a possibly different set of platform rules, using a defined mapping between the two.

open as a page

What techniques can a team use to actively detect when a design model (e.g., an architecture diagram or domain model) has drifted out of sync with the running code it's supposed to describe, rather than discovering the mismatch by accident?

level: seniorimportance: must knowfreq 55%

basics

~20 s

You can automatically compare the model against the real code—like checking that every class the diagram shows still exists, and no new dependency was added that the diagram doesn't show—and fail a build or flag a warning when they don't match, instead of just hoping someone notices.

open as a page

In a layered Model-Driven Design approach, such as platform-independent versus platform-specific models, what does it mean for an abstraction layer to 'leak,' and why does that undermine the whole approach?

level: middleimportance: should knowfreq 45%

basics

~20 s

Leaking means details from a lower, more technical layer sneak into a higher, supposedly technology-neutral layer, like a database-specific trick showing up in the 'business' model, which defeats the whole point of keeping them separate.

open as a page

A platform team is deciding whether to build a new configuration language as an internal DSL embedded in their host language, or as an external DSL with its own grammar and parser. What factors should drive that decision, and what's a concrete failure mode of picking the wrong one?

level: middleimportance: should knowfreq 50%

basics

~20 s

Pick internal if the authors are your own developers and you want low build cost with full access to the host language. Pick external if you need a stricter, safer, or non-programmer-friendly notation. Picking wrong means either wasting months on a parser nobody needed, or handing power users a dangerously open language when a safe sandbox was required.

open as a page

UML is itself defined as a MOF metamodel. What does that mean concretely, and how does extending UML via a Profile differ from defining a brand-new MOF metamodel?

level: middleimportance: should knowfreq 28%

basics

~20 s

UML's own rules, like what a Class or Association is, are written using MOF's building blocks. If you just need to tag extra info onto UML diagrams, you use a lightweight Profile, like stickers on a form. If you need a whole new kind of diagram with its own rules, you define a brand-new metamodel instead.

open as a page

In round-trip engineering—where a tool must support both generating code from a model and updating the model when the code changes—what specifically makes preserving hand-written customizations across repeated round trips difficult?

level: middleimportance: should knowfreq 45%

basics

~20 s

The tool has to remember which parts of the code were auto-generated and which parts a person wrote by hand, every time it regenerates, so it doesn't delete a developer's custom logic. Keeping that separation correct through many edit cycles is the hard part.

open as a page

DSLs are often described as enabling 'executable' or 'generative' models. What's the practical difference between those two ways of making a DSL model actually do something, and what trade-offs come with each?

level: seniorimportance: should knowfreq 35%

basics

~20 s

An executable model is run directly by an interpreter that reads the model and carries out its behavior on the spot, like a script. A generative model is instead used as a blueprint to produce other code, such as Java or SQL, which is then compiled and run separately. Executable is more immediate; generative gives you inspectable, tunable output code.

open as a page

Bidirectional transformation frameworks (e.g., QVT, or lens-based approaches) aim to keep a model and a derived view (or two related models) consistent automatically in both directions. What makes writing a correct bidirectional transformation harder than writing two independent one-directional transformations, and when would a team be better off NOT attempting full bidirectionality?

level: seniorimportance: should knowfreq 35%

basics

~20 s

A good two-way sync tool must guarantee that going model→view→model (or the reverse) lands back exactly where you started, which is much harder to prove than just writing 'convert A to B' and separately 'convert B to A' by hand, because the two halves can quietly disagree. If changes are rare or one side is clearly authoritative, it's often not worth the complexity.

open as a page

Some code generators regenerate an entire application from a model on every run, while others generate only specific artifacts, like a base class or a data layer, and leave the rest to be written by hand. What drives the choice to generate only part of a system rather than the whole thing, and what does a team give up by choosing partial generation?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Full generation regenerates everything from the model every time, which only works if the model truly captures 100% of the system's logic. Partial generation, generating only the repetitive, well-understood parts (like data access boilerplate) and hand-writing the rest, is far more common because most real systems have logic that's awkward or impossible to fully model.

open as a page

QVT (Query/View/Transformation), the OMG standard, and ATL (ATLAS Transformation Language) are both widely used model-to-model transformation languages, but they take different approaches: QVT offers a declarative relational style (and an imperative operational style), while ATL is a hybrid declarative/imperative language. What are the practical consequences of choosing a more declarative versus a more imperative transformation style?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Declarative transformation rules say what relationship must hold between source and target, and let the engine figure out how to make it happen. Imperative rules say exactly what steps to run. Declarative is easier to reason about and check consistency; imperative gives more control for tricky, order-dependent logic.

open as a page

Given that treating the model as a governing, single-source-of-truth artifact sounds appealing, when would you deliberately choose NOT to adopt heavyweight Model-Driven Design — formal CIM/PIM/PSM layers plus generative tooling — for a project?

level: principalimportance: should knowfreq 35%

basics

~20 s

When the project is small, the platform won't change, or the domain is still being figured out — the cost of building and maintaining formal layered models and generators outweighs any benefit, and a lighter, code-first approach ships faster and adapts better.

open as a page

MDA promised that a platform-independent model plus automated transformations would insulate business logic from platform churn and reduce cross-platform maintenance cost. Given that promise, why did full-pipeline MDA see only narrow industry adoption, and under what conditions would you actually reach for it today?

level: principalimportance: should knowfreq 28%

basics

~20 s

Building and maintaining the models, transformation rules, and generators cost more than most teams ever got back, because most software only ever targets one platform and changes more through new features than through swapping technology — so the insurance MDA sold rarely paid out. It's worth it mainly for large, long-lived systems that genuinely target multiple platforms or where regulators demand traceable models.

open as a page

An organization has run a model-driven code generation pipeline in production for several years, with dozens of transformation rules and templates maintained by a rotating set of engineers. What are the characteristic ways such a pipeline degrades over time, and when should a team conclude that model-driven generation is no longer the right investment for a given system?

level: principalimportance: should knowfreq 25%

basics

~20 s

Over years, the transformation rules and templates themselves become their own hard-to-maintain codebase, generated and hand-written code quietly drift apart, and fewer people understand how the pipeline works. It stops being worth it once the generator's upkeep costs more than just writing the code directly.

open as a page

The term 'Model-Driven Design' is used both by the OMG's Model Driven Architecture community and by Eric Evans' Domain-Driven Design. What's the key difference between what each means by it, and why does the confusion matter in practice?

level: seniorimportance: nice to knowfreq 20%

basics

~10 s

OMG's version means machines transform formal models into code automatically; Evans' Domain-Driven Design version means developers hand-write code that mirrors an evolving domain model — same name, very different amount of automation and formality.

open as a page

Beyond textual DSLs, some domains use graphical DSL editors - diagram-based tools for things like state machines or process flows. What trade-offs make graphical notation the right call for some domains and the wrong call for others?

level: principalimportance: nice to knowfreq 20%

basics

~20 s

Graphical DSLs, diagrams you drag and connect, work well when the domain is naturally visual with a small, stable set of elements per screen, like state machines or flowcharts. They get unwieldy for anything with lots of detail, like complex logic or large data structures, where text is faster to write, search, and compare between versions.

open as a page

Triple Graph Grammars (TGGs) are one formal approach to specifying and incrementally maintaining consistency between two related models (e.g., a platform-independent model and a platform-specific model) as either one changes. How does TGG's approach to incremental synchronization differ from simply re-running a full bidirectional transformation from scratch after every change, and what does a team give up by adopting a rule-based grammar formalism like this?

level: principalimportance: nice to knowfreq 15%

basics

~20 s

Instead of throwing away and rebuilding the whole target model every time the source changes even slightly, TGGs figure out just the small piece that actually needs to update, using pre-defined pattern rules—like patching instead of rewriting. The cost is that you must formally define every allowed pattern up front, which is a lot of specialized modeling work.

open as a page