skip to content

Which UML diagram types remain genuinely useful for describing software architecture, what does each one express, and how do they relate to C4?

level: middleimportance: should knowfreq 45%

answer

  1. Component = lollipop provided / socket required
  2. Deployment = nodes + artefacts + multiplicity
  3. Sequence = lifelines, alt/opt/loop, sync vs async
  4. C4 = abstractions; UML = notation
  5. Class diagrams: generate or skip

basics

~20 s

Mainly three: component diagrams (units of functionality and the interfaces they provide/require), deployment diagrams (which artefacts run on which nodes), and sequence diagrams (the ordered messages of one scenario over time). They map onto C4's component, deployment, and dynamic views.

solid answer

~50 s

UML has 14 diagram types; for architecture only a few earn their keep. **Component diagrams** show units of functionality with *provided* and *required* interfaces (the ball-and-socket notation), making dependencies explicit as contracts rather than arrows — good for showing plug-points and substitutability. **Deployment diagrams** show nodes (devices, execution environments, e.g. regions, VMs, pods, browsers) with artefacts deployed onto them and communication paths between them, plus multiplicity — the only standard view that answers "where does this actually run, and how many of them are there?". **Sequence diagrams** show participants as lifelines and the time-ordered messages between them, with synchronous vs asynchronous arrows, activation bars, returns, and optional/loop/alt fragments — the best notation for one end-to-end scenario, including failure paths and timeouts. Class diagrams are useful mainly at code level; use-case, state, and activity diagrams answer behaviour questions, not structure. C4 reuses these ideas: its deployment diagram is essentially a UML deployment view, and its dynamic diagram is a sequence/collaboration view.

go deeper

for a junior

Name component, deployment, and sequence diagrams and say in one line what each shows. Knowing that sequence = time order is the key point.

for a middle

Add the notation details that carry meaning: lollipop/socket interfaces, nodes with artefacts and multiplicity, sync vs async arrowheads and alt/loop fragments — and state that C4 is abstractions while UML is notation.

for a senior

Choose views per question being asked, explain how deployment views feed threat modelling and resilience discussions, and cover generating sequence views from traces plus the drift risk of hand-drawn class and deployment diagrams.

for a principal

Define the org's minimum viable notation set, when strict UML is warranted (regulated/model-driven contexts), and how diagram views tie into ADRs, review gates, and automated dependency/infra checks.

## Background **UML** (Unified Modeling Language) is an OMG standard notation with two families: **structure** diagrams (class, component, composite structure, deployment, package, object, profile) and **behaviour** diagrams (use case, sequence, communication, activity, state machine, timing, interaction overview). Full UML is large and few teams use it strictly; the pragmatic position — and Simon Brown's, in C4 — is to borrow the handful of views that answer real architectural questions and to always include a key so readers do not need UML fluency. ## The three that pay for themselves ### 1. Component diagram — *what are the parts and what contracts join them?* - **Element:** a component — a modular unit with encapsulated contents, replaceable within its environment. - **Key notation:** **provided interface** drawn as a lollipop (a circle on a stick) and **required interface** drawn as a socket (a half-circle). A socket snapping onto a lollipop is an *assembly connector*: it says "this part needs *that contract*", not "this part needs *that box*". - **Why it matters architecturally:** it expresses dependency on an *interface* rather than on an implementation, which is precisely what makes a part substitutable (swap the payment provider, swap the persistence adapter). Ports and delegation connectors let you show a component's external boundary separately from its internals. - **Trade-off:** the ball-and-socket notation is unfamiliar to many engineers today; without a key it becomes decoration. C4's level-3 component diagram makes the same point with plain labelled boxes. ### 2. Deployment diagram — *where does it run, and how many?* - **Elements:** **nodes** (`<<device>>`, `<<execution environment>>`) which can nest — region → availability zone → Kubernetes cluster → pod → JVM; **artefacts** (deployable files: a jar, an image, a bundle) *deployed onto* nodes; **communication paths** (lines between nodes, labelled with protocol); and **multiplicity** (`[3]`, `1..*`) on nodes. - **Why it matters:** it is the only standard view that separates *logical* structure from *physical* placement. It answers redundancy, co-location, network-boundary, latency-domain, and blast-radius questions, and it is the natural input to threat modelling (every communication path crossing a trust boundary needs a control). - **C4 link:** a C4 deployment diagram is exactly this — container instances mapped onto infrastructure nodes, with the explicit reminder that one container may have many instances. ### 3. Sequence diagram — *what happens, in what order, for this scenario?* - **Elements:** participants across the top, vertical **lifelines** downward as time, **messages** as arrows between lifelines (solid filled arrowhead = synchronous call, open arrowhead = asynchronous, dashed = return), **activation bars** showing when a participant is busy, and **combined fragments** — `alt` (alternatives), `opt` (optional), `loop`, `par` (parallel), `ref` (reference to another diagram). - **Why it matters:** structure diagrams cannot show ordering, concurrency, retries, timeouts, or compensating actions. A sequence diagram of the checkout flow, including the failure branch, communicates more about a distributed design than any static picture. Also the standard way to document saga/orchestration flows and authentication handshakes (OAuth flows are almost always drawn this way). - **Trade-off:** one diagram covers one scenario. Pick the two or three scenarios that dominate the design; do not attempt exhaustive coverage. ## The rest, briefly - **Class diagram** — precise for a component's internals and for domain models; C4 puts it at level 4 and recommends generating it or skipping it. - **Package diagram** — useful for module/layer dependency rules; often better enforced by a tool (dependency checks in CI) than drawn. - **State machine diagram** — valuable when a core entity has a non-trivial lifecycle (order, payment, subscription) with legal and illegal transitions. - **Activity / use case / communication / timing** — occasionally useful, rarely the right first artefact for architecture. ## How C4 and UML relate They are not competitors at the same level. C4 defines a **set of abstractions and zoom levels** and is notation-agnostic; UML defines a **notation** (and metamodel). You can render C4 levels using UML shapes, using C4-PlantUML macros, or using plain boxes. Common pragmatic stack: C4 for levels 1–3, a UML-style deployment diagram for infrastructure, UML sequence diagrams for the important runtime scenarios, and no hand-maintained class diagrams at all. ## Edge cases and pitfalls - Using a **sequence diagram as a design substitute**: a beautiful 40-message diagram often signals a chatty, over-coupled design; the diagram is doing its job by making that visible. - **Half-remembered UML** (wrong arrowheads, association vs dependency confusion) is worse than plain boxes with a key, because it implies precision that isn't there. - **Generated sequence diagrams** from traces (distributed tracing exports) are accurate and cheap; treat them as evidence, and hand-draw only the intended/target flow. - **Deployment diagrams drift fastest** in cloud environments; consider generating them from infrastructure-as-code or keeping them at the level of stable topology (regions, tiers) rather than instance names.

  • If you could keep only one UML-style diagram for a distributed system, which and why?
    A deployment diagram, because it is the only view that answers where things run, how many instances exist, and which network/trust boundaries are crossed — the questions that drive availability, latency, cost, and security. A close second is a sequence diagram of the single most critical end-to-end flow.
  • What does the ball-and-socket notation on a UML component diagram express that a plain arrow does not?
    That the dependency is on a *contract* (a named interface), not on a specific implementation box. The socket says "I require this interface"; anything providing it can be plugged in, which is exactly what makes the component substitutable and testable.
  • Is UML dead?
    Strict, full UML modelling is rare outside regulated or model-driven contexts. But the vocabulary and three of its views survive everywhere — sequence diagrams especially — and tools like PlantUML and Mermaid keep them cheap. The pragmatic answer: borrow the useful views, always ship a key, and don't demand fluency from readers.

saying these in an interview costs you the question

  • Calling every rectangle-and-arrow picture 'a UML diagram'
  • Using solid vs open arrowheads inconsistently on sequence diagrams, so sync and async are indistinguishable
  • Drawing a deployment diagram without multiplicity, hiding whether anything is redundant
  • Treating C4 and UML as mutually exclusive — C4 is abstractions, UML is notation
  • Maintaining hand-drawn class diagrams that the code contradicts within a sprint

context