skip to content

Notation and Decomposition

The shapes a data-flow diagram gives you, process, data store, external entity and flow, and how deep to decompose. Interviewers judge whether your notation is disciplined or improvised.

on this pageshow

explore

questions

12

What does a level-0 context diagram show in a data-flow-diagram threat model?

level: juniorimportance: must knowfreq 70%

answer

  1. start by drawing nothing inside
  2. who talks to it, not how it works
  3. one process, external entities, flows
  4. the outer boundary and what crosses it
  5. level 0 of the leveling hierarchy

basics

~20 s

A level-0 context diagram draws the whole system as one process, every external entity that talks to it, and the flows between them. It fixes scope and shows what crosses the outer trust boundary, with no internal structure.

solid answer

~50 s

The context diagram sits at the top of the leveling hierarchy: a single process for the entire system under review, every external entity around it — end users, operators, partner systems, identity providers, batch feeds — and every flow in and out, annotated with protocol and data class. The one boundary it always carries is the outer one, so everything crossing it is untrusted input coming in or a disclosure going out. For a national driving-licence renewal service, level 0 is what a ministerial sign-off actually reads: applicants, the photo-and-identity checking service, the payment provider, the police records feed, the print bureau. It already yields real threats — an anonymous applicant spoofing another citizen, tampering with a submitted document in transit, repudiating a renewal they claim they never requested. What it cannot show is elevation of privilege between two internal components; that needs a lower level.

go deeper

for a junior

Be ready to draw it on request: one box for the system, external entities around it, labelled arrows, and a line for the outer trust boundary. Say out loud that nothing internal is shown.

for a middle

Explain what the level buys you — a scope everyone can dispute, and the outside-attacker entry points enumerated before any internal structure exists. Name the interfaces people forget: admin, batch, monitoring.

for a senior

Show that you can produce real threats from level 0 alone and can state its limit: no internal privilege escalation, no component-to-component boundary. Be able to say which pending decision would force you to decompose.

for a principal

Own the artifact politics: the context diagram is the page a sign-off actually reads and the one that stays true across releases. Decide who maintains it and what change to the external interface set triggers a re-review.

## The leveling hierarchy A data-flow diagram (DFD) threat model is not one drawing. It is a small hierarchy of drawings at increasing detail, and the convention borrowed from structured analysis names them by level. - **Level 0, the context diagram.** The entire system under review is drawn as a *single process*. Around it sit the external entities it exchanges data with, and between them the flows. Nothing inside the system is shown. - **Level 1.** That single process is opened up into its major internal processes and data stores, with the same external flows still arriving at the edge. - **Level 2 and below.** One level-1 process is opened up in turn, showing its internals. The context diagram is therefore the answer to a scoping conversation that has already happened: *this* is the thing we are modeling, and *these* are the parties it talks to. ## What belongs on it Four things and no more: 1. **One process** — the system, named as the business names it. 2. **External entities** — anything that interacts with the system but is not part of it, and that you do not control or model internally: citizens and customers, administrators and operators, partner services, identity providers, payment providers, regulators pulling reports, batch feeds, the monitoring backend you ship logs to. 3. **Flows** — every data path in and out, with a direction, a protocol and a data class (personal data, credentials, money instructions, telemetry). A flow with no annotation is half a flow. 4. **The outer trust boundary** — a line separating the system from everything around it. Every flow that crosses it is an entry point (untrusted input) or an exit point (a disclosure or an action taken on the outside world). What does **not** belong: internal components, internal databases, deployment detail such as hosts, subnets and load balancers, and message ordering. Those are the business of lower levels or of other notations entirely. ## Worked example A national driving-licence renewal service. The level-0 diagram has one box labelled *Licence Renewal Service*, and around it: the applicant (submits an application, a photo and a scanned proof of address; receives a status), the photo-and-identity verification provider, the payment provider, the police records feed that supplies disqualification data, the print bureau that receives an authorised production instruction, and the ministry's reporting users. That single page is what a non-technical sign-off reads and approves, and it is enough to enumerate a first tier of threats: | Flow crossing the boundary | Representative threat | | --- | --- | | Application submitted by applicant | Anonymous applicant impersonates another citizen (spoofing) | | Scanned document uploaded | Document altered before or during submission (tampering) | | Renewal confirmation | Applicant later denies having requested the renewal (repudiation) | | Instruction to the print bureau | An unauthorised instruction produces a genuine document (integrity of the physical artifact) | | Records feed from police | Availability loss stops all renewals (denial of service) | Notice the asset in play: personal data and the integrity of an official document, not a generic customer database. The context diagram is where you name that asset, because it is where you name who is on the other side of every line. ## Why the level exists at all Three jobs. **It fixes the boundary of the exercise in a form people can disagree with.** A one-page diagram makes a missing interface visible: someone in the room says *what about the call-centre agents who can amend an application?*, and you have found an entry point before you have modeled a single internal component. **It gives you the outside-attacker view for free.** The threats above are the ones an anonymous or lightly authenticated outsider can reach, and they are usually the highest-exposure ones. Starting inside the system, at level 1, tends to bury them. **It is the stable artifact.** Internals churn every sprint; the set of parties the service talks to changes rarely. A context diagram stays true for a long time, which is why it is the page you put in front of a sign-off. ## Common failures - Drawing internal components on it, so that it becomes a bad level-1 diagram. - Omitting the boring external entities — the monitoring egress, the nightly reconciliation feed, the admin console reached from an office network — because they feel internal. They are entry points, and they are frequently the least protected. - Treating it as too shallow to yield threats and skipping straight to level 1. - Confusing it with a deployment diagram. The context diagram is about *who exchanges what data with the system*, not about where anything runs. ## What it cannot tell you Anything that depends on internal structure: privilege escalation from one internal component to another, an internal store read by a component that should not read it, a boundary between two teams' services. If the pending question is one of those, you decompose. The context diagram is the start of the hierarchy, not the whole of it.

  • Do trust boundaries appear on a context diagram, or only at lower levels?
    At least one appears: the outer boundary around the single system process. Every flow crossing it is an entry point or an outward disclosure, which is exactly the analysis the level-0 diagram is for. Boundaries *inside* the system — between components under different credentials, accounts or teams — only become drawable once you decompose, and finding one is itself a reason to go a level deeper.
  • How do you know a context diagram is complete?
    Walk the interfaces rather than the architecture: who authenticates, who administers, who is billed, who receives exports, what runs on a schedule, where telemetry and backups go, and which third parties are called out. Each answer is an external entity or a flow. Any party a person can name that is not on the page is a gap, and the ones people forget are admin, batch and monitoring paths.
  • Can a context diagram be too big?
    It can, and the fix is usually scope rather than layout. Thirty external entities normally means you have drawn a whole department as one system. Either the thing under review is genuinely a platform — in which case group entities by role and accept the page — or you should be modeling one service inside it, with the neighbours becoming externals.

It is the diagram on the front of an embassy building: who may approach, which windows they queue at, and what they hand over. It says nothing about the corridors inside.

saying these in an interview costs you the question

  • Draws internal components and databases on the context diagram
  • Says a context diagram is too shallow to produce threats
  • Omits admin, batch and monitoring interfaces as internal
  • Confuses the context diagram with a deployment or network map
  • Leaves flows unlabelled, with no direction or data class

context

open as a page

On a threat-modeling data-flow diagram, what do process, data store, external entity and flow each represent?

level: juniorimportance: must knowfreq 76%

basics

~20 s

A process is running code that transforms data, a data store is passive data at rest, an external entity is a person or system you do not control, and a flow is data in motion between them.

open as a page

What does a data-flow diagram deliberately omit, and which threats hide in those omissions?

level: middleimportance: must knowfreq 62%

basics

~20 s

A data-flow diagram shows what data moves where, not when or in what order. It omits control flow, sequencing, retries and error paths, so replay, race, double-processing and failure-path leak threats stay invisible on it.

open as a page

Which sources do you rebuild a data-flow diagram from for an undocumented legacy service?

level: middleimportance: must knowfreq 62%

basics

~20 s

Rebuild the diagram from artifacts that describe the running system: infrastructure-as-code and its state, route and proxy configuration, outbound calls grepped from the repository, identity bindings and database grants, and request telemetry. Confirm each edge with an operator.

open as a page

Why does a data flow on a threat-modeling DFD need protocol, direction and data class on it?

level: middleimportance: should knowfreq 54%

basics

~20 s

An unlabelled arrow hides the three things the analysis needs: what channel protects the data, which way the data actually moves, and how valuable it is. Without those, every flow looks identical and threats collapse into guesswork.

open as a page

When does a deployment diagram reveal a threat that a component diagram cannot?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A component diagram shows logical separation; a deployment diagram shows where code actually runs. Co-location threats, such as two supposedly separate services sharing one process, one host and one credentials file, appear only in the deployment view.

open as a page

How do you decide how deep to decompose a data-flow diagram for a threat model?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Decompose only where more detail would change a decision. Stop when the next level would produce elements you would mitigate identically, when no boundary hides inside the box, or when its owners will model it themselves.

open as a page

Your recovered data-flow diagram has an edge telemetry proves but no configuration explains — what now?

level: seniorimportance: should knowfreq 44%

basics

~20 s

An observed flow is real evidence, so the configuration set is what is incomplete. Characterise the edge from telemetry — endpoints, identity, protocol, volume, schedule — then walk it with the on-call operator, keeping it drawn until someone accounts for it.

open as a page

In DFD leveling, what must a child diagram preserve from the process it decomposes?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

The child diagram must show the same flows in and out as the parent process it expands — same sources, same sinks, none added or dropped. Only internals are new, and numbering carries down: process 3 expands into 3.1, 3.2.

open as a page

When is a single process box on a data-flow diagram actually hiding several processes?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

Whenever the box covers code running at different privileges, reached by different callers, or handling different data classes. One process shape asserts one identity and one set of controls, so anything inside it that differs disappears.

open as a page

What does mandating a data-flow diagram for every design review cost you?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Standardising on one diagram makes reviews comparable and the practice teachable, but it silently caps what teams can find. Protocol, retry, error-path and co-location threats do not fit the notation, so nobody raises them and the gap reads as safety.

open as a page

How complete must a recovered data-flow diagram be before you start enumerating threats on it?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

A recovered diagram is ready when it answers the decision at hand: every trust-boundary crossing drawn, every data store's class named, every identity reaching a sensitive store enumerated, and every unverified edge labelled. Time-box the rest.

open as a page