skip to content

Data-Flow Diagrams and Trust Boundaries

Decomposing a system into a data-flow diagram and drawing the trust boundaries where privilege changes, then reading the attack surface off it. Live whiteboard exercises almost always start here.

on this pageshow

explore

questions

29

What does a level-0 context diagram show in a data-flow-diagram threat model?

level: juniorimportance: must knowfreq 70%

answer

  1. start by drawing nothing inside
  2. who talks to it, not how it works
  3. one process, external entities, flows
  4. the outer boundary and what crosses it
  5. level 0 of the leveling hierarchy

basics

~20 s

A level-0 context diagram draws the whole system as one process, every external entity that talks to it, and the flows between them. It fixes scope and shows what crosses the outer trust boundary, with no internal structure.

solid answer

~50 s

The context diagram sits at the top of the leveling hierarchy: a single process for the entire system under review, every external entity around it — end users, operators, partner systems, identity providers, batch feeds — and every flow in and out, annotated with protocol and data class. The one boundary it always carries is the outer one, so everything crossing it is untrusted input coming in or a disclosure going out. For a national driving-licence renewal service, level 0 is what a ministerial sign-off actually reads: applicants, the photo-and-identity checking service, the payment provider, the police records feed, the print bureau. It already yields real threats — an anonymous applicant spoofing another citizen, tampering with a submitted document in transit, repudiating a renewal they claim they never requested. What it cannot show is elevation of privilege between two internal components; that needs a lower level.

go deeper

for a junior

Be ready to draw it on request: one box for the system, external entities around it, labelled arrows, and a line for the outer trust boundary. Say out loud that nothing internal is shown.

for a middle

Explain what the level buys you — a scope everyone can dispute, and the outside-attacker entry points enumerated before any internal structure exists. Name the interfaces people forget: admin, batch, monitoring.

for a senior

Show that you can produce real threats from level 0 alone and can state its limit: no internal privilege escalation, no component-to-component boundary. Be able to say which pending decision would force you to decompose.

for a principal

Own the artifact politics: the context diagram is the page a sign-off actually reads and the one that stays true across releases. Decide who maintains it and what change to the external interface set triggers a re-review.

## The leveling hierarchy A data-flow diagram (DFD) threat model is not one drawing. It is a small hierarchy of drawings at increasing detail, and the convention borrowed from structured analysis names them by level. - **Level 0, the context diagram.** The entire system under review is drawn as a *single process*. Around it sit the external entities it exchanges data with, and between them the flows. Nothing inside the system is shown. - **Level 1.** That single process is opened up into its major internal processes and data stores, with the same external flows still arriving at the edge. - **Level 2 and below.** One level-1 process is opened up in turn, showing its internals. The context diagram is therefore the answer to a scoping conversation that has already happened: *this* is the thing we are modeling, and *these* are the parties it talks to. ## What belongs on it Four things and no more: 1. **One process** — the system, named as the business names it. 2. **External entities** — anything that interacts with the system but is not part of it, and that you do not control or model internally: citizens and customers, administrators and operators, partner services, identity providers, payment providers, regulators pulling reports, batch feeds, the monitoring backend you ship logs to. 3. **Flows** — every data path in and out, with a direction, a protocol and a data class (personal data, credentials, money instructions, telemetry). A flow with no annotation is half a flow. 4. **The outer trust boundary** — a line separating the system from everything around it. Every flow that crosses it is an entry point (untrusted input) or an exit point (a disclosure or an action taken on the outside world). What does **not** belong: internal components, internal databases, deployment detail such as hosts, subnets and load balancers, and message ordering. Those are the business of lower levels or of other notations entirely. ## Worked example A national driving-licence renewal service. The level-0 diagram has one box labelled *Licence Renewal Service*, and around it: the applicant (submits an application, a photo and a scanned proof of address; receives a status), the photo-and-identity verification provider, the payment provider, the police records feed that supplies disqualification data, the print bureau that receives an authorised production instruction, and the ministry's reporting users. That single page is what a non-technical sign-off reads and approves, and it is enough to enumerate a first tier of threats: | Flow crossing the boundary | Representative threat | | --- | --- | | Application submitted by applicant | Anonymous applicant impersonates another citizen (spoofing) | | Scanned document uploaded | Document altered before or during submission (tampering) | | Renewal confirmation | Applicant later denies having requested the renewal (repudiation) | | Instruction to the print bureau | An unauthorised instruction produces a genuine document (integrity of the physical artifact) | | Records feed from police | Availability loss stops all renewals (denial of service) | Notice the asset in play: personal data and the integrity of an official document, not a generic customer database. The context diagram is where you name that asset, because it is where you name who is on the other side of every line. ## Why the level exists at all Three jobs. **It fixes the boundary of the exercise in a form people can disagree with.** A one-page diagram makes a missing interface visible: someone in the room says *what about the call-centre agents who can amend an application?*, and you have found an entry point before you have modeled a single internal component. **It gives you the outside-attacker view for free.** The threats above are the ones an anonymous or lightly authenticated outsider can reach, and they are usually the highest-exposure ones. Starting inside the system, at level 1, tends to bury them. **It is the stable artifact.** Internals churn every sprint; the set of parties the service talks to changes rarely. A context diagram stays true for a long time, which is why it is the page you put in front of a sign-off. ## Common failures - Drawing internal components on it, so that it becomes a bad level-1 diagram. - Omitting the boring external entities — the monitoring egress, the nightly reconciliation feed, the admin console reached from an office network — because they feel internal. They are entry points, and they are frequently the least protected. - Treating it as too shallow to yield threats and skipping straight to level 1. - Confusing it with a deployment diagram. The context diagram is about *who exchanges what data with the system*, not about where anything runs. ## What it cannot tell you Anything that depends on internal structure: privilege escalation from one internal component to another, an internal store read by a component that should not read it, a boundary between two teams' services. If the pending question is one of those, you decompose. The context diagram is the start of the hierarchy, not the whole of it.

  • Do trust boundaries appear on a context diagram, or only at lower levels?
    At least one appears: the outer boundary around the single system process. Every flow crossing it is an entry point or an outward disclosure, which is exactly the analysis the level-0 diagram is for. Boundaries *inside* the system — between components under different credentials, accounts or teams — only become drawable once you decompose, and finding one is itself a reason to go a level deeper.
  • How do you know a context diagram is complete?
    Walk the interfaces rather than the architecture: who authenticates, who administers, who is billed, who receives exports, what runs on a schedule, where telemetry and backups go, and which third parties are called out. Each answer is an external entity or a flow. Any party a person can name that is not on the page is a gap, and the ones people forget are admin, batch and monitoring paths.
  • Can a context diagram be too big?
    It can, and the fix is usually scope rather than layout. Thirty external entities normally means you have drawn a whole department as one system. Either the thing under review is genuinely a platform — in which case group entities by role and accept the page — or you should be modeling one service inside it, with the neighbours becoming externals.

It is the diagram on the front of an embassy building: who may approach, which windows they queue at, and what they hand over. It says nothing about the corridors inside.

saying these in an interview costs you the question

  • Draws internal components and databases on the context diagram
  • Says a context diagram is too shallow to produce threats
  • Omits admin, batch and monitoring interfaces as internal
  • Confuses the context diagram with a deployment or network map
  • Leaves flows unlabelled, with no direction or data class

context

open as a page

On a threat-modeling data-flow diagram, what do process, data store, external entity and flow each represent?

level: juniorimportance: must knowfreq 76%

basics

~20 s

A process is running code that transforms data, a data store is passive data at rest, an external entity is a person or system you do not control, and a flow is data in motion between them.

open as a page

When you enumerate entry points on a data-flow diagram, what counts as one?

level: juniorimportance: must knowfreq 68%

basics

~20 s

An entry point is anywhere data or a request crosses into the system from outside its trust boundary. Not just HTTP endpoints: queue consumers, file drops, inbound email, webhooks, scheduled jobs that fetch remote data, and admin or support consoles all count.

open as a page

On a data-flow diagram, what does a trust boundary mark, and what decides where it is drawn?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A trust boundary marks a data flow where the two sides do not trust each other equally. Draw it wherever the principal, privilege, tenant, code owner or execution context changes - not wherever the network or a firewall changes.

open as a page

What does a data-flow diagram deliberately omit, and which threats hide in those omissions?

level: middleimportance: must knowfreq 62%

basics

~20 s

A data-flow diagram shows what data moves where, not when or in what order. It omits control flow, sequencing, retries and error paths, so replay, race, double-processing and failure-path leak threats stay invisible on it.

open as a page

A data-flow diagram names every process and store but draws no trust boundaries — what do you conclude?

level: middleimportance: must knowfreq 72%

basics

~20 s

Conclude that threat enumeration has not started. A trust boundary marks where the privilege of whoever handles the data changes; with none drawn, every flow on the page looks equally trusted and there is nothing to analyse.

open as a page

Which sources do you rebuild a data-flow diagram from for an undocumented legacy service?

level: middleimportance: must knowfreq 62%

basics

~20 s

Rebuild the diagram from artifacts that describe the running system: infrastructure-as-code and its state, route and proxy configuration, outbound calls grepped from the repository, identity bindings and database grants, and request telemetry. Confirm each edge with an operator.

open as a page

Why annotate every data-flow diagram element with the principal and privilege it runs as?

level: middleimportance: must knowfreq 62%

basics

~20 s

Writing the principal an element acts as, and the privilege it actually holds, makes every privilege change visible on a data-flow diagram. Without those labels, boundary placement is guesswork and high privilege hides in plain sight.

open as a page

Why is tenancy a trust boundary on a data-flow diagram even when both tenants share one subnet?

level: middleimportance: must knowfreq 72%

basics

~20 s

A trust boundary marks where trust or privilege changes, not where the network changes. Two tenants sharing one subnet still have different owners and different rights, so anything crossing from one to the other crosses a boundary.

open as a page

A threat model's boxes are 'Team Alpha', 'Team Bravo' and 'Platform' — what is wrong with it?

level: juniorimportance: should knowfreq 40%

basics

~20 s

The boxes are organisational units, not system elements. A data-flow diagram is decomposed by where data moves — external entities, processes, stores, flows — so an org chart leaves an adversary nowhere to stand and no boundary to draw.

open as a page

Why does a data flow on a threat-modeling DFD need protocol, direction and data class on it?

level: middleimportance: should knowfreq 54%

basics

~20 s

An unlabelled arrow hides the three things the analysis needs: what channel protects the data, which way the data actually moves, and how valuable it is. Without those, every flow looks identical and threats collapse into guesswork.

open as a page

Why enumerate a system's exit points, not just its entry points, when threat modeling?

level: middleimportance: should knowfreq 45%

basics

~20 s

Exit points are where data leaves the trust boundary: logs, exports, backups, restore copies, crash dumps and telemetry. They carry the asset outward without anyone breaking in, and they usually skip the authorization the primary request path enforces.

open as a page

In a single-page trading app whose trader-vs-supervisor route guard ships in downloaded JavaScript, where does the trust boundary actually belong?

level: middleimportance: should knowfreq 58%

basics

~20 s

The boundary belongs at the server API, not inside the browser. Everything you ship - route guards, hidden buttons, role flags - runs on a machine the user controls, so both roles sit in one untrusted zone.

open as a page

When does a deployment diagram reveal a threat that a component diagram cannot?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A component diagram shows logical separation; a deployment diagram shows where code actually runs. Co-location threats, such as two supposedly separate services sharing one process, one host and one credentials file, appear only in the deployment view.

open as a page

How do you decide how deep to decompose a data-flow diagram for a threat model?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Decompose only where more detail would change a decision. Stop when the next level would produce elements you would mitigate identically, when no boundary hides inside the box, or when its owners will model it themselves.

open as a page

A 200-element data-flow diagram and a two-box 'monolith to database' one both fail review — why?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Both defeat the walk. The 200-element page spreads a reviewer's attention evenly so boundary crossings vanish into noise; the two-box page hides every privilege change inside a box, so each threat comes out as an unactionable 'tamper with the monolith'.

open as a page

Your entry-point list is ranked by internet exposure — what does that ranking miss?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Exposure is only one axis. The others are privilege — what the path can do once used — and attribution: whether a use binds to a person and would be noticed. A rarely used break-glass admin route can outrank a public form.

open as a page

Your recovered data-flow diagram has an edge telemetry proves but no configuration explains — what now?

level: seniorimportance: should knowfreq 44%

basics

~20 s

An observed flow is real evidence, so the configuration set is what is incomplete. Characterise the edge from telemetry — endpoints, identity, protocol, volume, schedule — then walk it with the on-call operator, keeping it drawn until someone accounts for it.

open as a page

A support console calls the accounts API under its own service identity - what does the threat model lose?

level: seniorimportance: should knowfreq 52%

basics

~10 s

The agent's identity dies at that boundary. Downstream sees one caller holding the union of every agent's access, cannot make any per-agent decision, and records one name for every read, so attribution is lost.

open as a page

A payroll bureau's own staff can read your employee records - where does the trust boundary go?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Draw the boundary around granted access, not around the company. Bureau operators who can read every payroll record sit inside the zone that handles employee data, so operator misuse is an in-scope threat your model has to name.

open as a page

On a managed database service, which threats stay on your side of the shared-responsibility line?

level: seniorimportance: should knowfreq 58%

basics

~20 s

The provider owns the host, the patching, the storage and the backup machinery; you own the schema, the credentials and every access rule. A cross-tenant read caused by a missing tenant filter is your threat, not theirs.

open as a page

In DFD leveling, what must a child diagram preserve from the process it decomposes?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

The child diagram must show the same flows in and out as the parent process it expands — same sources, same sinks, none added or dropped. Only internals are new, and numbering carries down: process 3 expands into 3.1, 3.2.

open as a page

When is a single process box on a data-flow diagram actually hiding several processes?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

Whenever the box covers code running at different privileges, reached by different callers, or handling different data classes. One process shape asserts one identity and one set of controls, so anything inside it that differs disappears.

open as a page

Your pricing engine loads another team's plugin into its own process - where does the trust boundary go, and can one exist inside a process?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Trust changes at the code-ownership seam: foreign code inside your process shares your address space, secrets and privileges. A boundary belongs there, but nothing enforces it in-process - so mark it unenforced, or move the plugin into a separately privileged process.

open as a page

What does mandating a data-flow diagram for every design review cost you?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Standardising on one diagram makes reviews comparable and the practice teachable, but it silently caps what teams can find. Protocol, retry, error-path and co-location threats do not fit the notation, so nobody raises them and the gap reads as safety.

open as a page

A data-flow diagram predates a move to three cloud regions and a mobile client — what can you still conclude?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Very little about security, and nothing by omission. The data classes and business flows probably still hold; every claim the page makes about boundaries, surface and the adversary set has expired, and absence from it proves nothing.

open as a page

How complete must a recovered data-flow diagram be before you start enumerating threats on it?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

A recovered diagram is ready when it answers the decision at hand: every trust-boundary crossing drawn, every data store's class named, every identity reaching a sensitive store enumerated, and every unverified edge labelled. Time-box the rest.

open as a page

A reviewed threat model annotates every component with the same operator credential - what do you conclude?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

A uniform principal is a finding, not a tidy diagram. With one credential everywhere, no edge can show a privilege change, so the boundaries the system needs have been erased by its identity design rather than proven absent.

open as a page

In a white-label banking app, each partner assumes the other runs fraud checks - how do you close that gap?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Neither model covers the seam, because each stops at its own edge and draws the missing control on the far side. Close it by modeling the shared boundary jointly and naming exactly one owner for every control on it.

open as a page