skip to content

What does a data-flow diagram deliberately omit, and which threats hide in those omissions?

level: middleimportance: must knowfreq 62%

answer

  1. structure, not sequence
  2. arrows carry data, not time
  3. what happens on the second attempt
  4. the unhappy path is never drawn
  5. dead-letter stores and crash dumps

basics

~20 s

A data-flow diagram shows what data moves where, not when or in what order. It omits control flow, sequencing, retries and error paths, so replay, race, double-processing and failure-path leak threats stay invisible on it.

solid answer

~50 s

A data-flow diagram is a static picture of processes, data stores, external entities and the flows between them, cut by trust boundaries. That projection is chosen so you can walk every element and ask what could go wrong with it, and it deliberately throws away four things: control flow and ordering, retry and delivery semantics, error and fallback paths, and time or state. So a whole family of threats has nowhere to appear: replay of a captured message, a check-then-use race, at-least-once double-processing, a state step reached out of order, and raw payloads spilling into failure sinks nobody drew. In an event-driven order pipeline, a consumer that dead-letters failed messages into an object bucket puts customer personal data somewhere the diagram never showed. The fix is not to abandon the DFD but to walk each flow's unhappy path and draw a message-ordering view where the answer is interesting.

go deeper

for a junior

Be ready to say what the four DFD element types are and that flows show data movement, not order of execution. Knowing the diagram is drawn from the happy path is enough at this level.

for a middle

An interviewer expects you to name the omissions concretely, ordering, retries, error paths and state, and to pair at least two of them with a threat that becomes invisible as a result. This is the tier where the mechanics of the notation are yours to explain.

for a senior

Show that you actively hunt the gaps: walk each flow asking what happens on failure, on repeat and out of order, then pull the failure sinks you find back onto the diagram. Talk about a real spill into a dead-letter or log store you had to model.

for a principal

Own the framing that every diagram is a projection and the notation caps the threat surface a team can report. Be ready to say when a second view is worth its maintenance cost and when the honest answer is a one-off sketch that gets thrown away.

## What a data-flow diagram is actually for A data-flow diagram (DFD) is a static picture of where data lives and where it travels: **processes** that transform data, **data stores** that hold it, **external entities** outside your control, and **flows** connecting them, cut by **trust boundaries** where the level of trust changes. That projection is chosen deliberately. Element-by-element enumeration methods work directly off this element list: STRIDE walks each element asking about spoofing (violating authentication), tampering (violating integrity), repudiation (violating non-repudiation), information disclosure (violating confidentiality), denial of service (violating availability) and elevation of privilege (violating authorization). A DFD is the notation that makes that walk mechanical. Every diagram is a projection, and a projection is defined as much by what it discards as by what it keeps. ## The four families a DFD discards 1. **Control flow and ordering.** An arrow means data moves in that direction. It does not mean this call happens first, or once, or only after that other call. There is nowhere on a DFD to say what state the process was in when the flow arrived. 2. **Retries, timeouts and delivery semantics.** Whether a flow is at-most-once, at-least-once or exactly-once is not expressible. Neither is what the caller does when it times out. 3. **Error, exception and fallback paths.** DFDs are drawn from the happy path. The branch taken when validation fails, when a dependency is down, or when deserialization throws, usually never gets drawn. 4. **Time, state and volume.** Session lifetimes, token expiry, rate, and how long data lingers in a store are all outside the notation. ## The threats that hide in the gaps - **Replay.** A message captured and re-sent is accepted again, because nothing on the diagram says a flow may only be honoured once. - **Time-of-check to time-of-use races.** The check and the use are two flows on the diagram with no ordering relationship drawn between them. - **Double-processing.** At-least-once delivery plus a non-idempotent handler produces a double credit or a double shipment. The DFD shows one arrow either way. - **State-machine skipping.** A step reached without its prerequisite, because the diagram never claimed a prerequisite existed. - **Failure-path spill.** Raw payloads written into a dead-letter store, a log line or a crash dump that appear on no diagram and in no data inventory. - **Insecure fallback.** A downgrade to a weaker path when the primary one fails, invisible because the failure branch was never drawn. ### A worked case An event-driven order pipeline models cleanly: a checkout process, an order service, an order store, one boundary at the internet edge. The consumer cannot deserialize a malformed message, so the platform dead-letters the raw message into an object bucket for operators to inspect. That bucket is on no diagram. The raw message carries the customer's name, delivery address and contact details, and every operator with read access to that bucket now holds a copy that never entered the data inventory or the retention policy. The adversary here is an insider with entirely legitimate access; the asset is customer personal data. No amount of staring at the DFD surfaces this, because the notation has no place for error handling. Walking the failure path of each flow surfaces it immediately, and the remedy is to draw the failure sink back onto the DFD as a real data store with its own boundary crossing. ## What to do about it Treat the DFD as the entry point, not the whole model. For each flow, ask three unhappy-path questions: what happens if it fails, what happens if it is repeated, and what happens if it arrives out of order. When any answer is interesting, draw a message-ordering view of that one flow, showing each message in sequence with its freshness or idempotency protection annotated, and enumerate against that instead. Then draw whatever failure sinks you discovered back onto the DFD, because that is what the DFD is good at: showing where data comes to rest and which boundaries it crosses. The mistake to avoid is the opposite over-correction, replacing the DFD with a behavioural view. An ordering view of one interaction tells you nothing about the eleven data stores elsewhere in the system or where your boundaries are. The two views answer different questions, and a competent model uses the cheapest view that makes the threat you are hunting visible.

  • A wallet top-up double-credits the balance whenever the client retries after a timeout. Why would the DFD never show that, and what do you draw instead?
    The DFD shows one flow from client to wallet service to balance store, with no way to express that the flow can happen twice for one user intent. Draw a message-ordering view of the top-up: request, timeout, retry, and the second credit. That view makes the missing idempotency key and the at-least-once semantics visible, and it names the adversary precisely, an authenticated low-privilege user who learns to force the timeout, with money as the asset.
  • Does needing a second view mean the DFD was the wrong diagram to start with?
    No. The DFD did its job: it inventoried the data stores, the external entities and the boundary crossings, which is exactly what an ordering view cannot do. Needing a supplement is normal, because every diagram is a projection. The judgment being tested is whether you notice which threats your projection cannot express, and reach for a second view only for the specific flows where that matters.
  • How can you make error paths visible without abandoning the DFD notation?
    Draw them as first-class elements. A dead-letter queue, an error bucket, a crash-dump location and a verbose log sink are all data stores, so give them a box, a flow into them, and a boundary crossing if the readers differ from the readers of the primary store. Annotating each flow with what happens on failure turns an implicit branch into something the element walk will pick up.

A DFD is a plumbing schematic: it shows every pipe and tank, but nothing about water hammer, what happens when a valve sticks, or where the overflow goes.

saying these in an interview costs you the question

  • Says a DFD shows the order in which operations happen
  • Reads arrows on a DFD as call sequence or request-response
  • Threat-models only the happy path of each flow
  • Thinks adding timestamps to a DFD fixes ordering threats
  • Does not treat a dead-letter queue or log sink as a data store
  • Proposes replacing the DFD with a sequence view entirely

context