Why does a data flow on a threat-modeling DFD need protocol, direction and data class on it?
answer
- an unlabelled arrow asserts almost nothing
- arrows follow the data, not the caller
- request and response are not one edge
- the channel may already solve two threats
- what is on the wire sets the stakes
basics
~20 sAn unlabelled arrow hides the three things the analysis needs: what channel protects the data, which way the data actually moves, and how valuable it is. Without those, every flow looks identical and threats collapse into guesswork.
solid answer
~50 sAn arrow with no label asserts only that two boxes talk. The **protocol** tells you what the channel gives you for free and what it does not, so you know whether tampering and disclosure on that hop are already handled. The **direction** is about data movement, not who calls whom, which is why a request and its response are two flows carrying completely different data. The **data class** tells you what is actually at stake on that hop, which is what separates a flow worth spending a sprint on from one that is not. A school photo-ordering site once drew a single arrow labelled `notifications` between the site and parents. It was two flows in opposite directions: outbound mail carrying pupil names, and an inbound bounce and complaint webhook the site parsed and trusted. One arrow, one label, and the inbound untrusted-input flow was invisible until someone annotated the edge.
go deeper
Be ready to say that an arrow needs a label and to name the three things worth writing on it. Knowing that the arrow follows the data rather than the caller already puts you ahead of most candidates at this level.
Expect to explain why a request and its response are two flows with different threats, and what a named protocol does and does not guarantee. Practise annotating a small design out loud.
Demonstrate that annotation drives prioritisation: use the data class to argue which of several flows deserves the sprint, and spot an inbound flow from a trusted partner that your code parses without validation.
Own the annotation standard. Decide the minimum every team must label, in vocabulary their engineers already use, and defend why more detail than that costs you diagram freshness without buying analysis.
## The arrow is the cheapest place to lie to yourself Boxes get scrutiny in design reviews. Arrows get drawn last, fast, and unlabelled. But the arrow is where the asset actually travels, where it leaves one operator's control and enters another's, and where the majority of the interesting exposure lives. Three annotations turn a decorative line into an analysable one. ## Protocol: what the channel already does for you Writing the protocol on the flow answers a question you would otherwise assume: does this hop already have confidentiality, integrity and peer authentication, or not? A flow marked as mutually authenticated TLS is a different threat surface from the same arrow carrying an unauthenticated plaintext file transfer, a message on an internal queue with no transport encryption, or a signed webhook over the public internet. The protocol annotation also exposes asymmetries that architecture diagrams smooth over. A hop can be encrypted in transit but terminate at an intermediary that sees plaintext. A file drop can be encrypted at the transport layer while the file itself is readable by anyone with access to the destination directory. Naming the protocol on the arrow is what makes someone ask which of those is true. ## Direction: data movement, not call direction This is the annotation people get backwards. A DFD arrow shows the direction the **data** moves, not the direction the **call** goes. A service reading a file from a bucket initiates the call, but the file flows from the store to the service, so the arrow points at the service. The convention matters because the two directions of an exchange carry different payloads, cross the boundary in different directions, and attract different threats: - Outbound from you: exposure of data you hold. Who can see it, where does it end up, is it more than the receiver needs? - Inbound to you: untrusted input. Who can inject into it, what does your code do with it, what happens if it is malformed or attacker-chosen? That is why a two-way exchange is normally drawn as two arrows rather than one double-headed one. A single bidirectional arrow silently merges an exposure question and an input-validation question into one line that answers neither. ## Data class: what is at stake on this hop The data class is the annotation that lets you prioritise. Write what is genuinely on the wire: session tokens, salary amounts, national identifiers, pupil names, unwatermarked master files, firmware images, aggregate counters. Two arrows between the same two boxes can differ by orders of magnitude in what a compromise costs, and only the data class shows it. It also catches over-collection. A flow annotated as carrying a full employee record when the receiver needs an account number and an amount is a design defect visible on the diagram, before anyone enumerates a single threat. ## A worked example: the monthly payroll run Draw a payroll run as elements. An HR data store holds employee records. A payment-file process reads from it, builds a bank payment file, and drops that file to the bank over an outbound file transfer. Unannotated, this is three shapes and two arrows, and it looks boring. Annotate the edges: | Flow | Direction | Protocol | Data class | | --- | --- | --- | --- | | HR store to payment-file process | store to process | database session over the internal network | full employee records: names, bank details, salary amounts | | Payment-file process to bank | outbound, to a third party | scheduled file transfer over SFTP | payment instructions: account numbers and amounts, signed or not | | Bank to payment-file process | inbound acknowledgement | file fetched from the bank drop | settlement results parsed and trusted by your process | Now the analysis writes itself, and note who the adversary is: not an anonymous internet attacker, but the payroll clerk who is *supposed* to operate this run. The first flow says the process pulls far more of the record than a payment file needs. The second says the money instructions are protected by the transport only, so whether the file itself is signed decides whether an operator who can reach the drop directory can add a line to it. The third says an inbound file from outside is parsed by your code, which is untrusted input arriving on a path nobody thinks of as an input path because the bank is a trusted party. None of that was visible when the arrows were unlabelled. All of it is visible in one table. ## How much annotation is enough Enough that the diagram answers the questions you would otherwise have to ask a person. In practice: protocol and whether it authenticates both ends, direction as separate arrows when the two directions differ, and a data class named in the domain's own words rather than a classification tier nobody can decode. Annotating every arrow with a full schema is waste; annotating none of them is why the diagram gets ignored.
- When is a single double-headed arrow acceptable rather than two separate flows?When both directions carry the same data class, cross the same boundary, and sit on the same channel, so the analysis of one is genuinely the analysis of the other. A symmetric heartbeat qualifies. A request and its response almost never do, because one carries a query and the other carries the records, and the inbound one is untrusted input while the outbound one is exposure.
- The protocol on a flow is TLS. Does that close out tampering and disclosure on that hop?Only in transit, and only if both ends are authenticated. It says nothing about who terminates the connection, what an intermediary sees, whether the payload is signed independently of the channel, or whether the data is protected once it lands. State what the protocol gives you and what it leaves open, rather than treating the label as a mitigation.
- How do you keep flow annotations from rotting as the system changes?Keep them coarse enough to stay true. Protocol, direction and a domain-language data class change far more slowly than payload schemas, so they survive normal churn. Tie a refresh to the events that actually invalidate them, such as a new integration, a new data class entering the system, or a change of protocol, rather than promising a periodic sweep nobody will do.
An unlabelled arrow is a shipping route with no manifest: you know two ports are connected, but not what is in the container, which way it sailed, or whether the box was sealed.
saying these in an interview costs you the question
- Points the arrow at whoever initiates the call
- Draws every exchange as one double-headed arrow
- Labels flows with transport only and stops there
- Assumes an inbound flow from a partner is trusted input
- Uses classification tiers nobody in the room can decode