On a threat-modeling data-flow diagram, what do process, data store, external entity and flow each represent?
answer
- four shapes, four different assumptions
- one of them is passive by definition
- one you cannot control from inside
- only one of them runs as an identity
- arrows are data in motion, not links
basics
~20 sA process is running code that transforms data, a data store is passive data at rest, an external entity is a person or system you do not control, and a flow is data in motion between them.
solid answer
~50 sA data-flow diagram gives threat modeling four building blocks. A **process** is code that is actually running and doing something to data: a service, a job, a function. A **data store** is passive: a database, a bucket, a queue, a log file, a config file. An **external entity** is a person or system that sits outside what you build and operate, so you cannot put a control inside it. A **flow** is data in motion between two of the above, and it is the thing that actually crosses boundaries. The classification is not cosmetic: it decides which threats you enumerate and where a control can even be placed. On a podcast-hosting platform, the listener app, the delivery network and the ad exchange are external entities; the origin bucket and the transcode cache are data stores; ingest and packaging are processes; the arrows between them are flows.
go deeper
Be ready to name the four element types and give one concrete example of each from a system you have worked on. Interviewers at this level mostly want to see that you would not draw a free-form architecture sketch and call it a threat model.
Expect to explain the tests that separate the types: passive versus executing, and inside versus outside your control. Be able to defend a genuinely ambiguous box, such as a managed database or another team's service.
Show that classification drives the analysis. Say what each type commits you to asking and where a control can actually be placed, and be ready to classify an unfamiliar design live on a whiteboard without stalling on the ambiguous elements.
Own the consistency question. Decide what your organisation treats as external by default, whether managed services are drawn as stores or entities, and how much notation discipline is worth imposing before it costs you engagement from delivery teams.
## What the notation is for A data-flow diagram (DFD) in threat modeling is not architecture art. It is a deliberately small vocabulary chosen so that a room full of people can agree on what the system is before they argue about what can go wrong with it. Its power comes from the fact that there are only a few shapes, and each shape carries a different set of assumptions about who is running the thing, what privileges it holds, and what you are allowed to do about it. ## The four element types **Process.** Running code that receives data, does something to it, and emits data. A web service, a background worker, a scheduled job, a serverless function, a daemon on a device. Conventionally drawn as a circle (or a rounded rectangle). A process is the only element that has a *privilege*: it runs as some identity, with some set of permissions, on some host. That makes it the richest source of threats, because it can be impersonated, subverted, made to leak, or made to do work on an attacker's behalf. **Data store.** Data at rest. A relational database, an object bucket, a message queue, a cache, a log, a file on disk, a configuration store. Conventionally drawn as two parallel horizontal lines. The defining property is that it is **passive**: it does not act on its own. This is the test people get wrong most often. If the box has triggers, stored procedures, a replication agent, or a scanner that reaches out to other systems, then something there is executing code, and that part is a process, not a store. **External entity.** A person or a system that interacts with your system but that you do not build, run or control: an end user, an operator, a partner's API, a third-party service you call, a device in someone else's hands. Conventionally drawn as a rectangle. The defining property is that you **cannot place a control inside it**. You cannot patch it, you cannot audit its code, you cannot make it validate anything on your behalf. Everything you do about an external entity happens at your edge of the connection. That is also why every flow to or from an external entity crosses a boundary by definition — the entity is, by construction, on the other side of one. **Data flow.** Data moving from one element to another, drawn as an arrow. A flow is not a function call and not a network link; it is *data in motion*, and it is what actually leaves your control and travels somewhere. Flows carry the payload that matters — the credentials, the salary figures, the media file, the firmware image — which is why the annotation on a flow (its protocol, its direction, what class of data it carries) does more analytic work than the arrow itself. Most teams also draw **trust boundaries** as dashed lines over the top of the diagram. That is a mark on the diagram rather than a fifth kind of thing in the system: it says the elements on either side are run, owned or trusted differently. ## Why the classification changes the analysis The element type is a shorthand for what you may assume and what you may do: - Something you classify as a **process** invites questions about the identity it runs as and what happens when it is subverted. - Something you classify as a **store** invites questions about who can read and write it directly, bypassing the process that is supposed to guard it. - Something you classify as an **external entity** ends the conversation about fixing it internally and starts a conversation about what you validate, authenticate and log at your own edge. - Something you classify as a **flow** invites questions about the channel: who can see it, alter it, replay it, or stop it. Misclassify, and the whole enumeration skews. Calling an in-house recommendation service an external entity because a different team owns it quietly excuses you from asking what it runs as. Calling a queue a process hides that anyone with the right credentials can publish into it directly. ## A worked classification Take a podcast-hosting platform. The listener application on a stranger's phone is an **external entity**: you ship it, but you do not run it and you cannot trust anything it asserts. The content delivery network in front of your origin is an **external entity** too — a third party you configure but do not operate. The advertising exchange your packager calls for a mid-roll spot is an **external entity**. The origin object bucket holding master audio files and the transcode cache holding derived renditions are **data stores**. The upload ingest service, the transcoder and the manifest packager are **processes**. Everything between them is a **flow**. Drawn that way, the interesting question surfaces on its own: the master audio in the origin bucket is the asset, an anonymous listener is the cheapest adversary in the world, and the path from that listener through a third party you do not run to a store full of unwatermarked masters is now three labelled hops instead of one hand-wave. ## Getting it right in an interview Do not recite shapes. Say what the shape *commits you to*: a process has a privilege, a store is passive, an external entity is beyond your reach, a flow is the thing that leaves. Then classify the system on the whiteboard in front of you and defend one hard call.
- A team argues a database is a process because it has stored procedures and a replication agent. Are they right?Partly, and it is worth splitting. The stored data is a data store; the stored procedures and the replication agent are code that executes and reaches out, so they behave as processes. Draw them separately when the distinction changes the analysis, for example when replication ships data across a boundary the store's own access controls never see. If the procedures are trivial lookups, leave it as one store and note the simplification.
- Another team is inside your company but not on your team. External entity or process?Ask whether you can place a control inside it and whether you can see how it behaves. If you can require them to change code, read their configuration and audit their identity, model it as a process inside your scope. If in practice you can only defend your own edge of the connection, treat it as an external entity and be explicit that you are choosing to distrust an internal system.
- Why is a flow drawn as its own element rather than just a line connecting two boxes?Because the flow is where the asset actually travels and where most of the interesting exposure lives. Two boxes can be perfectly hardened while the data between them moves in the clear, goes to the wrong place, or carries far more than the receiver needs. Treating the arrow as a first-class element forces someone to say what is on it and over what channel.
Think of a warehouse: workers are processes, shelves are data stores, the courier who drops off a pallet is an external entity, and the pallets moving around are the flows.
saying these in an interview costs you the question
- Calls every box a process because everything is software
- Treats an external entity as something you can harden
- Draws a queue or cache as a process rather than a store
- Says the shapes are cosmetic and only boundaries matter
- Describes flows as network links rather than data in motion