When is a single process box on a data-flow diagram actually hiding several processes?
answer
- a box is a claim about uniformity
- one identity, one set of callers
- product names are not processes
- the notation has a marker for not-atomic
- internal edges become the analysis
basics
~20 sWhenever the box covers code running at different privileges, reached by different callers, or handling different data classes. One process shape asserts one identity and one set of controls, so anything inside it that differs disappears.
solid answer
~50 sA process shape is a claim: one thing, one identity, one set of controls, no internal edges worth drawing. The claim breaks when the box is really a product name. A warehouse robot fleet has a `fleet controller` that is genuinely three programs: a route scheduler, an over-the-air firmware updater and a telemetry sink. They differ in privilege — only the updater can sign and ship code to machines that move — and in who reaches them. Drawn as one circle, the flows between them vanish, and so does the question of whether the telemetry sink, which anything on the floor can talk to, shares an identity with the updater. Classic notation has a marker for this: a complex process, conventionally two concentric circles, meaning this box is not atomic and has its own diagram beneath it. Use it, or split the box.
go deeper
Be ready to say that one process shape means one running thing with one identity, and to notice when a box is really a product or team name covering several programs.
Expect to name the signals that force a split: different privilege, different callers, different data classes, different operators. Be able to point at one on a design you have worked on.
Show the judgment live. Interrogate a plausible-sounding box, split it where privilege differs, and articulate the threats on the newly visible internal flows rather than just redrawing shapes.
Own where the organisation stops. Set the rule for when a component may stay whole, keep unexamined internals visibly marked rather than silently assumed clean, and defend that against both over-splitting and diagrams that flatter the architecture.
## What a single process box promises When you draw one circle and give it a name, you are asserting several things at once: this runs as one identity, its parts trust each other, there are no edges inside it that carry data anyone would want, and a control placed on it covers everything it contains. Those assertions are usually fine. When they are false, they are false invisibly, which is the worst way for an assumption to be wrong. ## The signals that a box is lying Four signals, in rough order of how often they matter. **Different privilege.** Parts of the box hold materially different power. One part can only read a queue; another can sign artefacts, write to production, or issue credentials. Merging them on the diagram means the highest privilege inside the box quietly becomes the privilege of the whole box, and nobody asks whether the low-privilege part could reach the high-privilege one. **Different exposure.** Different parts are reachable by different callers. If one component of the box accepts connections from anything on a local network and another is only reachable from an internal admin path, they do not share a threat profile, and the diagram should not imply they do. **Different data classes.** One part handles bulk low-value telemetry, another handles signing keys or personal data. The asset at stake differs, so the cost of a compromise differs. **Different lifecycle and operator.** One part is deployed weekly by a product team, another is a vendor component updated rarely by whoever remembers. Different change rates mean different assumptions rot at different speeds. A useful heuristic: if you cannot describe the box's *single* identity and *single* set of inbound callers in one sentence, it is more than one process. ## A worked case: the fleet controller A warehouse runs a fleet of autonomous robots, and the design shows one box called `fleet controller`. Ask what runs inside it and you get three answers: - a **route scheduler** that computes paths and sends movement instructions to robots, - an **over-the-air firmware updater** that builds, signs and pushes firmware images to those robots, - a **telemetry sink** that ingests position, battery and fault streams from every unit on the floor. These fail every test above. The updater holds signing material and can change what physical machines do; the scheduler can move machines but not change their code; the sink holds nothing precious and accepts input from every device on the floor, including any device an attacker has managed to get onto the network. Yet the box implies they share an identity and a blast radius. Redraw it as three processes with the flows between them made explicit, and the interesting questions appear immediately. Does the telemetry sink write into a store the scheduler reads, so that fabricated telemetry can steer routing decisions? Does the updater pull build artefacts from a pipeline whose inputs include third-party dependencies, so that a compromised dependency becomes signed firmware on a moving machine? Can something that reaches the sink reach the updater's signing path because they run under one identity on one host? Note what is at stake here, because it is not a database of customers. The asset is the **safety of a physical process**, and the adversary who matters most is not a person typing at a browser but a compromised component in the update path. ## The notation for it Classic DFD notation has an answer that is neither pretending the box is atomic nor exploding everything. A **complex or multi-process element**, conventionally drawn as two concentric circles, marks a process that is known to contain several processes and has its own diagram beneath it. The marker is honest bookkeeping: it says on the face of the diagram that this box has not been analysed internally, so a reader does not mistake absence of internal threats for evidence there are none. Use it when the internals genuinely do not affect the decision in front of you — say, a mature component whose parts share an identity and an exposure anyway. Split the box when any of the four signals above applies, because then the internal edges *are* the analysis. ## The opposite failure Splitting everything is its own defect. A diagram with forty processes because someone enumerated every class in the codebase is unreadable, and unreadable diagrams do not get argued with. The discipline is not maximum resolution. It is that every box you leave whole is a box whose internals you have consciously decided are uniform in identity, exposure and data class — and that you can say so out loud when someone asks. ## In an interview The tell of a strong answer is that the candidate treats a suspiciously well-named box as a prompt rather than a fact. `Fleet controller`, `platform service`, `integration layer` are product names, not processes. Ask what identity it runs as and who can reach it. If there is more than one answer, the box needs splitting or marking.
- How do you decide between marking a box as complex and actually splitting it?Ask whether the internal edges change a decision you are making now. If the parts share an identity, exposure and data class, mark it complex and move on: the marker records honestly that you have not looked inside. If any part holds materially more privilege or accepts input from a different population, split it, because the flows between those parts are exactly what you would otherwise miss.
- What stops this from degenerating into a diagram with forty boxes?A rule about why you split, not how far. Split on differences in identity, exposure, data class or operator, and stop when the remaining internals are uniform on all four. Splitting to mirror the code structure produces resolution nobody can read and arguments about class boundaries rather than about threats.
- You inherit a diagram where every box is a team name. What is your first move?Treat each one as an open question and ask the owning team two things: what identity does this run as, and who can send it a request. Team names track org charts, not privilege, so the answers usually split at least one box immediately. Record the splits you make and the ones you deliberately deferred as a complex process.
A single box labelled with a product name is like one door labelled 'Operations': behind it sit the mailroom and the vault, and one lock on the door tells you nothing about which of them a visitor reached.
saying these in an interview costs you the question
- Treats a box name as evidence the box is one thing
- Splits processes to mirror code structure rather than privilege
- Assumes co-located components share a threat profile safely
- Never marks unexamined internals, implying they are clean
- Refuses to split because the diagram would look busy