Why annotate every data-flow diagram element with the principal and privilege it runs as?
answer
- The diagram shows what, not who
- Every box has an identity behind it
- Granted access, not designed intent
- A boundary is a difference between two labels
basics
~20 sWriting the principal an element acts as, and the privilege it actually holds, makes every privilege change visible on a data-flow diagram. Without those labels, boundary placement is guesswork and high privilege hides in plain sight.
solid answer
~50 sA data-flow diagram shows what moves; the principal annotation shows who it moves as. On each process I write the identity it executes under and the privilege that identity actually holds, not the role the design intended. Take a nightly ETL job: drawn plainly it is one box with an arrow into the ledger database. Annotated it reads `etl_runner` - database-administrator rights on the production ledger, write to every table - and that arrow now crosses a privilege change that was invisible a moment earlier. The label reframes the threats: anything that can influence that process, such as a dependency baked into its image or an analyst who edits its SQL, inherits ledger write access, so tampering with money and with the record of who moved it is in scope, not just a bad report. Boundaries are placed where the principal or its privilege changes, so unannotated elements make placement arbitrary.
go deeper
Be ready to say what a principal is - the account something runs as - and that a data-flow diagram does not show it unless you write it on. Knowing that a batch job has an identity too is most of what is asked here.
Explain the mechanics: what you write on a process, a store and a flow, and why the granted privilege rather than the intended role is the thing recorded. Expect to be handed a small diagram and asked to add the labels out loud.
Show the judgment of an annotator who has done it on real systems: chase the account nobody can describe, notice that a job holds standing write access it only needs at midnight, and use the labels to order findings by what each principal can reach.
Own the question of how annotations stay true across many teams and models - who is accountable for the grant behind a label, how staleness is detected, and how you keep a review from degenerating into boundaries drawn around network zones.
## Principal and privilege, defined A **principal** is the identity that a running piece of software or a human actor presents when it acts: a service account, a machine identity, a database login, an operating-system user, a scheduled-job account, a signed-in person. A **privilege** is what that principal is actually permitted to do. A **data-flow diagram** (DFD) - the processes, data stores, data flows and external entities that describe how data moves through a system - shows neither by default. A plain DFD says a nightly job reads a source table and writes the ledger. It does not say the job holds administrator rights on the ledger database and could rewrite any row in it. That missing half is what the annotation supplies, and it is the reason this step exists as its own discipline rather than as a footnote on the drawing. ## Write the granted privilege, not the intended role The annotation has two halves and both carry weight. The identity half names the account concretely - `etl_runner`, `console-svc`, `meter-ops` - because a label like *the application* covers a dozen components that may be one account or twelve, and you cannot tell which without asking. The privilege half records what that account can actually do today. Granted privilege routinely exceeds designed need: a job that must append to two tables is very often running under a login that can write all of them, drop them, or read every other schema on the instance. If nobody in the room can say what the account holds, that uncertainty is itself the first finding of the session, and it is worth stopping to resolve. ## What to write on each element type Process runtime identity + the privilege it effectively holds Data store which principals read it, which write it, at what privilege Data flow the identity presented on the wire for this call External entity the human or system role, and whether it is authenticated at all The data-flow line matters more than people expect: the identity a process presents on an outbound call is frequently not the identity of whoever caused the call, and a diagram that only annotates boxes will never show that. ## Worked example A nightly ETL job pulls transactions from an operational database and writes summarised rows into the production ledger. Unannotated, it is a box between two cylinders and looks like plumbing. Annotated, the box reads `etl_runner`, and the store it writes reads *ledger - `etl_runner` has write on all tables, `reporting_ro` has read*. Two things immediately change. First, the flow into the ledger crosses a privilege change, so a trust boundary belongs there. Second, the population of adversaries widens beyond the obvious: whoever controls the job's runtime controls a ledger-writing identity, which puts a compromised transitive dependency baked into the job's image, and the analyst with commit rights to the job's SQL, on the same footing as an intruder. The assets at stake are money and the truthfulness of the financial record, not merely a data leak. ## Why boundary placement depends on it A trust boundary belongs where trust or privilege changes. You cannot honestly place one without the labels, because the thing you are looking for *is* a difference between two labels. With annotations present, placement becomes close to mechanical: wherever adjacent elements carry different principals, or the same principal at materially different privilege, something crosses a boundary. Without them, teams fall back on drawing boundaries around network segments, which answers a different question entirely. ## What it does to enumeration and priority Per-element threat enumeration reads differently once each element is labelled: a threat against a process is evaluated in light of everything that process's principal can reach, which is the difference between *someone corrupts a report* and *someone rewrites the ledger*. The annotations also give you an ordering for limited review time - start with the elements whose principal holds the widest privilege over the most valuable asset, because any threat that lands there inherits all of it. ## Common annotation mistakes - Naming the owning team or the engineer instead of the runtime identity. - Recording the role the design document assigns rather than the grant that exists. - Collapsing many components into one label such as *the service* without checking. - Treating *runs as an administrator* as deployment trivia outside the model. - Skipping scheduled jobs and batch processes because no human is present, which is exactly where the widest standing privilege tends to live. ## Keeping the labels honest Annotations rot faster than the diagram does, because grants change without the architecture changing. Treat each principal label as a claim to re-verify when the model is refreshed, and note when it was last checked. A model whose principals were accurate a year ago places its boundaries where they used to be.
- The team tells you the job runs as an application account. Is that a good enough annotation?No. Name the account and state what it can actually do: which databases, which tables, read or write, and whether it can grant to others. The value of the annotation is comparability - you are looking for a difference between adjacent labels, and two elements both labelled the application account are indistinguishable by construction. If nobody knows the grant, record that as an open finding rather than writing a guess.
- How do you annotate a data store rather than a process?List every principal that touches it and at what privilege - who reads, who writes, who can change its structure. A store is usually where principals with very different rights meet, so its threats depend on the union of its accessors rather than on any one caller. That list also tells you which inbound flow crosses the sharpest privilege change and therefore deserves a boundary.
- Does the annotation change how you prioritise the threats you found?Yes. Order by what the principal can reach: a threat against an element whose identity holds broad write access to a valuable asset inherits that entire reach, while the same threat against a narrowly scoped element is contained. That gives a defensible ordering when a review turns up more findings than a team can work, without inventing numbers.
A floor plan shows the doors; the annotation is writing on each door whose key opens it. Only then can you see which doorway is a real threshold.
saying these in an interview costs you the question
- Records the role the design intended instead of the access actually granted
- Annotates the owning team or engineer rather than the runtime identity
- Labels a dozen distinct accounts as one application principal
- Treats runs-as-administrator as deployment detail outside the model
- Places boundaries at network edges because no principals were written down