skip to content

How do you draw a data-flow diagram for a serverless architecture with no host to put on it?

level: middleimportance: must knowfreq 64%

answer

  1. same four elements, different nouns
  2. no host does not mean no boundary
  3. queues and config are data stores
  4. boundary follows the calling identity
  5. one process per execution role

basics

~20 s

Use the same four DFD elements: functions and managed-service calls are processes; buckets, tables and queues are data stores; event sources are external entities. Draw trust boundaries where the calling identity changes, not where a machine or network ends.

solid answer

~50 s

The absence of a host is not the absence of boundaries. I keep the four data-flow-diagram elements and change what fills them: a process is any unit of code running under one execution role, a data store is a managed store such as an object prefix, a table, a queue or the function's own configuration, and every event source is an external entity. The trust boundary answers one question on this diagram: which principals can make this call. Two functions with different execution roles are separated by a boundary; two functions sharing one role are a single element no matter how many boxes I drew. I label each flow with how it is authenticated and authorised: a signed call from a named role, a queue policy, a presigned URL with an expiry, an unauthenticated public endpoint. Role assumption gets its own arrow, because it is the crossing.

go deeper

for a junior

Be ready to name the four data-flow-diagram elements and give a serverless example of each: an uploader as an external entity, a function as a process, a bucket or queue as a data store, an invocation as a flow.

for a middle

An interviewer expects you to explain why the boundary follows the execution role rather than the network, and to show that a queue and a function's configuration are data stores with reader and writer sets you must write down.

for a senior

Show you can drive this at a whiteboard: annotate each flow with what authenticates it, draw role assumption as a crossing, and say out loud which boxes collapse into one element because they share a role.

for a principal

Own the standard: decide what a team's serverless diagram must always show - roles, event sources, configuration stores, control-plane crossings - so models from different teams can be compared and reviewed rather than each being an artistic choice.

## Why the diagram feels impossible at first A data-flow diagram (DFD) has four element types: **external entities** (actors outside your control that talk to the system), **processes** (code that transforms data), **data stores** (where data rests), and **data flows** (the arrows between them). **Trust boundaries** are drawn across flows wherever the level of trust changes, and they are the reason a DFD is worth drawing at all: a threat is most often something crossing a boundary in a way you did not intend. On a classic deployment the boundaries almost draw themselves, because they line up with things you can point at: the internet edge, the DMZ, the application subnet, the database subnet, the machine you administer versus the machine you do not. A serverless workload takes all of that away. The compute exists for the duration of one invocation, you never log into it, there is no long-lived process to attach to, and often no network diagram to trace at all. Many engineers conclude the system has no boundaries. It has the same number as before. They are drawn along **identity** rather than along wiring. ## Map the four elements onto managed services ``` Element Typical serverless representative ----------------- ------------------------------------------------------------- External entity a browser or mobile client, a partner system, an uploader holding a presigned URL, a scheduler, another team's account Process a function, a managed transform or workflow step, a container task - anything that runs under exactly one execution role Data store an object bucket or prefix, a table, a queue or topic, a parameter or secret store, and the function's own configuration Data flow an invocation, a read or write, an event delivery, and a role assumption ``` Two of those entries surprise people. A **queue is a data store, not an arrow**: something writes to it, something else reads from it later, and it has its own access policy that decides who may do each. And **configuration is a data store**: environment variables, parameter entries and the role attached to a function are data at rest, with their own reader and writer sets, and they are frequently where a secret actually lives. ## Identity is the perimeter On a managed-services diagram, the useful boundary is the set of things one principal can reach. Concretely: - Two functions with **different execution roles** are separated by a trust boundary, even though they sit in the same account, the same repository and the same deployment. - Two functions with the **same execution role** are not separated by anything. Compromise one and you have the other's access. Draw them as one element for threat-enumeration purposes even if you keep two boxes for readability. - **Role assumption is a boundary crossing.** Wherever code assumes another role, that arrow changes what the code can reach. Draw it, and label who may assume it, under what condition, and what the assumed role reaches. This is why the network picture is the wrong picture. A function inside a private network segment that holds a role with read access to a customer data store reaches that data regardless of the segment; a subnet drawn as the only boundary hides exactly the reachability that matters here. ## Annotate the flows, not just the boxes For each arrow write down how the call is authenticated and what authorises it: a signed request from a named role, a queue resource policy naming one producer, a presigned URL with a short expiry, a public endpoint with no authentication at all. Unlabelled arrows are where teams silently assume that an internal-looking hop is trusted. There is also a second plane of flows worth marking even if you do not model it in depth here: the **control plane**. Someone can change a function's code, its environment variables or its role without touching the data path at all, and the set of principals who can do that is usually different from the set who can call it. Mark that crossing on the diagram; modelling the build and deploy path as a system in its own right is a separate exercise. ## A worked sketch A media-transcoding workload, drawn the way this section argues: ``` (External uploader, holds a presigned URL) | PUT object v [ uploads/ prefix ] --object-created event--> ( Transcoder, role: transcoder ) | write v [ output/ prefix ] Boundaries: anonymous internet | transcoder role | readers of output/ ``` There is no host anywhere on that picture and it is still a complete model. It tells you the uploader is anonymous and bounded only by the presigned URL's scope and expiry; that the object-created event is an entry point carrying attacker-chosen bytes; that everything the transcoder role can reach is inside one blast radius; and that the interesting question about `output/` is who else can read it. ## Common failure modes - Drawing infrastructure - regions, network segments, service icons - instead of data flows, and ending up with a deployment diagram that supports no threat enumeration. - Collapsing every function into one box called *the app*, which erases exactly the role differences that are the boundaries. - Leaving configuration and secrets off the diagram, so nobody asks who can read them. - Treating a managed store as somebody else's problem rather than a data store with a reader set you chose.

  • If there is no host, where does data at rest actually live on this diagram?
    In managed stores and in configuration. Object prefixes, tables, queues and parameter entries are the obvious ones; the less obvious one is the function's own configuration, because a third-party API secret sitting in a plaintext environment variable is disclosed to every principal in the organisation who can describe that function. Put configuration on the diagram as a store and write down its reader set.
  • Should the deploy and configuration path appear on the same diagram as the request path?
    At minimum as a labelled crossing. Anyone who can change a function's code, its environment or its attached role changes what the system does without ever sending it a request, and that principal set is usually different from the callers. Mark the flow and who may make it; modelling the delivery path itself as the system under analysis is a separate exercise with its own diagram.
  • How granular should a process be - one per function, or one per service?
    One per distinct permission set. If five functions all run under the same execution role, they are one element for threat purposes, because a threat that reaches any of them has the access of all of them. If one function assumes a second role for part of its work, that is two elements joined by a role-assumption arrow.

It is like mapping a building by who holds which keys rather than by where the walls are: two rooms sharing a key are one room to a burglar.

saying these in an interview costs you the question

  • Says a serverless system has no trust boundaries because there are no servers
  • Draws network segments as the only boundary and stops there
  • Puts every function in one box because they are all the same application
  • Treats queues as arrows rather than as stores with their own access policy
  • Leaves function configuration and secrets off the diagram entirely
  • Never asks who can change the function's code or role

context