skip to content

Recovering an Undocumented System

Where the elements and flows come from when nobody wrote a design down: infrastructure-as-code, route config, repo egress calls and request telemetry. Interviewers ask how you know it is right.

on this pageshow

questions

3

Which sources do you rebuild a data-flow diagram from for an undocumented legacy service?

level: middleimportance: must knowfreq 62%

answer

  1. evidence, not memory
  2. what is permitted vs what happened
  3. config over-states, telemetry under-states
  4. grep the repo for egress hosts
  5. walk the draft with the operators

basics

~20 s

Rebuild the diagram from artifacts that describe the running system: infrastructure-as-code and its state, route and proxy configuration, outbound calls grepped from the repository, identity bindings and database grants, and request telemetry. Confirm each edge with an operator.

solid answer

~50 s

I never start from memory or from someone's slide. I start from five kinds of evidence and treat each as partial. Declarative infrastructure config and its state file give me provisioned networks, queues, buckets and managed databases. Route and reverse-proxy configuration gives me inbound paths and which upstream serves each. Grepping the repository for hostnames, clients, connection strings and webhook URLs gives me outbound edges. Identity role bindings and database grants give me who is *permitted* to reach what. Request telemetry, access logs and flow logs give me what was actually *used*. The two directions cross-check each other: config over-states (dead rules, orphaned grants) and telemetry under-states (only the observation window). I draw the union as candidate edges, then walk the draft with the people who still operate the thing, and mark anything they cannot confirm as an explicit assumption on the diagram rather than deleting it.

go deeper

for a junior

Be ready to name concrete places facts come from: the infrastructure config, the route file, the code's outbound calls, the grants, the logs. Naming three real sources beats a vague answer about "looking at the architecture".

for a middle

An interviewer expects you to explain the mechanics of each source and its blind spot — that config states intent while telemetry states behaviour, and that the two fail in opposite directions. Show the cross-check, not just the list.

for a senior

Show how you sequence the work under time pressure and how you convert a draft into confirmed facts by walking it with operators. Demonstrate that you annotate each edge with its evidence and date rather than presenting a recovered diagram as settled truth.

for a principal

Own the question of what recovery is worth across an estate: what evidence your platform should emit so this exercise is cheap next time, and how you keep recovered models honest about their own confidence instead of laundering guesses into architecture documents.

## The situation A twelve-year-old freight-forwarding customs-declaration broker still runs in production. It submits declarations to a government gateway and pulls tariff and consignee data from a third-party customs-data vendor. The last engineer who wrote it left years ago; there is no diagram, no design document, and two operators who know how to restart it. You have been asked to threat-model it. The asset that matters most here is not customer records, it is **audit truth and non-repudiation**: a declaration that was altered, replayed, or attributed to the wrong filer is a legal problem, not just a data problem. A plausible adversary is the vendor at the far end of one of those flows being compromised. You cannot threat-model a system you cannot describe. Recovery comes first, and recovery is an evidence exercise. ## The five evidence sources, and what each one lies about **1. Declarative infrastructure config and its state.** The repository that provisions the estate, plus the state it records, tells you which networks, subnets, load balancers, queues, object stores and managed databases exist and how they are wired. It is the best single source for *topology*. What it misses: everything created by hand outside it, everything that drifted after the last apply, and everything provisioned before the team adopted the tooling — for a twelve-year-old system, that is often the older and more interesting half. **2. Route and reverse-proxy configuration.** The route file, the ingress rules, the API route table: these enumerate the inbound paths that reach the application and which upstream serves each one. This is where the entry side of the picture comes from. What it misses: everything that is not an HTTP request — a queue consumer, a scheduled job, a file drop, an operator with a shell. **3. The repository itself, read for egress.** Grep for hostnames, base URLs, SDK client construction, connection strings, webhook targets, SMTP and SFTP endpoints. This is usually the only way outbound third-party edges surface at all, because nothing in the infrastructure config has to mention a host you merely call over the internet. What it misses: whether the code path is live, which environment the endpoint belongs to, and any destination that arrives at runtime from configuration or a database row. **4. Identity bindings and grants.** Service-account role bindings, cross-account trust policies, database `GRANT`s and key policies tell you *who is permitted to reach what*. A grant is not proof that a flow happens — but it is proof that a flow **could** happen, which for threat modeling is nearly as important. Draw it as a candidate edge and label it: the permission is itself attack surface. **5. Request telemetry.** Access logs, flow logs, request traces and metrics tell you what genuinely moved, when, how often, and to where. This is the only source that reflects behaviour rather than intent. What it misses: anything outside the retention or observation window. A quarterly reconciliation run leaves nothing in a 30-day log, and its absence is not evidence it does not exist. ## Cross-checking, not merging The useful discipline is to keep the two directions of error apart: | Source type | Systematic error | |---|---| | Config, code, grants | **Over-states**: dead rules, orphaned grants, unreachable code | | Telemetry | **Under-states**: only the window, only what fired | So: the union of both is your candidate edge set; the intersection is the set you can call confirmed without asking anyone. Everything in the difference is a question for a human. Config with no traffic may be dead — or may be the quarterly job. Traffic with no config is the more alarming direction and deserves its own investigation. ## Then talk to the operators The draft is a hypothesis. Walking it edge by edge with the two people who still run the broker converts guesses into facts faster than any amount of further reading, and it surfaces the flows that leave no artifact at all: the analyst who pulls a monthly extract by hand, the vendor who mails a certificate once a year, the emergency path someone uses when the gateway rejects a batch. Record who confirmed each edge and when — provenance on the diagram is what lets the next reader judge it. ## What goes on the paper Annotate each edge with the evidence that produced it — config, telemetry, code, operator statement, or inference — and mark unverified edges visibly rather than omitting them. An edge you silently drop becomes a threat nobody ever considers; an edge drawn with a question mark on it becomes a task. For the customs broker, the vendor flow drawn as a real, confirmed edge with its protocol and data class is what lets you ask the question that matters: if that vendor is compromised, what can it do to the truth of a declaration?

  • Why is an identity role binding weaker evidence of a data flow than a log entry?
    A binding proves permission, not occurrence — the grant may be orphaned from a decommissioned caller. A log entry proves the flow was actually taken. But do not discard the binding: an unused permission that would let one component read a store is a latent edge, and it is exactly what an attacker who lands on that component gets to use. Draw it as a candidate edge and label its evidence as "grant only".
  • You grep the repository and find an outbound call to a vendor host. What do you still not know about that edge?
    Whether the code path is reachable at all, which environment the host belongs to, what protocol and authentication it uses, which direction the sensitive data actually moves, what data class it carries, and how often it fires. Resolve those from route config, credential storage and telemetry, and ask an operator for the rest. Until then it is a drawn edge with an incomplete annotation, not a fact.
  • Telemetry shows nothing for a flow the configuration clearly permits. Do you draw it?
    Yes, and mark it unconfirmed. Absence of traffic in a retention window is not absence of a flow: quarterly jobs, disaster-recovery paths and break-glass access all look identical to dead config in thirty days of logs. Check schedulers and job definitions, ask the operators, and only retire the edge once someone can say what replaced it — recording that decision so the next reviewer does not re-derive it.

It is closer to surveying an old building than to reading its blueprints: you measure what is standing, and you mark on the plan the walls you could not get behind.

saying these in an interview costs you the question

  • Treats the infrastructure-as-code repository as complete and authoritative
  • Assumes no traffic in the logs means no flow exists
  • Treats a permission grant as proof the flow is in use
  • Never speaks to the people who actually operate the system
  • Reconstructs from a years-old slide deck instead of running artifacts
  • Silently omits an edge that could not be verified

context

open as a page

Your recovered data-flow diagram has an edge telemetry proves but no configuration explains — what now?

level: seniorimportance: should knowfreq 44%

basics

~20 s

An observed flow is real evidence, so the configuration set is what is incomplete. Characterise the edge from telemetry — endpoints, identity, protocol, volume, schedule — then walk it with the on-call operator, keeping it drawn until someone accounts for it.

open as a page

How complete must a recovered data-flow diagram be before you start enumerating threats on it?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

A recovered diagram is ready when it answers the decision at hand: every trust-boundary crossing drawn, every data store's class named, every identity reaching a sensitive store enumerated, and every unverified edge labelled. Time-box the rest.

open as a page