How does Dynatrace build its live topology map, and why does automated root cause need one?
answer
- Derived continuously, never drawn by hand
- Hosts and processes, not only services
- Edges from observed connections and traced calls
- The graph is versioned in time
- Causation walks edges, correlation ranks charts
basics
~20 sDynatrace's agents continuously report hosts, processes, containers and the connections between them, and the platform keeps that as a time-versioned entity graph. Automated root cause needs it because causation is a walk over dependency edges, not a correlation of charts.
solid answer
~50 sThe map is derived, not drawn. Agents on each host report the processes and containers running there, the requests traced between them, and — crucially — the connections they simply **observe** one process opening to another. From that Dynatrace maintains a graph of entities (hosts, containers, process groups, services, databases, queues) joined by *runs on* and *calls* relationships, and it keeps that graph **versioned in time**, so an investigation can ask what depended on what at the moment of the degradation, not only now. Because agents observe connections, components nobody instrumented — a database, an appliance, a legacy process — are still real nodes. The graph is the precondition for causation because timing alone yields correlation: forty entities degraded at once. The graph turns that flat list into a directed search — which candidates share a dependency path, which end of the path moved first, and what changed on those nodes.
go deeper
Recall that the map is discovered rather than maintained by hand, and that it contains infrastructure as well as services: hosts and containers appear on it, not just the applications you deploy.
Explain both sources of edges — traced calls between processes and connections the agent simply observes — and why the second is what puts uninstrumented databases and appliances on the map.
Show why the graph is operationally load-bearing: it scopes an incident's blast radius, orders candidates along a path, and lets analysis descend from a service into the host it shares with others.
Argue about what the model is worth across an estate: entity identity as the thing that makes cross-signal correlation mechanical, and the cost of the blind edges the graph will never contain.
## What the model actually contains Dynatrace does not draw a picture of your architecture from a diagram someone maintained; it derives one continuously from what its agents observe. The result is a graph of **entities** — hosts, containers, process groups, services, databases, queues, and the applications in front of them — connected by **relationships**: *runs on*, *is part of*, *calls*. Every signal the platform holds hangs off one of those nodes, so a metric, a trace, a log line and a deployment event about the same service are attached to the same identity rather than to four independent naming schemes that only a human can reconcile. That last point is quietly the most important one. In an estate whose metric store already holds 2.3 million series, the hard problem is not collecting data; it is knowing that *this* series, *that* trace and *those* log lines describe the same running thing. Stable entity identity is what makes cross-signal reasoning mechanical instead of a judgement call. ## Where the edges come from Two sources, and the second is the one candidates forget: 1. **Traced calls.** When a request crosses from one instrumented process to another, that is a caller-to-callee edge with timing attached. 2. **Observed connections.** The agent sits on the host and can see which process opened a connection to which endpoint, whether or not anything on the far end is instrumented. The second source is why the map is not limited to what somebody remembered to instrument. A managed database, a network appliance, a legacy service in an unsupported runtime and a third-party endpoint all appear as real dependencies, and the infrastructure layers beneath every service — process, container, host — are modelled as entities in their own right rather than as tags on a span. | | Topology derived from agent observation | Topology assembled only from calls that instrumented services report | |---|---|---| | Uninstrumented database or appliance | present as an observed dependency | present only as a client-side call, if at all | | Host and container beneath a service | modelled as linked entities | absent entirely | | A service nobody instrumented | visible as a process with connections | invisible | | Async handoff through a queue | edge only where the handoff is recognised | edge only where context propagates | | Cost of coverage | onboard the machine | instrument every service | ## Why it is time-versioned The graph carries history: what depended on what at 14:07 yesterday, not merely what depends on what now. This is not archival tidiness, it is a correctness requirement. Estates move constantly — pods rescheduled, versions rolled forward, traffic shifted between regions — and an investigation always runs **after** the degradation. A current-state graph asked about a forty-minute-old symptom would reason over edges that did not exist at the time and miss ones that did, which produces confident nonsense rather than an obvious error. ## Why causation needs a graph at all Without a dependency model, an automated analysis has only timing, and timing alone gives you correlation: forty things got worse in the same three minutes. Ranking them by how badly they moved is a popularity contest, and the loudest signal is usually the most user-facing one — that is, the symptom. The graph converts that flat list into a **directed search**: - **Scoping.** Take the entities that deviated and ask which of them lie on a single dependency path. Ones that do are candidates for a shared cause; ones that do not are probably a separate story. - **Ordering.** Along a path, ask which end started deviating first. Upstream-and-earlier is a cause-shaped position; downstream-and-later is a symptom-shaped one. - **Layer traversal.** Because the process and host beneath a service are entities too, the search can descend out of the application tier into the infrastructure that every service on that box shares. - **Aggregation.** Many degraded entities on one path become a single finding with a stated blast radius, instead of each threshold breach announcing itself independently. - **Context attachment.** Deployment and configuration events are attached to entities, so "this started two minutes after that node changed" is a query over the graph rather than a human recalling a change window. None of those five steps is expressible without knowing the shape of the estate. That is the whole answer to "why is the map not just a nice picture": it is the index that makes causation a walk rather than a search over everything. ## What the map still will not do It is a model of structure, not of intent. It tells you that the ordering service calls a menu database; it does not know that one of those calls is the only path by which 8,900 pupils get a meal booked. Nor does the graph rescue you where an edge is genuinely invisible — an asynchronous handoff nobody propagated context across, a dependency reached through a third party, or two services that share a hypervisor with no edge between them at all. When the cause hides in a missing edge, the analysis will confidently blame the nearest node it can see.
- How does a graph built from observed connections differ from one assembled only from calls that instrumented services report?An observation-derived graph includes components nobody instrumented — a managed database, an appliance, a third-party endpoint — because the agent sees the connection from the client side of the machine. It also models the process, container and host beneath each service as entities, so an analysis can descend below the application tier. A graph built only from reported calls contains exactly the services that emit them and nothing about what they run on.
- Why must the topology graph carry history rather than only the current state?Because every investigation happens after the fact and the estate has moved in between: pods rescheduled, versions rolled, traffic shifted. Reasoning about a forty-minute-old degradation with today's edges attributes the symptom to dependencies that did not exist then and overlooks ones that did. A time-versioned graph lets the analysis reconstruct the estate as it was at the moment the deviation started.
- How does the graph collapse many simultaneous alerts into one finding?Degraded entities that sit on a single dependency path are attributable to one suspected cause, so the platform opens one finding with a stated blast radius rather than one announcement per threshold breach. The value is not fewer messages for their own sake — it is that the affected set is stated as a consequence of a hypothesis, which is exactly what an engineer needs to judge whether the hypothesis is plausible.
Forty alarms sounding at once tell you only that something is wrong. The dependency graph is the building's floor plan, which is what turns a chorus of alarms into a direction to walk in.
saying these in an interview costs you the question
- Thinks the dependency map is drawn by hand or imported from a service catalogue
- Assumes the graph contains only services that somebody instrumented
- Treats the map as current-state only, with no history to reason over
- Confuses two things degrading at the same moment with a causal dependency
- Believes the map is the deliverable rather than the input that makes causation possible