skip to content

What is a Dynatrace automated root-cause verdict actually worth on call, and when is it wrong?

level: seniorimportance: should knowfreq 34%

answer

  1. A hypothesis with a blast radius attached
  2. It knows structure and timing, not intent
  3. Correctness failures move no timing signal
  4. No edge in the graph, no attribution
  5. Best answer is what else is affected

basics

~20 s

A Dynatrace verdict is a ranked hypothesis, not a diagnosis. From dependency edges, timing and change events it can say where a degradation started and what sits downstream. It stays blind to correctness bugs and causes outside the monitored estate.

solid answer

~50 s

The engine has three inputs: the dependency graph, per-entity deviation and timing, and change events attached to entities. From those it can legitimately conclude that one entity started deviating first, lies upstream of the others on a shared path, and that the rest are consistent with being downstream of it — a hypothesis with a blast radius, genuinely valuable at 02:00. It cannot say *why* that entity misbehaved. It is wrong in predictable classes: a shared resource with no edge in the graph, a correctness failure where nothing is slow, slow saturation with no crisp onset, a cause behind an uninstrumented hop, a cause outside your estate, and a change rolled out everywhere at once. So you keep owning business-outcome metrics, explicit context across handoffs the automation cannot see, and a small set of user-facing alerts. Use the verdict to scope blast radius, then confirm with a signal it did not use.

go deeper

for a junior

Recall the shape of the output: one finding naming a suspected entity plus the set of things affected, rather than a wall of independent alerts. Know that it is a suggestion to check, not a proven cause.

for a middle

Explain what the engine reasons over — dependency edges, deviation timing, change events — and therefore why it can order candidates on a path but can never say what is wrong inside the named entity.

for a senior

Name the failure classes from experience: shared resources with no edge, correctness bugs that move no timing, slow saturation, uninstrumented hops, causes outside the estate. Say what you instrumented to cover each.

for a principal

Decide how much of the estate's alerting rests on automation versus signals your teams own, and set the standard for what a team must instrument itself so the automation's blind spots are not organisational blind spots.

## What the verdict can honestly claim An automated causation engine has three inputs: the dependency graph, timing and deviation data per entity, and change events attached to those entities. From those it can make a genuinely strong claim, and it is worth stating precisely because candidates either overstate or dismiss it: > Among the entities that deviated, **this one appears to have started first and lies upstream of the others on a shared dependency path**, and here is the set of entities whose behaviour is consistent with being downstream of it. That is a hypothesis with a blast radius attached, and it is worth a great deal at 02:00. Two things it is emphatically **not**: it is not a statement about which line of code is wrong, and it is not a statement that the cause is inside the estate at all. The graph contains structure and timing; it contains no notion of correctness, intent or the world outside the machines it monitors. ## Where it will be wrong 1. **Shared invisible resource.** Two services with no edge between them degrade together because they sit on the same hypervisor, storage fabric or availability zone. With no edge to walk, the engine picks whichever moved first and names a peer symptom as the cause. 2. **Correctness failures.** Everything is fast, nothing deviates, and the data is wrong. No timing signal moves, so nothing is opened at all. This is the single largest category the automation cannot reach. 3. **Slow saturation.** A connection-pool leak or a queue filling over four days has no crisp onset, so the first-mover heuristic has nothing to latch onto and the deviation may be absorbed into the baseline as normal. 4. **Blind spots in the graph.** When the real cause sits behind an unrecognised protocol, an asynchronous handoff or an unsupported runtime, the nearest observed neighbour is blamed with full confidence. 5. **Causes outside the estate.** A third-party API, a DNS resolver, an expired certificate, a change shipped in a mobile client — all visible only as symptoms at your edge. 6. **Simultaneous fan-out.** A configuration change rolled everywhere at once produces dozens of near-simultaneous onsets, and ordering by first-mover becomes noise at millisecond distances. 7. **Chronically bad baselines.** An engine reacting to deviation from learned normal is quiet about a service that has always been slow, because nothing deviated. ## A worked case A school-meal ordering service running across 47 hosts took a support escalation: schools were reporting orders accepted for pupils with allergies that should have blocked them. The platform had opened nothing at all — request rates held at about 1,340 orders per minute, latency was flat, no entity had deviated. When the team forced the question, the nearest automated finding from earlier that morning pointed at a message broker with a brief queue-depth excursion, which was true and irrelevant. The actual cause was a nightly menu-import job that had written an empty allergen list for 118 dishes. Nothing was slow; a validation rule had been silently skipped. This is category 2 above, and it is exactly the incident nobody could explain from the existing dashboards: every chart was green because every chart measured time and volume, and the failure was in meaning. It was found by a business metric the team added afterwards — accepted orders whose allergen list is empty — not by anything the platform could have inferred. ## What you keep owning yourself | The finding gives you | You still have to establish | |---|---| | Where the deviation appears to have started | Why that entity misbehaved | | Which entities are on the blast path | Whether users are actually harmed | | That a change event coincides with the onset | Whether that change is causal | | That something deviated from normal | That "normal" was ever acceptable | Concretely, the instrumentation that stays yours: - **Business outcomes.** Orders accepted per minute, payment success rate, submissions rejected for validation. These are the only signals that move during a correctness failure, and no platform can infer which counter means "children get fed". - **Explicit context across the edges the automation cannot see** — queue handoffs, batch jobs, custom protocols — so the graph has an edge where causation would otherwise stop. - **A small set of symptom alerts you actually own**, expressed on user-facing outcomes, so the automated verdict is an input to an investigation rather than the thing that decides whether anyone is woken. - **Black-box checks from outside the estate**, which are the only way to see a failure whose cause is not on any monitored machine. ## How to use it on call Treat the verdict as a ranked hypothesis and, above all, as a **scoping** tool. Its most reliable output is not "what caused this" but "what else is affected", and on a large estate that answer is often the more valuable one. Check the stated onset against your own change record rather than accepting the attached events as causal, and confirm with one signal the engine did not use before acting on the verdict — a business counter, a log line, a manual request. The failure mode to avoid in an interview and in production is the same: repeating the verdict as a conclusion when it was only ever the fastest available hypothesis.

  • Why is a correctness failure invisible to automated causation, however good the dependency graph is?
    Because causation is inferred from deviation in timing and rates along edges, and a correctness bug moves neither. Requests succeed at normal latency and normal volume while carrying wrong values, so no entity deviates and nothing is opened. Only a signal expressed in domain terms — orders rejected, allergen lists empty, balances mismatched — moves at all, and that signal has to be written by someone who knows the domain.
  • The engine blames a service that turns out to be a peer symptom, not a cause. What does that usually indicate?
    Almost always a missing edge. Two entities that share something the graph does not model — a hypervisor, a storage fabric, an availability zone, an uninstrumented hop — degrade together with no relationship between them, so the analysis falls back on whichever moved first. The fix is to make the shared dependency visible as an entity, not to distrust the engine generally.
  • How would you use an automated verdict without letting the team stop thinking?
    Treat it as scoping first and attribution second: its most reliable output is the affected set, not the cause. Require one confirming signal the engine did not use — a business counter, a log line, a manual request — before acting, and check the claimed onset against your own change record rather than accepting attached events as causal. Verdicts that turn out wrong are worth reviewing for the missing edge that caused them.

saying these in an interview costs you the question

  • Treats the verdict as a diagnosis rather than a ranked hypothesis to confirm
  • Assumes the engine finds correctness bugs where nothing is slow
  • Thinks a change event attached to the onset proves the change was causal
  • Believes automated causation removes the need for business-outcome metrics
  • Dismisses the verdict entirely instead of using it to scope blast radius