skip to content

What does the Dynatrace OneAgent do to a running application, and what does it still not see?

level: middleimportance: must knowfreq 58%

answer

  1. No code change, no library added
  2. One install per host, not per service
  3. It discovers processes and instruments them
  4. Deep visibility usually waits for a restart
  5. Timing it knows, business meaning it does not

basics

~20 s

Dynatrace's OneAgent installs once per host, discovers the processes running there, and injects instrumentation into supported runtimes so traces, code-level timings and host metrics appear with no code change. It cannot supply business meaning or cover unsupported runtimes.

solid answer

~50 s

Dynatrace's OneAgent is installed **once per host**, not per application, with no dependency added to any build. It discovers the processes there, decides which are supported technologies, and loads instrumentation into managed runtimes such as the JVM **as the process starts** — which is why host and process metrics appear immediately but method-level detail usually waits for a restart. From then on it reports request entry points, outbound database and HTTP calls, code-level timings, host and container signals, and the connections it observes between processes, all keyed to the same entities. What it cannot produce is meaning: it knows a call took 812 ms, not that the call was a lunch-order submission that must never fail. It also misses asynchronous handoffs it cannot correlate, unrecognised protocols and batch work with no inbound request. Those gaps are what your own spans, attributes and business metrics exist to close.

go deeper

for a junior

Be ready to say where the agent lives: one install on the host, not a dependency in your build file, and traces plus host metrics show up without you editing code.

for a middle

Explain the mechanics. Process discovery, instrumentation loaded into supported runtimes at process start, why a restart is usually needed for method-level data, and which signals arrive without one.

for a senior

Show where the automation stops. Name blind spots you have actually hit — queue handoffs, unrecognised protocols, batch jobs, unsupported runtimes — and say what you instrumented by hand to close each.

for a principal

Own the boundary. Decide what the platform captures for every team for free versus the small set of domain signals each team must add, and keep that hand-written half consistent without doubling telemetry volume.

## One agent for the host, not one library per service Dynatrace's OneAgent is installed on a **host** — a virtual machine, a bare-metal box, or as a node-level component on a container platform — and is not added to any application's dependency list. Once running it enumerates the processes on that host, classifies them by technology, and keeps watching as processes come and go, so a container scheduled onto that node forty minutes later is picked up with nobody doing anything. That is the difference that surprises engineers arriving from a per-application agent: the unit of deployment is the machine, so the coverage question becomes *"is this host onboarded?"* rather than *"did this team add the library?"*. For a managed runtime such as the JVM or .NET, code-level visibility comes from the agent loading its own instrumentation into the process, and that happens **as the process starts**. A freshly onboarded host therefore reports CPU, memory, disk, network and process lifecycle immediately, but shows no method-level timings for services that were already running until they are restarted. Teams that miss this conclude on day one that the agent "does not work". Some compiled languages need a build-time step or a linked library instead, which is worth confirming before promising fleet-wide coverage. ## What arrives without anyone writing code Once instrumentation is in the process, the agent recognises the frameworks and clients inside it and instruments their boundaries: - **Inbound entry points** — the request handlers of recognised web frameworks, so each request becomes an end-to-end traced transaction (Dynatrace calls one a PurePath). - **Outbound calls** — database clients, HTTP clients, remote-call libraries and message-broker clients, which supply the caller-to-callee edges that stitch one trace across several processes. - **Code-level timing inside the process** — which methods on the request path consumed the time, at a granularity nobody had to anticipate in advance. - **Host, container and process signals** — utilisation, limits, restarts, and the lifecycle of every process on the box. - **Observed connections** — which process talks to which, including to components nobody instrumented at all. - **Log files written by the processes it already knows about**, which is how log lines land on the same entities as the traces. The property that makes this more than a pile of signals is that everything is keyed to the same entities: a slow request, the process that served it, the container it ran in and the host underneath are one connected model rather than four disconnected dashboards. ## Where the automation stops Automatic capture answers *what happened and how long it took*. It does not answer *what it meant*. | Question about a request | Captured automatically | Needs instrumentation you write | |---|---|---| | How long did the call take? | yes | — | | Which method burned the time? | yes, on supported runtimes | — | | Which process called which? | yes, from observed connections | — | | Was this a school lunch-order submission? | no | span attribute or custom metric | | Was the allergen list on the order correct? | no | business metric or validation event | | Did this queue consumer's work belong to that request? | often not | explicit context propagation | | What is happening in an unsupported runtime? | process-level signals only | your own telemetry | The recurring blind spots are worth naming explicitly, because they are the ones that cost hours: 1. **Asynchronous handoffs.** A request that ends by putting a message on a queue and a consumer that picks it up eight seconds later are two traces unless the handoff is recognised or you propagate context yourself. 2. **Unrecognised protocols and runtimes.** A custom binary protocol between two services shows as an observed connection with timing, not as a stitched call with parameters. 3. **Batch and scheduled work.** There is no inbound request to hang a trace on, so the automation has nothing to anchor to. 4. **Meaning.** The agent can tell you a call took 812 ms; it cannot tell you that the call was the only path by which 8,900 pupils get a meal booked, or that the response was fast and wrong. ## What that means in practice On a school-meal ordering estate of 47 hosts, onboarding gave the team an accurate map and full request timings within a day of restarts, and it genuinely removed the "please add the tracing library" negotiation from forty teams. What it did not remove was the need for a small number of hand-written signals: an attribute marking which requests are order submissions, a counter for accepted orders per minute, and an explicit link across the menu-import job's queue handoff. Those three additions are what later made an incident explicable that nobody could account for from the existing dashboards, in an estate whose metric store already held 2.3 million series. The honest summary for an interview is: automatic instrumentation buys you breadth, uniformity and coverage of things you would never have instrumented by hand — and it buys you none of your domain. Breadth is exactly what a platform can automate; meaning is exactly what it cannot.

  • Why does code-level detail for a Java service usually appear only after the process is restarted?
    Because the agent gets inside a managed runtime by loading its instrumentation as the process starts. A process that was already running keeps the code it started with, so nothing has been instrumented inside it. Host, container and process-level signals need no restart — they are observed from outside — which is why a half-onboarded estate shows utilisation everywhere but method timings only where services have cycled.
  • What kinds of failure leave the automatically stitched trace incomplete?
    Asynchronous handoffs are the big one: a request that enqueues work and a consumer that picks it up later become two traces unless the handoff is recognised or context is passed explicitly. Custom or unrecognised protocols appear as an observed connection with timing but no stitched call. Batch and scheduled jobs have no inbound request to anchor a trace to, and unsupported runtimes give process-level signals only.
  • What does the agent still give you on a host that runs no supported application runtime at all?
    Host and process signals — utilisation, limits, restarts, process lifecycle — plus the network connections it observes between processes. That is what keeps databases, appliances and third-party components present in the dependency model even though nothing was instrumented inside them, and it is why coverage of those components costs nothing beyond onboarding the machine.

It is a master key kept at the building's front desk rather than a lock fitted to every apartment door: one install and every tenant process is reachable, but the key tells you who came and went, never why it mattered.

saying these in an interview costs you the question

  • Claims each service must add a vendor library to its build for the agent to work
  • Thinks one agent is installed per application process rather than per host
  • Says code-level detail appears instantly on processes that were already running
  • Assumes automatic instrumentation removes any need to write your own spans or metrics
  • Believes the agent can tell you why a business value is wrong, not just that a call was slow