skip to content

When is leaving a sampling profiler always on in production worth its cost, versus profiling on demand?

level: seniorimportance: nice to knowfreq 20%

answer

  1. You cannot profile a process that is gone
  2. Overhead follows sample rate, not throughput
  3. Watch the tail, not the mean
  4. Symbols must outlive the profiles
  5. Decide separately per profile kind

basics

~20 s

Always-on profiling earns its cost when the incidents you must explain are already over — a restarted or rescheduled process cannot be profiled retroactively. On-demand capture suffices for reproducible workloads and costs nothing while idle.

solid answer

~50 s

Continuous collection buys one thing on-demand capture structurally cannot: a profile of a window that has **already ended**. By the time a human looks at an incident, the process is often gone, so attaching a profiler answers questions about a healthy system rather than the sick one. It also enables comparison — this build against last week's, this instance against its peers — and fleet-wide attribution of a code path's total cost, which no single ad-hoc capture gives. Against that: overhead is small but never zero, and it lands on tail latency as well as on average CPU. Symbols must be retained for as long as the profiles, and storage grows with fleet size times window times retention. On-demand still wins where the workload is reproducible, where the profile kind is expensive to collect, or where the runtime makes continuous capture risky. Measure the overhead on a subset before committing the fleet.

go deeper

for a junior

Know that profiling is not free and that a profiler has to be running while the problem is happening. Remember that a process which has already restarted cannot be profiled retroactively.

for a middle

Explain why sampling overhead follows the interrupt rate and stack depth rather than request volume, and why storage cost is fleet size times window count times retention.

for a senior

Argue the tradeoff on a real incident: what you could not answer because nothing was collecting, how you would measure the overhead on a subset of nodes, and which profile kinds you would leave off.

for a principal

Own it as a portfolio decision — which services justify continuous collection, what retention and symbol policy makes the data usable months later, and how you keep the standing cost proportionate to the incidents it actually resolves.

## What continuous collection actually costs The overhead is real but usually modest, and it comes from four places rather than one. - **CPU.** Cost tracks the interrupt rate and the depth of the stack being walked, not the amount of work the service does. Sampling tens of times a second rather than thousands is what makes always-on collection affordable in the first place. - **Tail latency.** Interrupts land unevenly, and a request unlucky enough to absorb several stack walks pays for them. Judging the overhead by average CPU alone is the classic mistake; watch the high percentiles. - **Storage and symbolisation.** Aggregated profiles merge identical stacks, so they compress well, but the bill is fleet size times window count times retention. Worse, the symbol data that maps addresses to function names must be kept for as long as the profiles it explains, or old profiles become unreadable. - **Operations.** Something extra runs on every node or inside every process, and it must be deployed, upgraded and reasoned about when it misbehaves. None of those numbers is knowable from a document. The only honest answer to 'what does it cost' is that you measured it on your workload. ## What it buys that on-demand cannot 1. **Retrospective incidents.** This is the argument that decides it. A process that has been restarted, rescheduled or scaled away cannot be profiled after the fact. Continuous collection is the only way to have data from a window nobody was watching. 2. **Comparison over time and across builds.** 'Was this code path this expensive last week?' is unanswerable from a single capture, and it is the question that separates a regression from a long-standing cost. 3. **Fleet-wide attribution.** The total cost of a library across every instance is a different quantity from its cost on the one host somebody profiled, and it is the quantity that justifies engineering effort. 4. **Steady-state waste.** Continuous data surfaces permanent inefficiency — a serializer costing a fifth of the fleet's CPU — that nobody would ever open an incident about. ## Where on-demand still wins - **Reproducible workloads.** A batch job, a load test or a staging repro can simply be run again with a profiler attached, at zero standing cost. - **Expensive profile kinds.** Periodic stack sampling for CPU is cheap. Profilers that hook every allocation or every lock acquisition are a different order of cost and are better enabled on demand, or continuously at a heavily reduced rate. - **Constrained runtimes and environments.** Where stack walking is unreliable, where symbols cannot be shipped, or where the platform forbids the required privileges, continuous collection buys unreadable data. - **Small estates.** With a handful of long-lived services and someone always on hand, the retrospective argument weakens considerably. ## A worked decision A cheese-ageing inventory platform brings up a new region, and the 7-node cluster there goes live before the observability rollout reaches it. At 02:14 the nightly maturation-schedule recompute drives p99 from 340 ms to 1.9 s for 26 minutes on three of the seven nodes, then recovers. By the time anyone looks in the morning the pods have been rescheduled twice. With metrics alone the team knows *that* it happened and roughly *where*. With continuous profiles they would open the 02:14 window, filter to the three affected instances, and read which code path took the CPU — a question that no longer has an answer once the processes are gone. Without them, the only plan is to wait for a recurrence with a capture standing by, which is a plan that costs a second incident. That asymmetry is the whole case. The cost of always-on is paid continuously and is measurable; the cost of not having it is paid in incidents that stay unexplained. ## How to decide responsibly - **Measure, do not assume.** Run representative load with and without the profiler on a subset — two of those seven nodes — and compare throughput and tail latency, not just mean CPU. - **Enable per profile kind.** CPU sampling continuously is a different decision from allocation or contention profiling; make them separately. - **Set the rate deliberately.** Lower rates cost less and need longer windows to reach a confident conclusion. Choose the tradeoff rather than inheriting a default. - **Plan symbol retention with profile retention.** Profiles you cannot symbolise are storage without value. - **Roll out progressively** and be willing to turn it off for a service where the measurement says the tail cost is unacceptable. ## What interviewers listen for - Leads with the retrospective argument rather than with overhead numbers. - Knows overhead scales with sampling rate and stack depth, not with throughput. - Mentions symbolisation and retention, which almost everybody forgets. - Says how they would measure the cost instead of quoting a figure.

  • How would you convince yourself the always-on profiler's overhead is acceptable before a fleet rollout?
    Measure it rather than quote it. Run the same representative load against nodes with and without collection enabled, and compare throughput and the high latency percentiles, not the mean — stack walking lands unevenly and shows up in the tail first. Repeat per profile kind, since allocation and contention profiling cost far more than CPU sampling, and roll out progressively.
  • Which profile kinds would you hesitate to leave enabled continuously?
    The ones that are not periodic stack samples. Profilers that hook every allocation or every lock acquisition scale their cost with how much the program does, rather than with a fixed interrupt rate, so their overhead grows exactly when the system is busiest. Enable those on demand, or continuously at a heavily reduced rate, and keep CPU sampling as the always-on default.

It is the difference between a security camera that is always recording and one somebody switches on after hearing a noise: the second is cheaper and useless for anything that has already happened.

saying these in an interview costs you the question

  • Assumes profiling overhead is zero because sampling is cheap
  • Would enable every profile kind at full rate everywhere
  • Cannot say what always-on buys that on-demand cannot
  • Thinks a process can be profiled after it has exited
  • Judges overhead by average CPU and ignores the tail
  • Forgets that symbols must be retained alongside profiles