skip to content

How do you choose the sample rate for continuous Python stack sampling in production?

level: principalimportance: should knowfreq 24%

answer

  1. Two budgets pull against each other
  2. Cost is rate times stack walk
  3. Count samples per unit of work
  4. Always-on is broad, attach is deep
  5. Stack cardinality drives storage cost

basics

~20 s

Work backwards from the decisions the profile must support. Overhead is sample rate times stack-walk cost, so pick the lowest rate that still lands enough samples in each interesting code path, then measure that cost rather than guessing it.

solid answer

~50 s

Treat it as a budget, not a preference. Overhead is roughly `rate x cost_per_sample`, and cost per sample grows with stack depth and the number of threads walked — a few microseconds for a shallow stack, more for a deep one, so a hundred hertz is well under a percent of one core. Measure it on your own workload. Then check the other side: how many samples land inside the thing you need to see? At 100 Hz a 20 ms request contributes about two samples, so per-request answers are hopeless and only fleet-wide aggregation over minutes is meaningful. Rate buys resolution on short work; continuous profiling buys breadth instead. Decide the clock too — wall-clock for latency, CPU for saturation — then run always-on at a low rate and escalate to a high rate on one instance when you have a real question.

code

python · 19 lines
python
import sys
import threading
import time
import traceback


def cost_of_one_sample(reps=2000):
    tid = threading.get_ident()
    start = time.perf_counter()
    for _ in range(reps):
        frame = sys._current_frames()[tid]
        traceback.extract_stack(frame)
    return (time.perf_counter() - start) / reps


per_sample = cost_of_one_sample()
print(f"{per_sample * 1e6:.1f} us per stack sample")
for hz in (10, 100, 1000):
    print(f"{hz:>5} Hz -> {per_sample * hz * 100:.3f}% of one core")

go deeper

for a junior

The takeaway is that a profiler's sample rate is a dial with a cost on one side and detail on the other, and that production profiling normally runs slowly on purpose.

for a middle

Be able to state the overhead model — rate times stack-walk cost — and work out how many samples a request of a given duration actually produces at a given rate.

for a senior

Show that you would measure the cost on the real workload, separate the always-on rate from an investigation rate, and recognise which blind spots no rate can close.

for a principal

Own the whole capability: the standing budget, the escalation path, storage and cardinality, tool dependency across interpreter upgrades, and who is permitted to attach to a production process and under what audit.

### Frame it as two opposing budgets Sample rate sits between a cost budget and a statistics budget, and a defensible answer prices both. **The cost side.** Overhead is `rate x cost_per_sample x threads_walked`. Cost per sample is a stack walk, roughly linear in stack depth, plus whatever aggregation and storage the sampler does. This is measurable in minutes on your own workload rather than argued about: time a few thousand stack captures, multiply by the rate. A few microseconds per sample for a shallow stack and tens for a deep one means 100 Hz costs a small fraction of one core; the same sampler at 1,000 Hz costs ten times that, and on a deep framework stack with many threads it stops being free. An out-of-process sampler shifts most of that cost off the target, which is one of the strongest arguments for that design in an always-on deployment. **The statistics side.** Samples per unit of interesting work is what determines whether the profile can answer anything. At 100 Hz, a request that takes 20 ms yields about two samples — useless individually, fine once 10,000 requests are aggregated. At 1,000 Hz the same request yields twenty, which starts to make a single trace readable. So the question "what rate?" is really "what is the shortest unit of work I need resolution on, and am I answering per-request or per-fleet?" ### Continuous versus on-demand These are different products with different rates, and conflating them is the usual mistake. - **Always-on, low rate, broad.** Tens of hertz across every process, aggregated over minutes. It answers "where did the fleet's CPU go this week" and "what changed after the deploy". It cannot answer "why was *this* request slow". - **On-demand, high rate, narrow.** Hundreds to a thousand hertz, one instance, for a few minutes, attached during an investigation. It answers the specific question and then goes away. The mature posture is both: cheap continuous coverage that tells you *where* to look, plus a rehearsed, privileged path to attach hard to one instance when it tells you. ### The costs that are not CPU Rate also drives data volume, and this is where continuous profiling quietly gets expensive. Each sample is a full stack; the number of *distinct* stacks — the cardinality — is what storage and query cost track, and it explodes with deep frameworks, dynamic dispatch and per-request identifiers leaking into frame data. Ten times the rate is ten times the ingestion, not ten times the insight, because the extra samples mostly reinforce stacks you already had. Aggregating at the agent, truncating very deep stacks, and keeping retention short for raw profiles while keeping longer aggregates are the levers. ### What the rate cannot fix Some blind spots do not close by sampling faster: - **Rare paths.** A slow function on one request in ten thousand will not appear at any rate you can afford; that is a tracing question, not a profiling one. - **Native frames.** Faster sampling of a boundary you cannot see through yields more samples of the same boundary. - **The clock choice.** A CPU-time profile at any rate will not show a worker that is waiting. Picking wall-clock or CPU time shapes conclusions far more than the rate does. - **Sampler starvation.** An in-process sampler thread needs the GIL; raising its rate while a C extension holds the GIL for long stretches just queues up delayed samples and skews their timing. ### The organisational half of the decision Continuous profiling is a standing capability, so the choices that outlast the rate are governance ones. Who may attach to a production process, and is it audited? Is interpreter-level remote execution permitted at all in your environment, or must sampling be read-only? Do frame names or file paths in stored stacks leak anything sensitive? What is the documented escalation from the always-on view to a high-rate attach, and has anyone rehearsed it? On CPython 3.14 there is a further question worth settling deliberately: the standard library ships no sampling profiler, so this is a dependency on external tooling, and the free-threaded build changes the threading picture the sampler must interpret. Committing to a tool is committing to its support burden across interpreter upgrades. ### How to answer Give the overhead formula, give the samples-per-unit-of-work test, separate the always-on rate from the investigation rate, and name a cost that is not CPU — usually storage cardinality. Finish on the governance point. There is no single right number, and an interviewer at this level is listening for whether you know which quantities the number falls out of.

  • Why does raising the sample rate not help find a slow path that occurs on one request in ten thousand?
    Sampling is proportional: a path that consumes a negligible share of total time collects a negligible share of samples however fast you sample, and its few hits are indistinguishable from noise. Rare-but-slow is a tracing or logging problem — capture the outlier request end to end — not a profiling one. Profilers find what is common; tracing finds what is exceptional.
  • What limits how long you keep raw sampled stacks from a fleet?
    Cardinality and volume. Storage and query cost track the number of distinct stacks far more than the number of samples, and deep framework stacks with dynamic dispatch generate many. The usual answer is to aggregate at the agent, truncate very deep stacks, keep raw profiles for days and rolled-up aggregates for months, and treat frame data as potentially sensitive when deciding retention.
  • How do you justify the overhead of always-on profiling to a team that resists it?
    Measure it rather than argue it: time a stack walk on the real workload, multiply by the rate, and present the number as a fraction of one core alongside the cost of the last incident that was diagnosed by guesswork. Offer the low-rate always-on tier as the default and the high-rate attach as an opt-in during investigations, so the standing cost is small and bounded.

It is the same call as choosing a security camera's frame rate: high enough to see what you actually need to identify, low enough that you can afford to keep the footage from every camera all year.

saying these in an interview costs you the question

  • Picks a rate by habit without measuring overhead
  • Expects per-request detail from fleet-wide sampling
  • Ignores storage and stack cardinality costs
  • Believes a faster rate reveals rare slow paths
  • Runs the same high rate everywhere permanently
  • Never decides who may attach to production

context