What does a sampling profiler record, and how does a profile differ from a trace or a metric?
answer
- Interrupted on a timer, not on every call
- One capture is a stack sample
- Identical stacks merged into counts
- Keyed by code, not by request
- Statistical over a window, not exhaustive
basics
~20 sA sampling profiler interrupts the program at a fixed rate and records the call stack at that instant, merging identical stacks into counts. The result attributes a resource — CPU time or allocated bytes — to code paths rather than to requests.
solid answer
~50 sA sampling profiler arranges to be interrupted on a timer and, at each interrupt, walks the stack of the code that was running and stores the sequence of frames. That single record is a **stack sample**; identical stacks are merged and counted, so the stored profile is a tree of distinct stacks with sample counts, covering a window of time. That makes a profile different in kind from the other signals by what it is *keyed by*. A metric is keyed by a name plus a few dimensions and answers *how much* or *how often*. A trace is keyed by a request and answers *where that request spent its time and on which hop*. A profile is keyed by the **call stack**, and answers *which code consumed the resource*, aggregated across everything running in the window. It is also statistical, not exhaustive — it never claims to have seen every call.
code
pseudocode · 6 linesevery 1/rate seconds:
stack = walk_stack_of_running_thread()
counts[stack] += 1
# stored profile = the counts map, not the individual captures
# counts[["main","handleUpload","parseCsv"]] == 4127go deeper
Recall the mechanism in one sentence: interrupt on a timer, record the call stack, count identical stacks. Be able to say the result is about code paths rather than about individual requests.
Explain why overhead depends on the interrupt rate and stack depth rather than on throughput, name a profile other than CPU, and describe how merging identical stacks keeps the stored artifact small.
Show judgement about when a profile is the right reach — CPU burn inside one service, allocation driving pauses, a regression invisible at endpoint level — and be honest about how many samples a conclusion needs.
Own how profiles fit the signal portfolio: what you fund continuously versus on demand, how profiles are correlated with the other signals without pretending they are per-request, and what you stop collecting because nobody queries it.
## What the profiler is actually doing A sampling profiler does **not** observe every function call. It arranges to be interrupted periodically — by a timer signal, a hardware performance-counter overflow, or a hook inside the language runtime — and at each interrupt it walks the stack of the code that was executing, recording the sequence of frames from the entry point down to the function currently running. That one record is a **stack sample**. Nothing else is captured: no arguments, no wall-clock timeline, and by default no request identity. Two consequences follow immediately. - The profiler's cost tracks the **interrupt rate and the stack depth it must walk**, not how much work the program does. Doubling throughput does not double profiling overhead. - The output is **statistical**. A function consuming 40% of the CPU will be on the stack in roughly 40% of samples, with an error that shrinks as samples accumulate. Before storage, identical stacks are merged and counted, so the artifact is not a list of thousands of captures but a tree: each distinct stack with the number of samples that landed on it. That merging is why profiles stay small even over long windows. ## Which resource is being sampled 'Profile' is a family, not one thing. What varies is the resource and the trigger; the shape — stack, then a count — stays the same. - **CPU profile** — on-CPU time, triggered on a CPU timer. - **Allocation profile** — bytes or objects allocated, triggered every so many bytes allocated. - **Live-heap profile** — what is currently retained, rather than what was ever allocated. - **Lock-contention or blocking profile** — time spent waiting, captured at the blocking primitive. - **Off-CPU or wall-clock profile** — where a thread waited rather than ran. A candidate who can only name CPU has usually never chased a memory or contention problem. ## How a profile differs in kind | Signal | One data point is | Keyed by | Answers | |---|---|---|---| | Metric | a number at a timestamp | a name plus a few dimensions | how much, how often, is it trending | | Trace | a span for one operation | a request, service and operation | where *this* request spent its time | | Profile | a count of stack samples | the call stack itself | which code consumed the resource, over a window | The keying is the whole point. A profile is the only common signal whose primary key is **code**. Three things follow: 1. A trace tells you which service and which operation was slow; a profile tells you which **function**. A span reading 1.4 s on one service is where the trace's resolution ends unless somebody instrumented deeper. A CPU profile over the same window can say that most of it went into recompiling the same regular expression. 2. A profile sees code **nobody instrumented** — framework internals, serialization, garbage collection, third-party libraries, the runtime itself. Instrumentation only ever shows you what someone chose to mark. 3. A profile is normally **window-scoped and aggregate**, not per-request. Some profilers can attach a small set of dimensions to each captured stack so a profile can later be filtered, but the default artifact describes a process or a fleet over an interval, not one user's request. ## The shape of question only a profile answers 1. Where is the CPU going in a service that is busy but not obviously slow at any one endpoint? 2. Which code path allocates the garbage that is driving collection pauses? 3. What became more expensive between two builds, when both look identical at the level of endpoints and error rates? 4. Which library — often one nobody instrumented and nobody suspected — accounts for a fifth of the fleet's CPU bill? Notice that none of these is phrased around a single request. When the question is 'why was *this* request slow', a trace is the right tool; when it is 'what is this code costing us', it is a profile. ## Statistical error, stated honestly A short capture collects few samples. The hottest frame is usually solid; the third and fourth entries are frequently noise, and two captures of the same steady workload will disagree about them. Worse, a genuinely expensive code path that runs rarely can be missed entirely by a short window — the profiler was simply never interrupted while it was on the stack. Longer windows, or continuous collection, are the fix. Treating one thirty-second capture as authoritative about the whole service is a common junior mistake. ## What interviewers listen for - Says *samples the stack periodically*, not *records every call*. - Can name a profile other than CPU and say what it is triggered by. - States the keying difference against traces and metrics rather than describing profiles as 'more detailed traces'. - Treats the numbers as statistical, with a sense of how much data a conclusion needs.
- A span shows one service took 1.4 seconds. What can a profile add that the trace cannot?The trace stops at the boundary of what somebody instrumented, so it says only that this operation on this service was slow. A profile over the same window attributes the resource to functions, including framework, serialization, garbage-collection and third-party frames nobody marked up, so it can point at the specific code burning the time.
- Why might two profiles of the same steady workload disagree about the third-hottest function?Because sampling is statistical. Each entry's share is estimated from a finite number of stack samples, and the error is largest for frames with small shares. The top entry is usually stable, while the tail reshuffles between captures. Collecting over a longer window, or continuously, tightens the estimate.
It is a periodic photograph of what the program is doing rather than a log of everything it did: take enough photographs and the subject that appears most often is the one taking up the time.
saying these in an interview costs you the question
- Says a profiler records every function call it observes
- Assumes a profile can be attributed to one request by default
- Describes a profile as a more detailed distributed trace
- Treats a very short capture as authoritative about everything
- Cannot name a resource other than CPU that profiles measure