skip to content

Why does /debug/pprof/profile?seconds=30 block for thirty seconds while /debug/pprof/heap answers at once?

level: middleimportance: should knowfreq 42%

answer

  1. two families of endpoint, not one
  2. some sample over time, some dump state
  3. the window is the request duration
  4. does the window fit the server's WriteTimeout?

basics

~20 s

The /debug/pprof/profile endpoint turns CPU profiling on, waits out the requested window (30 seconds by default) and then streams the samples, so the request lasts the whole window. The /debug/pprof/heap endpoint needs no window and dumps current allocation records immediately.

solid answer

~40 s

`/debug/pprof/profile` is a *window* endpoint: the handler starts CPU sampling, sleeps for `seconds` (default 30 if the parameter is missing or unparseable), stops, and writes the collected profile as the response body. The request necessarily takes that long, and `go tool pprof http://host:6060/debug/pprof/profile?seconds=30` just sits there fetching it. `/debug/pprof/heap` is a *snapshot* endpoint — it serialises the allocation records the runtime already holds and returns straight away. Two consequences bite in practice. First, the requested duration must fit inside the serving `http.Server`'s `WriteTimeout`, or the handler rejects the request immediately rather than producing a truncated body. Second, only one CPU profile can be active in a process, so a second concurrent request for it fails. Passing `seconds` to a snapshot endpoint instead asks for a delta: two snapshots that far apart, subtracted.

code

text · 5 lines
text
# blocks for the whole window, then opens the pprof prompt on a 30s CPU profile
go tool pprof http://localhost:6060/debug/pprof/profile?seconds=30

# returns straight away: the allocation records the runtime already holds
go tool pprof http://localhost:6060/debug/pprof/heap

go deeper

for a junior

Know that /debug/pprof/profile takes time on purpose: it watches the process for the requested number of seconds, defaulting to 30, while snapshot endpoints answer immediately.

for a middle

Explain the window-versus-snapshot split, what the seconds parameter means on each kind of endpoint, and why the window has to fit inside the serving http.Server's WriteTimeout.

for a senior

Show operational judgment: sizing the window to the symptom, knowing a second concurrent collection fails, and insisting the URL address one instance rather than a balanced hostname.

for a principal

Own the operating envelope — whether profiling endpoints get their own http.Server with its own timeouts, and what limits stop an automated collector from holding profiling permanently open on a hot service.

## Two kinds of endpoint The handlers registered under `/debug/pprof/` fall into two families, and confusing them is the source of most surprises. **Window endpoints** must observe the program while it runs. `/debug/pprof/profile` is the CPU profile: there is no such thing as "the CPU profile right now", because a CPU profile is a set of samples accumulated over time. The handler therefore switches sampling on, blocks for the requested duration, switches it off, and writes the result. `/debug/pprof/trace` behaves the same way — it records events over a window. **Snapshot endpoints** serve state the runtime is already maintaining. `/debug/pprof/heap`, `/debug/pprof/goroutine`, `/debug/pprof/allocs` and friends are dumps of bookkeeping that exists whether or not anyone asks, so the handler can serialise and return immediately. ## The seconds parameter `seconds` is a query parameter, not part of the path: `/debug/pprof/profile?seconds=30`. If it is absent, zero, negative or unparseable, the CPU profile handler uses 30 seconds. Choosing it is a judgment call about the symptom: a window shorter than the phenomenon you are chasing may miss it entirely, while a very long one holds a request open and averages a spike away into the background. On a snapshot endpoint the same parameter means something different. `/debug/pprof/heap?seconds=10` does not sample for ten seconds; it takes a snapshot, waits, takes a second one and returns the difference. That delta form is often what you actually want when you are asking "what changed while the load was running" rather than "what does the process hold right now". ## The WriteTimeout trap A production `http.Server` usually sets `WriteTimeout` to something in the seconds range. `WriteTimeout` bounds how long the server will spend writing a response, and a 30-second CPU profile blows straight through a 15-second budget. Rather than letting you collect a body that gets cut off mid-profile — which yields an unparseable file and a confusing error in the analysis tool — the CPU profile handler compares the requested duration against the serving server's write deadline and refuses up front with a client error explaining that the profile duration exceeds the server's `WriteTimeout`. This is a strong argument for the second-listener pattern: the operator-facing `http.Server` can have long or absent write timeouts precisely because it serves profiles, while the public one keeps the tight timeouts it needs. Trying to serve both from one `http.Server` forces you to pick a `WriteTimeout` that is wrong for one of them. ## One CPU profile at a time CPU profiling is a process-wide facility, so only one collection can be in flight. If a second client asks for `/debug/pprof/profile` while the first window is still running, it does not get a second profile and it does not queue politely behind the first — the request fails with an error saying CPU profiling is already in use. This surprises teams who put the endpoint behind something that retries, or who have a dashboard scraping it on a timer while a human is also debugging. ## Fetching it across the network The toolchain speaks these URLs natively: ``` go tool pprof http://localhost:6060/debug/pprof/profile?seconds=30 ``` It performs the HTTP fetch, waits out the window, and drops you at its prompt with the downloaded profile. Because the profile is fetched over HTTP, the request lands on exactly one process. Point that URL at a load-balanced address and you profile whichever instance the balancer happened to pick — likely not the one that is misbehaving, and not the same one twice. When you need a specific instance, reach it directly: a tunnel or port-forward to that instance's operator listener, then the same command against `localhost`. If you would rather keep the artefact, fetch it with any HTTP client and save the body — the response is a gzipped protobuf profile that the tool reads from a file just as happily as from a URL. That is also the practical answer when the window is long and you do not want to hold a terminal open. ## What an interviewer is checking That you understand "this endpoint costs wall-clock time and CPU on a live process" rather than treating every path under `/debug/pprof/` as a free read. The follow-ons — why the window has to fit the server's write timeout, why concurrent collection fails, why the URL must address one instance — all fall out of that.

  • What happens if two clients request /debug/pprof/profile at the same time?
    The second one fails. CPU profiling is process-wide and only one collection can be active, so the handler returns an error saying profiling is already in use rather than starting a second window. It is a real hazard when a scraper polls the endpoint on a timer and a human debugs at the same time.
  • Your production http.Server sets WriteTimeout to 10s. What does a seconds=30 request to /debug/pprof/profile return?
    An immediate client error saying the profile duration exceeds the server's WriteTimeout — not a truncated profile. The handler compares the requested window against the write deadline first. The clean fix is a separate operator http.Server whose timeouts are chosen for profiling, not for public traffic.
  • Why is pointing go tool pprof at a load-balanced hostname a mistake?
    The fetch is one ordinary HTTP request, so it profiles whichever instance the balancer routes it to — probably not the unhealthy one, and not the same one on a second attempt. Reach the specific instance through a tunnel or port-forward to its operator listener and fetch from localhost.

saying these in an interview costs you the question

  • Thinks seconds=30 returns the last thirty seconds of already-collected samples
  • Expects a truncated profile instead of an upfront rejection when WriteTimeout is short
  • Believes seconds sets a sampling rate rather than a window
  • Assumes several clients can collect CPU profiles concurrently
  • Profiles a load-balanced hostname and treats the result as one instance