skip to content

What does Go's /debug/pprof/threadcreate profile record, and how do you use it to find what creates threads?

level: middleimportance: nice to knowfreq 22%

answer

  1. keyed on an event, not on time
  2. stacks captured when a thread is born
  3. attribution, not a census
  4. no rate to enable, unlike block and mutex
  5. the total only ever grows

basics

~20 s

Go's threadcreate profile records the stack that led to each new OS thread, aggregated by call site, so it shows which code path keeps producing threads. It counts creation events, not live threads, and is known to under-report.

solid answer

~50 s

`/debug/pprof/threadcreate`, served once you import `net/http/pprof`, samples the call stack at the moment the runtime creates a new OS thread and groups those samples by stack. You read it exactly like any other pprof profile — `go tool pprof http://host:6060/debug/pprof/threadcreate`, or add `?debug=1` for plain text — and what you are looking for is the dominant stack: if it runs through the cgo entry path, blocking C calls are making your threads; if it comes from a package that wraps a blocking system call, that wrapper is. Two caveats matter in practice. It counts creation events, so it only ever grows and it never tells you how many threads are alive now. And it has long been known to under-report, so use it to point at a culprit and confirm the actual count from `runtime/metrics` or the operating system.

code

text · 5 lines
text
$ go tool pprof http://localhost:6060/debug/pprof/threadcreate
(pprof) top

$ curl -s 'http://localhost:6060/debug/pprof/threadcreate?debug=1' | head -1
threadcreate profile: total 1043

go deeper

for a junior

Know that Go serves more than the CPU and heap profiles, and that /debug/pprof/threadcreate exists once net/http/pprof is imported. Being able to fetch it and describe what it is about is enough here.

for a middle

Explain that it samples the stack at each thread creation and aggregates by call site, that there is no rate to enable, and that its total is a count of events rather than of living threads.

for a senior

Demonstrate the workflow: confirm the real thread count elsewhere first, use the profile only for the dominant stack, cross-check against the goroutine profile, and say out loud that this profile is known to under-report.

for a principal

Decide what the service exposes and to whom. Weigh leaving a pprof endpoint reachable in production against the incident time lost without it, and set the expectation that thread attribution is corroborated, never quoted from one profile.

## What the profile is Go's runtime keeps several built-in profiles. `threadcreate` is one of the quieter ones: whenever the runtime allocates a new **M** — its name for an operating-system thread — it records the call stack that led to that creation. Samples with identical stacks are merged and counted, exactly like a memory or CPU profile, so what you get back is a ranked list of "this code path caused N threads to be created". It is exposed three ways: - **Over HTTP.** A blank import of `net/http/pprof` registers `/debug/pprof/threadcreate` on the default mux. Point `go tool pprof` at it, or fetch it with `?debug=1` for a human-readable text form whose first line is `threadcreate profile: total N`. - **On the command line.** `go tool pprof http://host:6060/debug/pprof/threadcreate` gives you the usual `top`, `list` and web views, including the flame graph. - **In code**, through the profile registry in `runtime/pprof`, if you need to dump it from a signal handler or an admin endpoint rather than over HTTP. Unlike the block and mutex profiles, there is **no rate to enable**. There is no `SetThreadCreateProfileRate`; the profile is simply always being collected, at negligible cost, because thread creation is rare compared with anything else the runtime does. ## What you are actually looking for The profile answers one question: *what code path keeps causing the runtime to need a fresh thread?* The runtime creates a thread when it has runnable goroutines and no thread free to run them, so a stack that dominates this profile is almost always a stack where an existing thread got stuck. In practice you see one of three shapes. **Calls into C.** Stacks running through the runtime's cgo call path mean goroutines are sitting inside C code. While a goroutine is inside a C call the runtime cannot take that thread back for other work, so N concurrent C calls means at least N threads, however few goroutines you have. **Blocking system calls.** Anything that blocks in the kernel for a long time — a file operation on a slow or remote filesystem, a `syscall` wrapper a library calls directly, a process wait — holds a thread for the duration. **Ordinary scheduler warm-up.** Early in a process's life the runtime naturally creates a handful of threads. A profile whose total is in the low tens with no single dominant stack is telling you nothing is wrong. ## The two caveats that decide how you use it **It counts creations, not live threads.** Each sample is an event that already happened. The total never goes down, even if every one of those threads is now parked and idle, and it does not decrease when a thread exits. So it cannot answer "how many threads do I have right now" — for that you want the `/sched/threads:threads` gauge from `runtime/metrics`, `GODEBUG=schedtrace=1000`, or the operating system's own count. **It is known to be incomplete.** The threadcreate profile has a long-standing reputation for missing threads and under-reporting; it is the least-maintained of the built-in profiles. Treat it as a *hint about attribution*, corroborated by a real count, and never quote its total as the number of threads in an incident review. An engineer who says "threadcreate showed 40, so we are fine" while the kernel reports 1,200 threads has drawn exactly the wrong conclusion. ## A workflow that works 1. Establish the real thread count from `runtime/metrics` or the operating system, and put it next to `runtime.NumGoroutine()`. Confirm the shape of the problem first: threads climbing while goroutines stay flat. 2. Pull the threadcreate profile and look only at the *top stacks*, ignoring the total. You want the one call path that dominates. 3. Cross-check with the goroutine profile. If threads are being created because goroutines are stuck, those goroutines are parked somewhere identifiable, and the two profiles should tell the same story from opposite ends. 4. Fix at the choke point the profile named — typically by bounding how many of those blocking calls may be in flight at once, so thread creation becomes queueing instead. ## Why not just use the CPU profile Because thread creation is not a CPU-time phenomenon. A thread blocked in C or in a system call burns no Go CPU time and shows up nowhere in a CPU profile, and the goroutine driving it may be a single cheap line of code. The threadcreate profile is the only built-in view that is keyed on the event you care about.

  • Do you need to turn the threadcreate profile on, the way you do for the block or mutex profile?
    No. The block profile needs runtime.SetBlockProfileRate and the mutex profile needs SetMutexProfileFraction, but threadcreate has no rate knob — the runtime always records it, because thread creation is rare enough to be free. All you need is a way to read it, which usually means importing net/http/pprof.
  • The threadcreate total says 40 but the kernel reports 1,200 threads. Which do you trust?
    The kernel. The threadcreate profile is known to under-report and is only a creation log, so a low total does not mean few threads. Trust /sched/threads:threads or the operating system for the count, and use threadcreate only for the ranked stacks that suggest where the threads came from.
  • Which other profile would you pull alongside it, and why?
    The goroutine profile. Thread growth usually means goroutines are stuck somewhere that holds their thread, and the goroutine profile shows exactly where they are parked, grouped by call site. The two views describe the same blocking call from opposite ends and should agree; if they do not, one of your assumptions is wrong.

saying these in an interview costs you the question

  • Reads the threadcreate total as live thread count
  • Thinks a sampling rate must be set first
  • Expects the profile total to fall when threads idle
  • Trusts its total over the kernel's thread count
  • Looks for thread creation in a CPU profile