How do you split a Go CPU profile's hottest function by tenant using pprof labels?
answer
- label where the work runs
- the tool can filter on tags
- one flag keeps samples, another picks keys
- tagfocus narrows, tagshow displays
- check what is left untagged
basics
~20 sLabel each unit of work with the tenant using runtime/pprof.Do where it executes, collect a CPU profile, then in go tool pprof list the tags and use -tagfocus to keep only one tenant's samples, comparing totals across tenants.
solid answer
~50 sFirst make the samples carry the tenant: `pprof.Do(ctx, pprof.Labels("tenant", t), …)` around each unit of work, applied by the goroutine that actually runs it, not where the executor was started. Then collect a CPU profile over a representative window. In `go tool pprof`, the interactive `tags` command lists the label keys and values present with their sample weight — that alone often answers the question, since it shows how the profile's total time divides between tenants. To go deeper, `-tagfocus` restricts the profile to samples whose tags match a regular expression, so you can look at one tenant's stacks in isolation; `-tagignore` drops matching samples, and `-tagshow`/`-taghide` control which tag keys are displayed rather than filtering samples. Sanity-check the untagged remainder: samples with no tenant tag mean work running on goroutines that never had labels applied.
code
text · 11 lines# interactive: 'tags' lists label keys and values with their sample weight
go tool pprof cpu.pprof
# keep only samples whose tags match the regexp
go tool pprof -tagfocus=acme cpu.pprof
# drop that tenant to see what the rest of the traffic looks like
go tool pprof -tagignore=acme cpu.pprof
# consider and display only the tenant key, without dropping samples
go tool pprof -tagshow=tenant cpu.pprofgo deeper
Know the two halves: pprof.Do puts a tenant label on the work, and go tool pprof can then filter the profile by that tag instead of showing one combined total.
Be able to name the tag flags and say what each does — tagfocus keeps matching samples, tagignore drops them, tagshow and taghide only change which keys are displayed — and distinguish them from -focus on function names.
Show the discipline around the numbers: profile a representative window, account for the untagged remainder, normalise per-tenant CPU by the work submitted, and recognise inherited labels on long-lived goroutines as an artefact rather than a finding.
Decide what the attribution is for and what it may be used for: a sampled CPU profile supports capacity and optimisation arguments, and saying so before anyone builds a billing or chargeback story on it is part of owning the measurement.
## The question this answers A multi-tenant process — one binary running work on behalf of many customers — produces profiles that are correct and unactionable. The top entry says a decode function is 40% of CPU. Every tenant goes through it. You cannot bill for it, cap it, or decide whether it is worth optimising, because you do not know whether it is one pathological customer or the shape of all traffic. Labels plus the tag filters in `go tool pprof` are how that single number becomes a breakdown. ## Step 1: label where the work runs The attribution is only as good as the labelling. Wrap each unit of work in `pprof.Do` with the tenant as a label value, applied inside the goroutine that executes the work: ```go pprof.Do(ctx, pprof.Labels("tenant", j.Tenant, "kind", j.Kind), func(context.Context) { execute(j) }) ``` Two failure modes to avoid up front. Labelling where the executor goroutines are *created* labels nothing useful, because a goroutine's labels are a snapshot taken when it starts and at that moment there is no tenant. And a goroutine spawned inside a labelled region keeps those labels for its whole life, so background work started during one tenant's job will be charged to that tenant forever. Pick a second key alongside the tenant if you can — job kind, priority, endpoint. Cross-tabulating two low-cardinality keys is far more informative than one, and costs nothing extra. What you must not do is use a per-unit identifier such as a job id: each sample then carries a unique tag, the profile bloats, and nothing aggregates. ## Step 2: collect a profile over a window that matters A CPU profile is a statistical sample over a window. Attribution conclusions are only as representative as the window: profile during the interval you are actually arguing about — the peak, the incident, the batch run — and long enough that the smaller tenants have enough samples to be more than noise. A tenant with twelve samples out of ten thousand is a rounding error, not a finding. ## Step 3: filter and group in the tool Inside `go tool pprof`, the interactive `tags` command lists the label keys and values in the profile with the sample weight behind each. That is usually the first and most valuable view: it turns "40% in decode" into a table of tenants by CPU time. The flags that operate on tags are: - **`-tagfocus`** — restrict the profile to samples whose tags match a regular expression. This is the one that gives you a per-tenant view: focus on a tenant and everything downstream (the totals, the flame graph, the per-function breakdown) describes that tenant alone. - **`-tagignore`** — the inverse: drop samples whose tags match. Useful for setting aside a known-heavy tenant to see what the rest of the traffic looks like. - **`-tagshow`** — restrict which tag *keys* are considered and displayed. It does not drop samples; it reduces the noise when a profile carries several keys and you only care about one. - **`-taghide`** — the inverse of `-tagshow`, hiding matching keys. The usual mistake is reaching for `-focus`, which matches *function names in the stack*, when you meant `-tagfocus`, which matches *tags*. They are different axes: one narrows to a part of the call graph, the other narrows to a subset of the work. ## Step 4: read the result honestly Three checks before you take a number to anyone: 1. **Account for the untagged remainder.** Sum the per-tenant weights and compare against the profile total. A large untagged share is not "other" — it is work running on goroutines that never had labels applied, and until you find out which, every per-tenant figure is a lower bound. 2. **Normalise before comparing tenants.** Raw CPU time tracks volume. If the question is whether a tenant is *inefficient* rather than *large*, divide by the work they submitted; otherwise the biggest customer always looks like the problem. 3. **Watch for inherited labels on long-lived goroutines.** A tenant that improbably dominates background work such as flushing or compaction is usually an artefact of a goroutine that inherited their tag at creation and kept it. ## The limits Only the CPU profile and the goroutine profile carry labels, so this technique gives you CPU time and goroutine counts per tenant and nothing else. Per-tenant memory or lock contention needs different instrumentation. And a CPU profile remains sampled: it supports statements like "roughly two thirds of the decode cost is this tenant", not exact per-customer accounting. If the output is going to a billing conversation, say which of those two you have.
- What is the difference between -focus and -tagfocus in go tool pprof?`-focus` matches function names in the call stack, narrowing the profile to a region of the call graph. `-tagfocus` matches the sample's tags, narrowing to a subset of the work regardless of which functions it ran. They are different axes and are often combined: focus on the decode subtree, then tag-focus a tenant, to answer "how much of decode is this customer".
- You -tagfocus a tenant and get an empty profile. What do you check first?Whether the labels were applied by the goroutine that runs the work. Labels are inherited only at goroutine creation, so executors started at process boot carry an empty set forever and their samples have no tags to match. Also confirm the profile type carries labels at all — a heap profile never does — and that the label value spelling matches what your code sets.
- Why is labelling each unit of work with its unique job id a bad idea?Labels are a grouping dimension. With a unique value per unit of work, almost every sample carries a distinct tag: the profile grows, the tool slows, and no group has enough samples to mean anything statistically. Keep values low-cardinality — tenant, job kind, priority — and use a different instrument if you need per-request identity.
- A tenant has the largest CPU share in the profile. Is that enough to call them the problem?No. Raw CPU time tracks volume, so the biggest customer usually has the biggest share legitimately. Normalise by the work they submitted before claiming inefficiency, account for the untagged remainder so the shares actually sum, and remember the profile is sampled — it supports "roughly two thirds of decode is this tenant", not exact per-customer accounting.
saying these in an interview costs you the question
- Uses -focus expecting it to filter by tag value
- Reports per-tenant shares without accounting for untagged samples
- Calls the largest tenant inefficient without normalising by volume
- Labels the executor at startup and blames the tool for empty output
- Presents a sampled profile as exact per-customer billing data