skip to content

How do you decide whether net/http/pprof ships in a production build, and on which listener?

level: principalimportance: nice to knowfreq 30%

answer

  1. two legitimate parties, one placement
  2. not whether, but on which listener
  3. loopback inherits an existing access story
  4. compiling it out costs you the incident
  5. encode the decision in CI, not in review

basics

~20 s

Treat it as a policy with three settings: always on a loopback or authenticated operator listener, never on a public one, and compiled out only where a debug surface is genuinely unacceptable. Decide once per class of service, not per incident.

solid answer

~50 s

The tension is real: on-call wants live profiles from the process that is actually misbehaving, and the security reviewer does not want a debug surface in production. Ban both extremes. "Never ship it" sounds safe and quietly costs you the only tool that explains a CPU or memory incident on the real workload, pushing teams into redeploys under pressure. "Ship it on the public mux" is an information leak plus an on-demand CPU cost any caller can trigger. My default is a second `http.Server` bound to `127.0.0.1`, reached over a tunnel, with its own timeouts sized for profiling; where operators cannot get host access, an authenticated operator port outside the load balancer's target set. I reserve build-tag gating for services where any debug surface is disallowed, and I accept that it means a redeploy to profile. Then I make the decision enforceable in CI rather than reviewable per PR.

code

go · 16 lines
go
//go:build pprof

package main

import (
	"log"
	"net/http"
	_ "net/http/pprof"
)

// present only in builds made with: go build -tags pprof
func init() {
	go func() {
		log.Println(http.ListenAndServe("127.0.0.1:6060", nil))
	}()
}

go deeper

for a junior

Take away the safe default: profiling endpoints belong on a separate loopback listener, never on the port serving customers, and you pass an explicit mux to the public server.

for a middle

Be able to describe the placements — loopback listener, authenticated operator port, build-tag gating — and the mechanics each needs, such as separate timeouts for the profiling server.

for a senior

Argue the operational cost of the strict option out loud: what an incident looks like when live profiling is unavailable, and what you would put in place to compensate.

for a principal

Own the standard across services: name who decides and who can overrule, choose placement per class of service against what the surface discloses, and make it hold with build-time checks rather than review comments.

## Why this is a decision and not a rule Every party is arguing for something legitimate. The on-call engineer wants a CPU profile from the instance that is degrading, on the real workload, right now — the alternative is guessing, or shipping a build to reproduce, which is slow and changes the thing being measured. The security reviewer sees an unauthenticated surface that returns the process command line and every goroutine stack, and that will burn a profiling window of CPU for anyone who asks. Both are right, so the answer is a placement decision rather than a yes/no. ## The options, in order of how often they are right **1. A separate operator listener on loopback.** A second `http.Server` bound to `127.0.0.1`, serving a mux you built yourself with the pprof handlers registered explicitly. The public server keeps its own mux and its tight timeouts; the operator server can have the long write timeout that a 30-second CPU profile requires. Access is by tunnel or port-forward, which means access is already governed by whatever controls host access — you inherit an existing authorisation story instead of inventing one. This is my default for anything that runs on hosts an operator can reach. **2. An authenticated operator port.** Where operators have no host access, bind the operator server to a port that is not in the load balancer's target set and put real authentication in front of it. This is more moving parts and a new thing to get wrong, so I take it when option 1 is genuinely unavailable rather than as the standard. **3. Compiled out behind a build tag.** Put the import and the listener in a file guarded by `//go:build pprof` so a normal build cannot serve profiles at all. Honest about what it costs: to profile production you must ship a different binary, which is exactly the pressure you do not want during an incident, and the profile you finally get is of a build nobody was running. I take this only where the compliance position is that no debug surface may exist in a production artefact — and then I insist the team has an alternative story for CPU and memory incidents, such as a canary or staging instance carrying representative load with profiling enabled. **4. On the public mux.** Never. It is the two-line accident, not a choice. ## What makes the argument concrete rather than theoretical When the reviewer and the service owner disagree, three questions usually settle it. What does the surface actually reveal for *this* service — are there credentials in the command line, do goroutine stacks carry customer identifiers in argument positions? Who can reach the port, in the real network, not the diagram? And what is the recovery path for a memory or CPU incident if profiling is not available — if the honest answer is "redeploy with profiling and wait for it to recur", the reviewer is buying a slower incident, and that cost belongs in the decision. Be explicit about who can be overruled. The service owner picks the placement; the security reviewer can require a stronger option, and an organisation-wide standard can override both. Writing that down beats re-litigating it per service. ## Make it enforceable, not reviewable A policy that lives in a review comment decays. Two mechanical guards carry it instead: - Fail the build if `net/http/pprof` appears in the release binary's package list where the policy forbids it. `go list -deps ./...` gives you that list for a shell check. - Fail the build, or lint, on any listener started with a nil handler. That is the half of the accident that is invisible in a diff, and an explicit handler is better style anyway. Also worth a line in the runbook: which port profiles live on, how to tunnel to it, and what window to request. A control nobody knows how to use is functionally the same as not shipping it, and you paid for the surface either way. ## The migration case When you inherit a fleet already serving profiles from the public mux, sequence it. Add the operator listener first, prove the on-call path works end to end, and only then take the public one away — otherwise the first incident after the change is the one that discovers profiling is gone, and the change gets reverted at the worst moment. ## What an interviewer is checking That you can hold two legitimate positions at once and produce a placement with an enforcement mechanism, rather than reciting either "pprof is a vulnerability" or "profiling is non-negotiable". The tell of a strong answer is naming what the strict option costs and how you would pay for it.

  • On-call argues that removing pprof makes CPU incidents unresolvable. What do you offer instead of a flat no?
    A loopback operator listener reached by tunnel, which keeps live profiling while inheriting host-access controls; or an authenticated operator port outside the load balancer where host access is impossible. If policy forces compiling it out, I fund the alternative explicitly — a canary instance under representative load with profiling enabled — rather than leaving the gap unstated.
  • How do you keep the decision from regressing six months later?
    Encode both halves mechanically. A CI check on the release build's package list from go list -deps catches net/http/pprof re-entering where it is banned, including through a new dependency. A lint or review rule against a nil handler on any listener catches the other half. Review comments alone do not survive team turnover.
  • What do you actually weigh when deciding how sensitive this surface is for a given service?
    What the endpoints disclose for that specific process — credentials or paths in the command line, customer identifiers visible in goroutine stack arguments, symbol names that reveal unreleased work — plus who can genuinely reach the port on the real network, and the cost of the incident you cannot diagnose without it. Those three answers differ enough between services to justify different placements.

saying these in an interview costs you the question

  • Bans pprof in production with no story for CPU or memory incidents
  • Calls an unusual port or an undocumented path a control
  • Treats the choice as purely technical with no owner named
  • Leaves the policy to PR review with no build-time check
  • Removes the public surface before the operator path is proven