skip to content

What does the Go compiler do differently at a call site that a PGO profile marks hot?

level: middleimportance: must knowfreq 44%

answer

  1. distribution, not new semantics
  2. spend the budget where the time is
  3. a big callee gets inlined at one site
  4. one dominant concrete type behind an interface
  5. guarded direct call, indirect fallback

basics

~20 s

Two things: it inlines hot callees it would otherwise judge too expensive, and it devirtualizes hot indirect calls by emitting a direct call to the dominant concrete type behind a type check, with the original indirect call as the fallback. Cold code is left alone.

solid answer

~50 s

A Go PGO profile tells the compiler which call sites carry real CPU time, and it spends its optimisation budget there instead of spreading it evenly. The first effect is **hot call-site inlining**: a callee the compiler would normally decline to inline gets inlined anyway when the profile shows the call site is hot, which also exposes the inlined body to constant propagation and to escape analysis in the caller's frame. The second is **devirtualization**: for an indirect call — a method called through an interface, or a function value — where the profile shows one concrete implementation dominates, the compiler emits a type check plus a direct call to that implementation, falling back to the original indirect call otherwise. The direct call is then itself an inlining candidate. Everything the profile does not mark hot compiles exactly as before, and the effect applies across the whole build, dependencies and standard library included.

code

go · 17 lines
go
type Store interface {
	Product(ctx context.Context, id string) (Product, error)
}

// s.Product is an indirect call. If the profile shows this site is hot and the
// dynamic type is nearly always *pgStore, the compiler can emit a guarded direct
// call to (*pgStore).Product - which then becomes an inlining candidate.
func handle(s Store) http.HandlerFunc {
	return func(w http.ResponseWriter, r *http.Request) {
		p, err := s.Product(r.Context(), r.PathValue("id"))
		if err != nil {
			http.Error(w, "not found", http.StatusNotFound)
			return
		}
		_ = json.NewEncoder(w).Encode(p)
	}
}

go deeper

for a junior

Know the two headline effects by name: inlining call sites the profile shows are hot, and turning a hot interface call into a direct call to the type that dominates it. Cold code is unchanged.

for a middle

Be able to explain why the two compound — devirtualization creates a static call that inlining can then take — and why the guarded direct call keeps other implementations working. Mention the knock-on effect on escape analysis.

for a senior

Show you know the scope and the cost: the profile shapes code generation across dependencies and the standard library, invalidating build caches, and the realistic payoff is a few percent, not a multiple. Say what you would measure to confirm it.

for a principal

Own the expectation-setting. Argue where a few percent of CPU is worth a build-pipeline complication and where the same effort spent removing work from the hot path returns more, and be clear that PGO never substitutes for a design fix.

## The idea: spend the optimisation budget where the time is Without a profile the compiler has to decide, from the source alone, which calls are worth inlining and which indirect calls might be worth specialising. It cannot tell a function called once at startup from one called ten thousand times a second, so it applies uniform, fairly conservative rules — inlining aggressively everywhere would bloat the binary and hurt instruction cache behaviour for no gain. A PGO profile removes the guesswork for the small fraction of call sites that actually matter. ## Effect one: hot call-site inlining When the profile shows that a particular call site accounts for meaningful CPU time, the compiler is willing to inline the callee there even if it would normally reject it as too large. This is per call site, not per function: the same function can be inlined at the hot caller and left as a real call everywhere else, so the binary does not grow across the board. Inlining matters for more than the saved call instruction. Once the callee's body is in the caller, the compiler can propagate constants into it, fold branches that are provably not taken for that caller, and — importantly in Go — run escape analysis over the combined body, so a value that had to be heap-allocated because it was passed to an opaque call can sometimes stay on the stack. This is why PGO often reduces allocations as a side effect, not just cycles. ## Effect two: devirtualization of hot indirect calls Go code reaches for interfaces constantly, and a call through an interface is an indirect call: the target is loaded from the interface value's method table at run time, so the compiler cannot inline it and the CPU cannot predict it as well as a direct call. If the profile shows a hot interface call site whose receiver is overwhelmingly one concrete type, PGO turns it into a conditional direct call. Conceptually the compiler rewrites ```go row, err := s.Product(ctx, id) // indirect: s is an interface ``` into "if the dynamic type of `s` is `*pgStore`, call `(*pgStore).Product` directly; otherwise do the original indirect call". The guarded direct call preserves semantics exactly — a different implementation still works, it just takes the slow path — while the fast path becomes an ordinary static call that can then be inlined by the pass above. The two optimisations compound: devirtualization is often valuable mainly because it unlocks inlining at a site that could never be inlined before. ## What is left alone Cold code compiles exactly as it did without the profile. This is a deliberate property: PGO is meant to be safe to leave on, so it must not make the ninety-nine percent of the program that is not hot bigger or slower. It also does not change semantics — a `-pgo=off` and a `-pgo=auto` build of the same source are behaviourally identical, which is what makes an A/B comparison of the two binaries a valid experiment. ## Scope of the effect The profile applies to the entire build, not just to your own packages. Dependencies and the standard library are recompiled under the profile's influence too, which is why a hot path that spends its time inside a standard-library function can benefit. The flip side is build cost: because code generation for essentially every package now depends on the profile's content, adding or changing `default.pgo` invalidates cached package builds and triggers a wide rebuild. Later Go releases reduced that overhead considerably, but it remains the reason the first profile-guided CI run feels slow. ## How much to expect The Go project reports typical CPU improvements in the low single digits — commonly quoted as roughly 2 to 7 percent for a representative profile, with more on some workloads. That is a meaningful win for a service where you cannot change the code, and a poor trade if you were hoping for a multiple. The honest framing in an interview is that PGO is the cheapest few percent available — no source change, one file, one flag that is already the default — and that anything larger has to come from changing what the program does, not from how it is compiled. ## The mental model to carry A profile does not tell the compiler anything new about semantics; it tells it about *distribution*. Every optimisation PGO enables is one the compiler already knew how to perform but could not justify performing everywhere. That is why the quality of the profile decides the quality of the result: point it at the wrong workload and the compiler will faithfully optimise the wrong call sites.

  • If devirtualization emits a direct call, what happens when a different implementation shows up at run time?
    Nothing breaks. The compiler emits a type check guarding the direct call, with the original indirect call as the fallback path, so any other implementation is dispatched exactly as before — it just does not get the fast path. That is why the optimisation is safe to apply from a profile that may not describe every caller.
  • Why does PGO sometimes reduce allocations and not just CPU time?
    Because inlining a hot callee merges its body into the caller's frame, and escape analysis then runs over the combined code. A value that had to be heap-allocated because it was handed to an opaque call can now be proven not to outlive the frame and stay on the stack. The allocation reduction is a second-order effect of the inlining decision.
  • Does the profile only affect your own packages?
    No — the profile influences code generation for the whole build, including dependencies and the standard library, so a hot path that spends its time inside a standard-library function can benefit. The cost is that introducing or changing the profile invalidates cached builds for those packages and forces a wide recompilation.
  • How large a win is realistic?
    The Go project reports typical CPU improvements of roughly a few percent — commonly quoted as 2 to 7 percent for a representative profile. It is the cheapest few percent you can get, because it needs no source change, but it is not a substitute for fixing an algorithm or removing work from the hot path.

saying these in an interview costs you the question

  • Claims PGO rewrites the algorithm or removes work
  • Thinks devirtualization drops support for other implementations
  • Says the whole binary is inlined more aggressively
  • Expects a 2x speedup from a profile
  • Believes the profile only affects the main package