Who owns refreshing a Go service's committed default.pgo, and when would you ban the committed profile outright?
answer
- a second input to the build
- an unreviewable blob in the diff
- rot is silent, never an error
- who signs for the refresh cadence
- or fetch it by version instead of committing
basics
~20 sTreat default.pgo as a build input with a named owner, a documented capture source and a refresh cadence tied to releases. Ban it when nobody will own the refresh, when its provenance cannot be audited, or when the build-time and reproducibility cost outweighs a few percent of CPU.
solid answer
~50 sThe decision is about ownership, not about compilers. A committed `default.pgo` is an opaque binary blob in the source tree that changes the machine code of every package in the build, so it needs the same treatment as a pinned dependency: a named owner, a documented capture procedure (which environment, which traffic window, merged from how many instances), and a refresh cadence — most sensibly an automated step in the release pipeline rather than a human chore. I would ban the committed artifact when any of that is missing: nobody signs up to refresh it, so it rots and quietly costs build time for nothing; or its provenance is unreviewable, so a profile from a load test or a developer laptop can be committed with a green diff. I would also decline it where the build must be trivially reproducible from source alone, or where the wide rebuild it forces hurts a CI budget more than a few percent of production CPU helps.
go deeper
You are not expected to set this policy, but know that the profile is a real build input: it is committed, it changes the emitted code, and someone has to keep it current.
Be able to describe the practical costs — an unreviewable binary in the diff, invalidated build caches, a binary that depends on a past production run — rather than only the upside.
Show you would measure the win on your own service before adopting it, and propose the automated capture-and-refresh step instead of relying on anyone's memory.
Own the trade explicitly: name who refreshes it and how often, how provenance is audited, what the few percent is actually worth for this fleet, and the conditions under which you would decline the committed artifact or move it to a versioned artifact store instead.
## Why this is a policy question at all Enabling PGO is technically trivial — one file, and the flag that reads it is already the default. That triviality is what makes it a governance problem. A `default.pgo` checked into the repository: - is a **binary artifact** nobody can review in a pull request; the diff shows a size change and nothing else, - **changes the emitted machine code of the whole build**, including dependencies and the standard library, - **invalidates build caches** when it changes, so a commit touching no Go source can trigger a full rebuild, - makes the binary a function of *two* inputs — the source and a recording of a past production run — so the same commit built before and after a profile refresh produces different machine code. None of those is disqualifying. All of them are things a release pipeline owner has a legitimate right to be consulted about, and a legitimate right to overrule a service team on. ## The three questions that decide it **Who refreshes it, and when?** The only sustainable answer is "the pipeline, automatically". A profile refreshed by whoever remembers is a profile that is three releases old within a year — the decay is invisible, since a stale profile produces no error and no warning, only a slow evaporation of the benefit. A workable policy is: capture from production on a fixed cadence, merge captures across instances, open the refresh as its own commit, and require that commit to be evidence-backed rather than routine-rubber-stamped. **Where did the profile come from?** This is the provenance question, and it is the one most teams skip. A profile is committed as bytes; nothing in review distinguishes a merged capture from three production instances at peak from a capture off one developer's laptop running a synthetic load. Both build fine, and the second one teaches the compiler to optimise a workload the service does not have. The mitigation is procedural: the refresh is produced by an automated job, not by hand, and the commit records which environment and window it came from. **What is a few percent worth here?** A profile-guided build typically returns single-digit CPU improvements. For a fleet large enough that a few percent of CPU is real money, or a service running close to a capacity ceiling you cannot raise, that pays for the process. For a service running three replicas well below their limits, it does not — and the honest call is to decline PGO and spend the attention on something with a larger return. Refusing a real optimisation because its payoff is smaller than its carrying cost is a legitimate engineering decision, and being able to say so plainly is the point of the question. ## When I would ban the committed artifact - **No owner.** If no team will sign for the refresh, do not commit the file. A rotted profile is worse than none, because it costs build time and review confusion while returning nothing. - **Unauditable provenance.** If profiles are captured by hand and pasted in, the artifact is unreviewable and eventually wrong. Either automate the capture or drop it. - **Reproducibility requirements that the second input breaks.** Some release processes need a binary to be a pure function of the source tree at a tag, verifiable by an independent rebuild. A committed profile is still deterministic — the same source plus the same profile yields the same binary — but it adds an input that has to be preserved and shipped with the source to make that claim hold. If your process cannot carry that, say no. - **CI economics.** If profile refreshes force wide rebuilds on a constrained runner fleet often enough to slow every team down, the few percent is not worth it. ## The middle path worth proposing The alternative to committing is `-pgo=<path>`: the release build fetches the profile from an artifact store, keyed by version, and the source tree stays free of binaries. That preserves auditability and reproducibility — the profile is versioned and addressable — at the cost of developers no longer getting the same binary from a plain `go build`. Which side of that trade you land on depends on whether you care more about the source tree being self-contained or about the build being byte-reproducible from source alone. A principal-level answer names both options and the criterion that chooses between them, rather than asserting one is correct. ## What I would ask for before agreeing Before adopting PGO across a fleet I would want: a measured A/B result on the actual service rather than a quoted industry number, an automated capture and refresh job, a documented capture source, and an agreed cadence with a named owning team. If those four exist, committing the profile is cheap and defensible. If two of them are aspirational, the feature is not ready to be a policy — and "we will add the automation later" is the specific promise that turns into a three-release-old profile nobody remembers adding.
- Is a build using a committed profile reproducible?Deterministically, yes: the same source plus the same profile produces the same binary. What changes is that the binary is now a function of two inputs, so an independent rebuild must obtain the exact profile too. If your release process promises reproducibility from a source tag alone, the profile has to be carried and versioned as part of that tag.
- What is the alternative to committing the profile into the repository?Have the release build fetch a versioned profile from an artifact store and pass it with `-pgo=<path>`. That keeps binaries out of the source tree and makes provenance addressable, at the cost of a plain local `go build` no longer producing the same binary as the release. Choose based on whether self-contained source or byte-identical local builds matters more.
- How would you keep an unrepresentative profile from being committed?Make the capture a pipeline job rather than a human step: profile production instances on a schedule, merge the captures, and open the refresh as an automated commit that records the environment and window it came from. Once the artifact is produced by hand, review cannot tell a production capture from a laptop load test.
- A team asks to adopt PGO fleet-wide. What do you require first?A measured A/B on their own service rather than a quoted number, an automated capture and refresh job, documented provenance, and a named owning team with a cadence. Without those, the few percent evaporates within a year and the organisation is left carrying build-time cost and an artifact nobody understands.
saying these in an interview costs you the question
- Treats enabling PGO as purely a compiler flag decision
- Assumes someone will remember to refresh the profile
- Cannot say which environment the profile was captured from
- Ignores that the profile is a second, unreviewable build input
- Adopts fleet-wide on a quoted percentage, with no local measurement