Should every service build under GOFIPS140 and run with fips140=only, or only the regulated one?
answer
- scope is not yours to set alone
- strict mode across shared libraries
- two artefacts cost real money
- pin the snapshot, never track latest
- gate the shared libraries in CI
basics
~20 sScope it to the audited boundary. Build the regulated services with a pinned GOFIPS140 snapshot and run those with fips140=only, leave the rest off, and gate shared libraries with a CI job under strict mode so common code stays compatible.
solid answer
~50 sI scope the mode to whatever the compliance owner defines as the audited boundary, because that boundary is their call and over-applying it costs real outages. Turning strict mode on fleet-wide converts every non-security MD5 or SHA-1 in every shared library into a run-time failure in a service nobody was auditing, and those surface on a request rather than in a build. So the regulated binaries build with a pinned validated snapshot such as `GOFIPS140=v1.0.0` — never `latest`, which compiles the toolchain's live source and drifts with each upgrade — and run under `GODEBUG=fips140=only` behind a canary. Everything else builds `off`. The bridge is a required CI job that runs shared-library tests under strict enforcement, so common code stays clean without paying for two production artefacts everywhere. The audit evidence comes from the settings `go version -m` reports.
go deeper
Understand that this is a policy call made per service rather than a default someone sets once, and that the setting changes what the shipped binary is allowed to do.
Be able to explain why enabling strict enforcement everywhere is risky: shared libraries use MD5 and SHA-1 for non-security work, and those failures land at run time.
Show how you would roll it out safely — pinned snapshot, canary, a required strict-mode check on shared libraries — and where the evidence for the artefact comes from.
Own the boundary conversation with the compliance owner, name the cost of two build variants against the cost of fleet-wide enforcement, and state the condition under which you would reverse the call.
## The decision, stated plainly There are two coherent positions and a lot of incoherent middle ground. Position one: every artefact in the fleet is built with `GOFIPS140` set and runs with `GODEBUG=fips140=only`, so there is exactly one build recipe and no chance of shipping the wrong variant into a regulated slot. Position two: only the services inside the audited boundary get that treatment, and everything else builds normally. Both are defensible; which is right depends on facts you have to go and get rather than reason about from first principles. ## What decides it **Where the audit boundary actually runs.** This is not an engineering decision. A compliance or security officer defines the scope of the assessed system, and if their scope statement covers a service you left `off`, that service moves inside regardless of how inconvenient it is. What remains yours is the implementation: which snapshot is pinned, how strict mode is rolled out, how the evidence is captured. Getting that division wrong — deciding scope yourself and presenting it as a technical fact — is the failure mode that gets a lead overruled late and expensively. **The blast radius of strict mode across shared code.** This is the argument against the fleet-wide position and it is concrete. Internal libraries are full of legitimate non-security digests: an ETag from MD5, a SHA-1 cache key, a checksum pinned by an old wire format. Under `fips140=only` those stop working, some of them by panicking rather than erroring, in services that were never in scope. You have converted a compliance requirement into an availability risk for unrelated systems. **The cost of two variants.** Two artefacts to build, sign, store and roll back, a test matrix that runs both, and the possibility that a bug reproduces in only one of them. That is a real tax, and it is the argument for the fleet-wide position. It is worth paying when one service out of twenty is regulated. It stops being worth it when most of the fleet is in scope: at that point a single recipe is cheaper than the divergence. **Toolchain coupling.** A pinned snapshot ties the regulated build to a specific validated version of the Go Cryptographic Module. When the toolchain moves, somebody has to decide whether the pinned version is still shipped and still covered. Applying that coupling to every service in the organisation means every upgrade drags that question along with it. ## The position I would actually take, and why Scope to the boundary, and build the bridge in CI. - Regulated services build with a **pinned** `GOFIPS140` version. Not `latest`: `latest` enables the mode but compiles the toolchain's current cryptographic source, which carries no external validation and changes under you with each release. Pinning is the difference between "we run FIPS mode" and "we run this validated module version", and only the second is an answer to an auditor. - Those services run `GODEBUG=fips140=only`, introduced through a canary instance rather than a fleet flip, because some rejections are panics. - Everything else builds `off`. - **Shared libraries carry a required check** that runs their tests under `GODEBUG=fips140=only`. This is the load-bearing piece. It costs one CI job per library, and it moves the failure from a regulated service's production traffic to the library author's pull request. Without it, the two-variant policy quietly rots: someone adds a SHA-1 checksum to a common package and nobody finds out until the regulated service redeploys. - **Evidence is mechanical.** The release pipeline records `go version -m` output for the shipped artefact, which carries `GOFIPS140=<version>` and `DefaultGODEBUG=fips140=on`. An assertion in a runbook is not evidence; a setting stamped into the binary is. ## What I would not claim That the mode certifies anything by itself. It selects and enables a module that has been through validation and constrains what the standard library will do. It does not constrain vendored or hand-rolled cryptography compiled into the same binary, and it says nothing about key management, operational procedure, or the rest of what a regime asks for. Overstating this in a design document is how a team ends up surprised in an assessment, and it is the claim a reviewer should push on hardest. ## Revisiting the call Write down the trigger for reconsidering: if the share of services inside the boundary crosses roughly half the fleet, or if the shared-library CI check starts failing often enough that people route around it, the two-variant policy has stopped paying for itself and one recipe becomes the cheaper answer. A decision with no stated reversal condition is the one nobody revisits until it hurts.
- Who can overrule this decision, and on what grounds?The compliance or security owner defines the boundary of the assessed system, so if their scope covers a service you left `off`, that service moves inside and the cost argument does not survive it. What stays yours is the implementation: which snapshot is pinned, whether strict mode gets a canary, how build evidence is captured. Write both down, because the next assessment will ask who decided what.
- What is the cost of maintaining two build variants of the same service?Two artefacts to build, sign, store and roll back, a test matrix that runs both, and the risk that a bug reproduces in only one. It is worth paying when one service is regulated and twenty are not. It stops being worth it when most of the fleet is in scope: at that point one recipe, built with the mode on everywhere, is cheaper than maintaining the divergence.
- How would you keep a shared internal library from breaking the regulated service?Run that library's own tests under `GODEBUG=fips140=only` as a required check on every pull request. It is cheap — one job, one environment variable — and it moves the failure from a regulated service's production request to the library author's change. Pair it with a review rule against new imports of `crypto/md5`, `crypto/sha1`, `crypto/rc4` and `crypto/des` in shared packages.
- What would make you reverse the decision and build the whole fleet in FIPS mode?Two triggers. If the audited boundary grows past roughly half the services, the divergence costs more than the uniformity. And if the shared-library strict-mode check starts failing so often that teams route around it, the two-variant policy has already stopped working and one recipe is more honest. Stating the reversal condition up front is what makes the decision reviewable rather than permanent by default.
saying these in an interview costs you the question
- Turns strict mode on fleet-wide without a canary
- Pins GOFIPS140=latest and calls it validated
- Treats the audit boundary as an engineering decision alone
- Claims the mode by itself certifies the deployment
- Ignores shared libraries when scoping the policy
- Keeps evidence in a runbook rather than in the artefact