After a base-image change, a Go log shipper's tls.Dial fails with 'x509: certificate signed by unknown authority'. What is the real fix, and how do you keep the skip flag out of the repo?
answer
- which of the two x509 errors is it
- a static binary does not carry roots
- minimal images ship no CA bundle
- normal tests pass either way
- assert the wrong name fails
basics
~20 sThat error means no trust root reaches the collector's certificate, usually because the new image ships no CA bundle. Supply the roots, then guard it with a test that requires a mismatched name to fail.
solid answer
~50 sRead the error first: unknown authority is a trust-root problem, not a name problem, so `ServerName` is not the fix and neither is disabling verification. The usual cause after an image change is a minimal base with no CA certificate bundle, which leaves the system root set empty, or a collector that moved behind an internal CA the image never carried. The fix is to supply roots — install the CA bundle in the image, mount the company root and point `RootCAs` at it, or on Unix set `SSL_CERT_FILE` to the bundle's path. Then make the workaround impossible to keep: a check in CI that fails the build on `InsecureSkipVerify: true` outside an allow-listed test, and a negative test that dials with a deliberately wrong `ServerName` and requires the handshake to fail. That test is the real guard — if anyone disables verification again, the dial succeeds and the test goes red.
code
text · 2 linesx509: certificate signed by unknown authority
x509: certificate is valid for collector.internal, not 10.0.4.17go deeper
Recognise that an unknown-authority message is about missing trust roots, and that a minimal container image often ships no CA bundle at all.
Distinguish the two common x509 failure messages and name the fix for each, and explain where a Go client's roots come from when none are configured.
Walk the full incident: diagnose from the error, choose the narrowest fix, then add the build check and the negative test that stop the workaround coming back.
Decide how trust material reaches every service in the fleet and who owns rotating it, so that no on-call engineer is ever left with a broken deploy and a one-line temptation.
## Read the error before changing anything Two handshake failures dominate, they look alike on a dashboard, and they have opposite fixes. - `x509: certificate signed by unknown authority` — the client could not build a chain from the presented certificate to anything in its trust set. The issuer is the problem. - `x509: certificate is valid for collector.internal, not 10.0.4.17` — the chain verified fine, but the name you demanded is not in the certificate. The expected name is the problem. A log shipper that worked yesterday and fails today with the first message, right after a base-image change, is almost never a certificate problem at all. It is a client that lost its roots. ## Why an image change breaks trust With `tls.Config.RootCAs` left nil, Go verifies against the platform trust store. On Unix-like systems `crypto/x509` reads the CA bundle from the conventional filesystem locations, and honours the `SSL_CERT_FILE` and `SSL_CERT_DIR` environment variables as overrides. A `scratch` or otherwise minimal image contains none of those files, so the effective root set is empty and *every* outbound handshake fails with unknown authority — a fact that surprises people who assume a statically linked Go binary carries its own root list. It does not. The second common cause is that the collector's certificate is now issued by an internal CA. Public roots cannot chain to it by construction, so the image needs the company root regardless of how complete its public bundle is. ## The fixes, narrowest first 1. **Set `RootCAs` explicitly** to a pool holding the company root, loaded from a path the platform mounts. This is the narrowest option: it names exactly what this client trusts, it is visible in code review, and it is testable. For a shipper that only ever dials one collector, refusing every public CA is a feature. 2. **Install the CA bundle in the image** (add the certificates package to the base, or copy the bundle in during the build). Right when the binary talks to many destinations and should follow the host's trust policy. 3. **Point `SSL_CERT_FILE` at a mounted bundle.** Useful when you cannot rebuild, but it is process-wide, Unix-only, and silently does nothing if the path is wrong — a poor default, an acceptable emergency lever. All three keep verification on. The one-line alternative does not, and it is the one that gets copy-pasted: a debugging workaround committed at 3am appears in the next service that dials the same endpoint, then the one after that, because it visibly "worked" the last time. ## Making the workaround stick out Two mechanisms, and you want both. **A check that fails the build.** `crypto/tls`'s skip flag is a plain struct field, so a small analyzer over the AST — the kind you write with `golang.org/x/tools/go/analysis` — can flag any composite literal that sets it true, with an allow-list for specific test files. A CI grep is a cruder version of the same guard and still worth having on day one. The value is that the change stops being invisible: a reviewer sees a build failure and an explicit exception rather than one extra field in a struct literal. **A negative test.** This is the part people miss. Tests normally assert that things work, and a client with verification disabled passes every such test — that is exactly why the workaround survives review. So write the test that only passes while verification is on: dial with a deliberately wrong `ServerName` and require the handshake to fail, ideally asserting the failure really is a hostname mismatch rather than a connection refused. If someone later sets the skip flag, that dial now succeeds and the test turns red, which is the behaviour you want from a guard. A companion test that dials with an empty `RootCAs` pool and requires an unknown-authority failure covers the trust-root half in the same way. ## Handling the 3am version When the pager has gone off and the deploy is broken, the pressure is to take the one-line fix. The honest answer is that the correct fixes are also small: mounting a bundle and restarting, or shipping a config change that points `RootCAs` at a file, are the same order of effort as flipping a boolean. If you truly must ship the flag to stop the bleeding, treat it like any other emergency measure: scope it to the single call site, open the ticket before you merge, and put the negative test in the same change so the follow-up is red until it is done. The failure mode this leaf exists to prevent is not one deliberate exception — it is an undated workaround that quietly becomes the house style.
- Why is a normal integration test useless as a guard here?Because it asserts success, and a client with verification disabled succeeds too — more reliably, in fact, since nothing can make it fail. Every happy-path test stays green while the security property is gone. Only a test that requires a specific failure, such as a deliberately mismatched `ServerName` producing a hostname error, distinguishes a verifying client from one that accepts anything.
- Is SSL_CERT_FILE a good long-term fix?Rarely. It is honoured by `crypto/x509` on Unix-like systems only, it applies to the whole process rather than one destination, and a typo in the path fails open into an empty root set with no warning at start-up. It is a fine emergency lever when you cannot rebuild an image, but for one endpoint behind a private CA, setting `RootCAs` on that connection's `tls.Config` is narrower, reviewable and testable.
- How do you find every other service that already copied the workaround?Search the organisation's code, not just this repo — the flag is a literal struct field, so a code search across repositories finds it reliably, and the same analyzer you wire into this build can be run over the others. Then order the results by what the client dials: services talking to internal endpoints over a private CA are the ones the fix actually unblocks, and they are usually the ones that copied it first.
saying these in an interview costs you the question
- Treats unknown authority as a hostname problem
- Assumes a statically linked Go binary carries its own root list
- Adds the skip flag and calls the incident resolved
- Relies on happy-path tests to prove verification still works
- Sets SSL_CERT_FILE process-wide instead of scoping the trust set
- Leaves an emergency workaround with no ticket and no expiry