A Python ticket-triage bot hits ssl.SSLCertVerificationError against an internal service behind a private CA. How do you fix it?
answer
- The client was never told about that issuer
- Add an anchor, do not remove the check
- Additive on top, or a closed world
- load_verify_locations takes file, dir or data
- get_ca_certs proves what is actually loaded
basics
~20 sAdd the private CA to the client trust anchors with ssl.SSLContext.load_verify_locations, rather than disabling verification. Then decide deliberately whether that context should also keep the public roots, and make sure the CA file is actually deployed everywhere the bot runs.
solid answer
~40 sThe error means the chain does not reach an anchor this context trusts, so the fix is to give it the anchor. Call `ssl.SSLContext.load_verify_locations(cafile=...)` - or `cadata=` if the certificate arrives as PEM text from configuration - on the context the bot uses. Then choose the trust model on purpose: `ssl.create_default_context()` followed by `load_verify_locations` trusts the public roots **plus** your CA, which suits a bot that also calls the outside world; `ssl.SSLContext(ssl.PROTOCOL_TLS_CLIENT)` with only your CA loaded trusts nothing else, which is stronger for a purely internal client. Confirm what the process really has with `ssl.SSLContext.get_ca_certs()` and `ssl.get_default_verify_paths()`. Most "works on my laptop" versions of this are deployment: the CA file is on the developer machine and not in the image.
code
python · 9 linesimport ssl
ctx = ssl.SSLContext(ssl.PROTOCOL_TLS_CLIENT) # verification already on, no anchors
ctx.minimum_version = ssl.TLSVersion.TLSv1_2
try:
ctx.load_verify_locations(cafile="/etc/pki/internal-ca.pem")
except FileNotFoundError:
print("internal CA bundle is not deployed on this host")
print(ctx.check_hostname, ctx.verify_mode.name, len(ctx.get_ca_certs()))go deeper
Recall the shape of the fix: a private CA is added to the client with load_verify_locations, never worked around by switching verification off. Knowing which method to name is enough at this level.
Explain the mechanics of load_verify_locations - cafile, capath, cadata - and that calling it on a default context is additive while passing cafile to the factory replaces the system store. Know that hostname matching is a separate check.
Demonstrate the diagnosis, not just the call: inspect get_default_verify_paths() and get_ca_certs() in the failing environment, distinguish an unknown issuer from a name mismatch, and treat a missing CA in the container image as a deployment defect.
Own the trust model across services: whether internal clients run closed-world, how the CA reaches every runtime, how it rotates without an outage, and why a single shared context builder beats each team writing their own TLS setup.
## Read the error before changing anything `ssl.SSLCertVerificationError` says one specific thing: the certificate the peer presented could not be chained to a trust anchor this context holds, or it failed a validity check on the way. It does not mean the network is broken and it does not mean TLS is misconfigured on the server. For an internal endpoint issued by an organisation's own CA, the cause is nearly always that the client has never been told about that CA - the platform trust store contains public roots only. ## The correct fix Give the context the anchor: ```python import ssl ctx = ssl.create_default_context() ctx.load_verify_locations(cafile="/etc/pki/internal-ca.pem") ``` `ssl.SSLContext.load_verify_locations` takes three keyword arguments and they cover the ways a CA arrives in practice: `cafile` for a PEM bundle on disk, `capath` for a hashed directory of them, and `cadata` for a PEM string or DER bytes you already hold - which is the one to reach for when the certificate comes from a configuration system or a secret store rather than a file. Called on a context built by the factory it is **additive**: the public roots stay and yours is added. ## Choose the trust model deliberately There are two defensible shapes, and the difference matters: * **Public roots plus the private CA** - `ssl.create_default_context()` then `load_verify_locations`. Right when the same client also talks to external endpoints. * **Private CA only** - `ssl.SSLContext(ssl.PROTOCOL_TLS_CLIENT)`, which starts with an empty store and verification already on, then load just your CA. Now a certificate from any public CA is rejected for these connections. For a bot that only ever calls internal services, this is the stronger position: it removes several hundred third-party organisations from the set of people who can produce a certificate your client will accept. A third shape people reach for by accident is `ssl.create_default_context(cafile=...)`, which loads *only* that file and silently skips the system store - fine if that is what you meant, a surprise outage on the first external call if it is not. ## Hostname matching is a separate gate Adding the CA fixes chain validation and nothing else. If the certificate covers `triage.internal` and the bot dials an IP address or a service alias, verification still fails - now with a hostname mismatch rather than an unknown issuer. The fix is to reissue the certificate with the right subject alternative name, or to dial the name the certificate covers. Turning `check_hostname` off to get past it re-opens the hole you just closed for chain validation, because any certificate your private CA ever issues - to any service, on any host - then satisfies the client for every connection. ## Where these bugs actually live: deployment The reason this is a senior question and not a middle one is that the API call is trivial and the operational failure is not. The bot runs fine on a laptop where the developer imported the CA into a local store months ago, and fails in the container where the trust store is a minimal bundle. Diagnose with two facts, gathered inside the failing process: * `ssl.get_default_verify_paths()` - the file and directory OpenSSL will consult, plus the names of the environment variables (`SSL_CERT_FILE`, `SSL_CERT_DIR`) that override them. A container image, a virtual environment, or a third-party CA bundle package installed as a dependency can each move that answer. * `len(ctx.get_ca_certs())` - how many anchors this context actually holds. Zero means an empty context; a couple of hundred means the platform store; the number changing after your `load_verify_locations` call proves it took effect. Then fix it where it belongs: ship the CA in the image or through configuration management, not by editing code paths per environment. ## The organisational shape of the fix With an eleven-person team running several services against the same internal CA, the failure mode is eleven slightly different context constructions, one of which quietly ends up with verification disabled at 2am. The durable fix is a single small internal helper module that builds the context - factory, private CA, TLS floor - and is imported everywhere, with the CA path or PEM coming from one configuration key. That also gives you one place to rotate the CA, one place to audit, and one place a reviewer has to look. Be careful about where that helper lives: a helper module that reaches back into the service packages it is meant to serve is how a circular import at startup gets introduced, and a TLS helper should depend on nothing but the standard library and its configuration input. ## What never to do Do not set `verify_mode` to `ssl.CERT_NONE`, and do not add an "insecure TLS" configuration flag "just for staging". Certificate errors are the system telling you something true about identity; the answer is always to make the identity check succeed honestly.
- Should that context still trust the public roots?It depends on what the client talks to. A bot that only calls internal services is safer with a closed world: `ssl.SSLContext(ssl.PROTOCOL_TLS_CLIENT)` plus your CA alone, so no public CA can issue a certificate it will accept. A client that also reaches external endpoints needs the factory context with your CA added on top.
- After adding the CA you now get a hostname mismatch instead. What is the fix?Reissue the certificate with a subject alternative name covering the address actually dialled, or dial the name the certificate already covers. Disabling `check_hostname` is not a fix: it would let any certificate your CA has ever issued, for any service, satisfy every connection the client makes.
- It works locally and fails in the container. How do you diagnose that inside the process?Print `ssl.get_default_verify_paths()` and the length of `get_ca_certs()` on the context in both environments. That tells you which bundle OpenSSL is reading, whether `SSL_CERT_FILE` or `SSL_CERT_DIR` is redirecting it, and whether your CA actually made it into the context - which is usually a packaging problem, not a code problem.
You do not stop checking identity documents because a new employer issues its own badges; you add that badge design to the list the guard recognises.
saying these in an interview costs you the question
- Reaches for ssl.CERT_NONE to make the error go away
- Adds an insecure-TLS configuration flag for staging
- Thinks the cafile argument to the factory is additive
- Disables check_hostname to get past a name mismatch
- Never checks which trust store the process is actually reading
- Fixes it on the laptop and never ships the CA to the image