In TLS, why does a server that omits an intermediate from its certificate_list fail for only some clients?
answer
- the server sends, the client builds
- same server, different client abilities
- the gap is one hop below
- roots may be omitted, intermediates may not
- cold store and no route fails
basics
~20 sBecause clients that already hold the missing intermediate, or can fetch it from a location named in the certificate, complete the path anyway. Clients with neither cannot reach a trust anchor and reject the chain, so the same server looks healthy and broken at once.
solid answer
~50 sThe server puts a `certificate_list` on the wire; the client has to build a path from it up to a trust anchor it already holds. If an intermediate is missing from the list, three populations of client behave differently: one already cached that intermediate from an earlier connection and completes the path; one can retrieve it from a location named inside the certificate and completes the path after a fetch; and one can do neither — a freshly started process, a minimal runtime image, a constrained device — and rejects the chain. Nothing about the server changed between those connections, which is why the report is always "it works for most of them". The rule the sender owes is that the first entry is its own certificate, each following entry should certify the one before it, and only a certificate naming a trust anchor may be left out.
code
pseudocode · 17 lines# what the server should send
certificate_list:
entry[0] subject = portal host issuer = Intermediate A
entry[1] subject = Intermediate A issuer = Root R
# Root R omitted: clients are expected to hold it
# what the broken server sends
certificate_list:
entry[0] subject = portal host issuer = Intermediate A
# Intermediate A missing -> client must supply it or stop here
# what the client does with the broken list
for each entry in certificate_list:
if issuer of entry is not present:
if issuer already in local store: continue
if issuer fetchable from location named in entry: fetch and continue
reject chaingo deeper
Remember that a server sends a list of certificates, not one, and that leaving a link out is a server-side mistake even when most connections still work.
Explain the three parts of the ordering rule with their different strengths, and name the three client populations whose differing abilities produce the inconsistent outcome.
Diagnose it: read the chain actually on the wire rather than the file on disk, locate the issuer-subject gap, and reproduce from a cold client with no outbound route before declaring it fixed.
Treat chain completeness as a deployment invariant to be verified continuously rather than a one-time setup step, since issuer migrations and automated renewal both reintroduce the same gap silently.
## The rule the sender is supposed to follow The `certificate_list` inside the `Certificate` message is not a bag of certificates. The specification states three things about it, and they carry different weights: 1. The sender's own certificate **MUST** come first. Everything else in the handshake depends on this — `CertificateVerify` is signed with the key of the first entry. 2. Each following certificate **SHOULD** directly certify the one immediately preceding it, so the list reads leaf-first and upward. 3. A certificate that names a trust anchor **MAY** be omitted, on the understanding that the peers already possess it. Point 3 is the one that gets misread. "A trust anchor may be omitted" licenses leaving out the root — the certificate the client must already hold for trust to mean anything. It does not license leaving out an intermediate, which the client usually does not hold. ## Why the same omission produces three outcomes The server is doing the identical thing on every connection. The variance is entirely on the client side, and there are three populations: - **Already has it.** The client fetched or was shipped that intermediate earlier and kept it. It fills the gap from its own store and the path completes with no sign anything was missing. - **Can go and get it.** The client is willing to retrieve a missing issuer from a location named inside the certificate it already holds. The path completes, but only after an extra network round trip to a third party — which fails in turn if the client has no route out. - **Has neither.** A short-lived process with a cold store, a minimal container image carrying only roots, a constrained device with a fixed list — there is nothing to fall back on. The path stops one hop below the anchor and the chain is rejected. That is the mechanism behind "works in one runtime, fails in another": the two runtimes are not disagreeing about the certificate, they are disagreeing about what they can supply for themselves. ## Why it also looks intermittent on one client | Situation | Outcome with the intermediate omitted | |---|---| | Long-lived client, warm store | succeeds — the gap is filled from cache | | Same client, freshly restarted | fails, or succeeds only after a fetch | | Client on an isolated network segment | fails, because the fetch cannot leave | | Automated caller in a minimal image | fails consistently, and is usually the first reporter | The first row is what the person who deployed the certificate saw. The last row is what the integration partner sees. Both are accurate reports of the same server. ## Order, and how strictly it is enforced TLS 1.2 stated the ordering requirement flatly. TLS 1.3 keeps the leaf-first rule but is explicit that receivers **should not** treat an out-of-order or extra-entry list as a fatal error, because senders legitimately include a second, transitional intermediate during an issuer migration and others are simply misconfigured. That tolerance is a compatibility allowance for receivers — it is not permission for a sender to ship a list in arbitrary order, and plenty of validators are stricter than the tolerance suggests. ## Sending too much is a smaller problem than sending too little - Including the root wastes bytes on every handshake and can never be the thing that makes validation succeed, because a trust anchor is trusted for being in the store, not for arriving on the wire. - Including one extra transitional intermediate during an issuer change is deliberate and useful — it lets clients anchored at either issuer complete a path. - Chains grew large enough that TLS 1.3 gained a compression mechanism: `compress_certificate(27)` negotiates an algorithm and the chain then travels in a `CompressedCertificate` message, with `zlib(1)`, `brotli(2)` and `zstd(3)` defined. ## Confirming it rather than guessing 1. Print the chain the server actually puts on the wire with a command-line TLS client, rather than inspecting the certificate file on disk — the file is not what was sent. 2. Count the entries and check the first is the end-entity certificate. 3. Check whether each entry's issuer matches the subject of the next entry; a gap there is the missing intermediate. 4. Re-test from a client with an empty store and no route to fetch, which is the population that fails. The fix is always the same: configure the server to send the intermediate. It is not a client-side problem even though only some clients report it.
- What exactly does the ordering rule require, and how strict is each part?Three different strengths. The sender's own certificate **must** be the first entry — `CertificateVerify` is signed with its key, so this one is load-bearing. Each following certificate **should** directly certify the one immediately before it. A certificate naming a trust anchor **may** be omitted, since peers are expected to hold anchors independently. TLS 1.3 additionally asks receivers not to treat a misordered list as fatal, which is a tolerance for receivers rather than licence for senders.
- Adding the intermediate made every handshake larger. Can the chain be compressed?Yes, in TLS 1.3. A peer advertises `compress_certificate(27)` listing the algorithms it accepts — `zlib(1)`, `brotli(2)` and `zstd(3)` are defined — and the other side may then send a `CompressedCertificate` message in place of `Certificate`. Both sides must support it, so it is an optimisation rather than something to rely on, and it changes only the encoding: the same `certificate_list` comes out the other end.
- Would sending the root as well make the chain work for the clients that fail?No. A trust anchor is trusted because it is in the client's store, not because it arrived on the wire; a client that does not already hold the root will not start trusting it just because the server transmitted it. Sending the root costs bytes on every handshake and fixes nothing. The missing hop is the intermediate.
A courier hands over a signed permit and the counter-signature of the official who issued it, but leaves out the supervisor who authorised that official. Anyone who has met the supervisor before waves it through; anyone who has not is stuck one signature short.
saying these in an interview costs you the question
- Claims an omitted intermediate fails uniformly on every client
- Says the root must be sent for the chain to validate
- Thinks the server builds the validation path for the client
- Puts the intermediate first and the end-entity certificate last
- Treats it as a client bug because most clients work
- Inspects the certificate file instead of the bytes actually sent