A gateway decodes a JWT, re-serializes its JSON and forwards it; verification now fails. Why?
answer
- The signature covers bytes, not meaning
- No canonicalization step exists
- Key order and whitespace are not preserved
- Treat the token as opaque
- Changing a claim means re-signing
basics
~20 sA JWS signature covers the exact encoded bytes of the header and payload segments joined by a dot, not the abstract JSON. Re-serializing changes whitespace, key order or escaping, so the signing input differs and verification fails.
solid answer
~40 sThe signing input of a compact JWS is the ASCII string `BASE64URL(header) + "." + BASE64URL(payload)` — the encoded segments exactly as transmitted. The signature is a function of those bytes, not of the JSON's meaning. Any round trip through a JSON parser and serializer can reorder object members, change spacing, alter Unicode escaping or renormalize numbers; re-encoding then yields different Base64url text, so the recomputed signing input no longer matches what the issuer signed and verification fails even though every claim is semantically identical. The rule that follows is simple: treat a received token as an **opaque string**. Pass it through byte-for-byte, and if a claim genuinely must change, mint and sign a new token — which requires the signing key, and that requirement is precisely the security property you want.
code
text · 3 linessigning_input = BASE64URL(UTF8(JOSE header JSON)) + "." + BASE64URL(payload bytes)
signature = BASE64URL( sign(key, ASCII(signing_input)) )
token = signing_input + "." + signaturego deeper
Remember that the token is a string you pass along untouched; anything that rebuilds it from parsed claims breaks the signature.
State the signing input precisely — encoded header, dot, encoded payload — and name what a JSON round trip changes: order, whitespace, escaping, number formatting.
Diagnose it from symptoms: verifies at one hop and fails at the next, claims identical in logs. Compare raw strings and hunt the component that deserializes tokens.
Set the platform rule that tokens are opaque credentials, and design claim enrichment as a re-mint by a trusted issuer or a side-channel header rather than in-flight editing.
## What the signature is computed over In a compact JWS the signature covers a single string: the Base64url-encoded header, an ASCII dot, and the Base64url-encoded payload. That string — the *JWS signing input* — is what the algorithm named in `alg` consumes. Verification repeats the operation over the received first two segments and compares the result to the third. There is no canonicalization step anywhere: JOSE deliberately signs the serialized bytes so that no two implementations can disagree about what "the same JSON" means. ## Why re-serializing breaks it JSON has many textual forms for one logical document. A parse-then-serialize round trip can change: - **member order** — most serializers do not preserve the issuer's ordering; - **whitespace** — pretty-printing or compacting; - **string escaping** — a non-ASCII character emitted literally in UTF-8 versus as a `\u` escape; - **number formatting** — trailing zeros, exponent notation, or an integer widened by a language that has only floating-point numbers. Any one of those changes at least one byte, which changes the Base64url text, which changes the signing input, which changes the signature. The failure is total: a MAC or signature check is all-or-nothing, so a single byte drift produces exactly the same rejection as a forged token. This makes the symptom confusing in production — the claims look right in the logs, so engineers suspect key rotation or clock issues when the real cause is a component that helpfully "normalized" the token. ## Diagnosing it The tell is that the token verifies at the edge and fails downstream, or verifies when captured before a particular hop and fails after it. Capture the raw token string at both points and compare them as strings, not as decoded claims — a decoded-claims comparison will show them as identical, which is the trap. Then look for the component in between that deserializes tokens: an API gateway with a claims-enrichment feature, a service mesh filter, a logging or tracing layer that round-trips the header, a test fixture that rebuilds tokens from parsed claims, or a serialization framework applied to a DTO that holds the token as structured data rather than as a string. ## The design rule A token is a **credential in string form**. Carry it as a `String`, store it as a `String`, log it (redacted) as a `String`, and never model it as a parsed object that gets re-emitted. If a component needs to add information — a tenant, an enriched role, a correlation id — the options are to carry that data alongside the token in its own header, or to have a trusted component mint a *new*, separately signed token for the internal hop. "Edit the claims in flight" is not an option, and the fact that it is not is the point: if a middlebox could rewrite claims and keep verification passing, so could an attacker. ## The same rule, other symptoms The identical mechanism explains a family of issues: - **Padding re-added.** Some libraries re-encode segments with `=` padding; the segment bytes change and verification fails. - **Case or alphabet conversion.** Anything that upper-cases, trims or URL-decodes the token corrupts it. - **Whitespace from headers.** A stray newline or a leading space picked up from a header value or an environment variable changes the last segment's bytes. All of them have the same fix: preserve the bytes. ## The reassuring corollary Because the signature binds the exact encoded segments, nobody — not a proxy, not the client, not an attacker — can change a single claim and keep the token valid without the key. Verifying the signature is what turns the payload from attacker-controlled input into a statement you can act on, and this byte-exactness is the mechanism that makes that guarantee airtight.
- How would you confirm this diagnosis quickly in production?Capture the raw token string before and after the suspect hop and diff them as strings. Comparing decoded claims is the trap — they will match. A single differing byte in segment one or two, or a re-added '=' pad, confirms that something in between round-tripped the token.
- A service must add a tenant claim before calling downstream. What is the correct design?Either carry the tenant in its own request header alongside the untouched token, or have a trusted internal issuer mint a new, separately signed token for the internal hop. Never edit claims in flight: if a middlebox could do that and keep the token valid, the signature would guarantee nothing.
- Why does JOSE sign the encoded text instead of canonicalized JSON?Because canonicalization is a notorious source of disagreement — implementations differ on ordering, escaping and number formats, and any mismatch becomes a signature bypass or a false rejection. Signing the exact serialized bytes removes the ambiguity entirely at the cost of requiring byte-preserving handling.
saying these in an interview costs you the question
- Says the signature covers only the payload
- Assumes semantically equal JSON produces the same signature
- Blames clock skew when claims look correct
- Suggests re-signing in a proxy with a different key
- Models the token as a parsed object in a DTO