Your service publishes a request-signing scheme for bench instruments uploading results — what belongs in the canonical string, and what does an unsigned header mean?
answer
- both sides must build identical bytes
- coverage is the contract
- unsigned header, rewritable in transit
- digest the body exactly as received
- closed list of signed header names
basics
~20 sA canonical string fixes which bytes both sides sign: method, path, query in a defined order, a closed named set of headers, and a digest of the body as received. Anything outside it is unsigned and an intermediary may rewrite it undetected.
solid answer
~50 sThe canonical string is the agreed byte-for-byte rendering of the request that both sides run the keyed tag over, and its contents are the security boundary of the whole scheme. Publish it explicitly: the method, the path, the query parameters sorted and re-encoded into a defined order, a **named, closed** list of headers with normalised values, a digest of the body, and the signed timestamp that bounds replay. Verify over the bytes exactly as they arrived — a verifier that parses the body and re-serialises it before hashing will break perfectly valid signatures over key order, number formatting or whitespace. Anything you left out of that list is unsigned, so an intermediary in the request path can add or rewrite it and no verifier will notice. That is why the list is closed, published and versioned rather than "whatever the caller happened to sign".
code
pseudocode · 26 lines// verify side, in the order the checks must actually run
if content_length > MAX_SIGNED_BODY:
return 413 // refuse before buffering anything
raw = read_raw_body() // bytes as received, no parsing yet
signed_names = request.signed_header_list()
if not signed_names.contains_all(REQUIRED_HEADERS):
return 401 // downgrade attempt: required header unsigned
canonical =
request.method + "\n" +
normalise_path(request.path) + "\n" +
canonical_query(request.query) + "\n" + // sorted, re-encoded
join("\n", [lower(n) + ":" + trim(request.header(n)) for n in sort(signed_names)]) + "\n" +
join(";", sort(signed_names)) + "\n" +
hex(sha256(raw))
key = lookup_secret(request.key_id())
if key == null or not tag_matches(key, canonical, request.signature()):
return 401 // same answer for unknown key and bad tag
if abs(now - request.signed_timestamp()) > SKEW_WINDOW:
return 401
handle(parse(raw)) // parse only after the bytes are authenticatedgo deeper
Remember the one-line version: the signature only protects the parts of the request that were put into the canonical string, and everything else is ordinary untrusted input.
Be able to list the canonical string element by element and say what each one prevents, and explain why the verifier must hash the raw bytes rather than a re-serialisation of the parsed body.
Show the verify-side rules that a published document has to carry: a required minimum signed header set so the caller cannot downgrade coverage, a signed-body size cap, and a diagnostic path that lets a remote integrator find the mismatching line.
Weigh flexibility against a support burden you cannot escape. A caller-named header set and a strict fixed one differ mainly in who pays when the scheme changes, and with hundreds of unattended callers that cost lands on you as a multi-year version overlap.
## The property the whole scheme hangs on A request-signing scheme asks a caller to prove it holds a shared secret **without sending it**. The caller builds a *canonical string* — an agreed rendering of the request — computes a keyed tag over it with the secret (an `HMAC` over a hash function is the usual construction), and sends the tag along with a key identifier. The verifier rebuilds the same canonical string from the request it actually received, recomputes the tag with its copy of the secret, and compares. Everything rests on one property: **both sides must produce identical bytes for the same request, and different bytes for a different request.** So the canonical string is not a formatting detail. It *is* the coverage statement of the scheme. What is inside it cannot be altered in transit without the tag failing. What is outside it can be altered freely, and nothing downstream will know. ## What goes in, and what each piece is buying | Element | Why it is covered | What a captured, valid signature buys if it is left out | |---|---|---| | Method | The same bytes must not be replayable under a different verb | A captured upload replays as a delete against the same path | | Path | A result belongs to one instrument, not another | The identical body is filed under a different instrument's run | | Query, sorted and re-encoded | Callers and intermediaries reorder and re-escape parameters | Either intermittent false rejections, or a flag such as `overwrite=true` flipped undetected | | A named, closed header set | Only the headers you list are protected | Any other header is attacker- or proxy-controlled input | | Digest of the body | The body is the payload; hashing it keeps the canonical string small | Payload substitution under a valid signature | | Signed timestamp | Bounds how long a captured signature stays usable | The signature is valid forever | The query rule catches people out. If you sign the raw query string as it arrived, any intermediary that re-encodes a space or reorders parameters produces an unexplainable rejection on a fleet you cannot debug. Sorting by parameter name and re-encoding both sides to one published rule removes a whole class of support tickets. ## Sign the bytes as received, never a re-serialisation This is where every interoperability bug in this class actually lives. A verifier that deserialises the uploaded document, hands it to the handler, and *then* re-serialises it to compute the digest is comparing a tag over the caller's bytes with a tag over its own bytes. Key ordering, numeric formatting, unicode escaping, a trailing newline: any of them flips the digest while the request is entirely legitimate. So the verify path must capture the **raw body bytes before any parsing**, and hash those. Two consequences follow: 1. You must be willing to buffer what you hash, so publish and enforce a maximum signed body size and reject an oversize request *before* reading it, not after. 2. For a genuinely large upload, put the body digest in a signed header instead: the caller computes the digest, the header carrying it is inside the canonical string, and the verifier checks the signature immediately, then verifies the digest as it streams the body to storage. AWS Signature Version 4 does exactly this with `X-Amz-Content-Sha256`, which is why a signed multi-gigabyte upload does not need to be buffered to authenticate it. ## The signed header set is a closed list, and it is yours to close A scheme can fix the header set in the published version, or let the caller name the set it signed (SigV4 takes the second route, listing them in `SignedHeaders` inside the `Authorization` value). The second is more flexible and carries a trap: if the caller names the set, the verifier must **reject any request whose named set omits a header the scheme requires** — the timestamp, the body digest, the host. Otherwise an attacker takes a captured request, drops the timestamp from the named set and from the request, and the tag still verifies over what remains. That is a downgrade attack against your own flexibility, and it is a verify-side rule, not a caller-side one. ## What you actually publish - The exact concatenation, including separators, trailing newlines and case rules for header names and values. - The canonical ordering and percent-encoding rule for the query. - The required minimum header set, and the statement that headers outside the signed set are unauthenticated input. - The digest algorithm and where the digest travels. - A worked example: one key, one request, the intermediate canonical string, the expected tag. Instrument vendors integrating against you will debug against that example instead of opening a ticket.
- A vendor reports that signatures verify from their test harness but fail once requests go through the customer's corporate proxy. Where do you look first?At the difference between the bytes the caller signed and the bytes you received. Usual causes: the proxy re-encoded or reordered query parameters, normalised the path, rewrote a header that is inside the signed set, or a body-transforming gateway altered the payload. Log the canonical string you rebuilt on a rejection, behind a diagnostic flag, and diff it against the vendor's — the mismatching line names the culprit immediately.
- Should the host or authority be inside the canonical string?Yes. Without it, a signature computed for your staging endpoint is valid at your production endpoint, and a signature for one tenant's hostname replays at another's. Bind it in the required header set, and be explicit in the published document about which value the caller signs when a proxy terminates TLS and rewrites the host — otherwise you have written a rule your own deployment breaks.
- The scheme needs a new required header next year. How do you introduce it?As a new scheme version, carried in the request so both can be accepted at once. Adding a header to the required signed set is a breaking change to the coverage statement: every existing caller's canonical string becomes wrong on the day you enforce it. Run both versions over an announced window, measure which callers are still on the old one, then retire it.
A contract where each clause is initialled in the margin. The initialled clauses are binding and cannot be edited afterwards; a note scribbled on an uninitialled margin can be added by anyone who handled the envelope, and the initials say nothing about it.
saying these in an interview costs you the question
- Signing only the request body and trusting the method and path
- Parsing and re-serialising the body before computing the digest
- Assuming every header the caller sent is covered by the signature
- Letting the caller name the signed header set with no required minimum
- Treating canonical ordering of query parameters as cosmetic tidiness
- Publishing the scheme without a worked example and expected tag