skip to content

What does a published request-signing contract owe several hundred instrument sites you cannot upgrade in one night?

level: principalimportance: should knowfreq 28%

answer

  1. a scheme is a public interface
  2. no flag days are available
  3. two secrets valid simultaneously
  4. version the covered-bytes list
  5. publish test vectors and reason codes

basics

~20 s

Versioned coverage rules, two signing secrets accepted at once so rotation is rolling rather than a cutover, a secret scoped per purpose, an explicit at-least-once delivery promise with an idempotency key, diagnosable rejections, and published test vectors.

solid answer

~50 s

A signing scheme is a published interface, and the callers are firmware at sites with no on-call engineer. Everything that would otherwise need a flag day has to be designed as an overlap. Carry a **scheme version** in the request so v1 and v2 verify side by side, because adding a header to the covered set changes every caller's canonical string. Accept **two signing secrets per caller at once**, selected by the key identifier in the request, so a rotation is a rolling deployment rather than a coordinated cutover. Scope a secret to one purpose, so the results uploader's credential cannot drive an unrelated endpoint. State plainly that delivery is at-least-once and name the idempotency key you dedupe on. And make a rejection diagnosable from the far end — a reason code and your own clock — plus test vectors an integrator can check against offline.

code

json · 10 lines
json
{
  "scheme_version": "labsig-v2",
  "supported_versions": ["labsig-v1", "labsig-v2"],
  "required_signed_headers": ["host", "x-upload-date", "x-content-digest"],
  "skew_window_seconds": 300,
  "delivery": "at-least-once",
  "idempotency_key": "run_id",
  "active_key_ids": ["lab-214-2025-11", "lab-214-2026-09"],
  "retirement": { "labsig-v1": "2027-03-31", "lab-214-2025-11": "2026-12-31" }
}

go deeper

for a junior

Take away the framing: once other people's software signs requests to your published rules, those rules are a contract, and changing them is a release for every one of those callers.

for a middle

Be able to separate changes that are compatible from changes that alter the covered bytes, and describe the dual-acceptance mechanism — key identifier in the request, two secrets valid at once — that makes rotation rolling.

for a senior

Show the operational half: per-key and per-version acceptance counts that tell you when a retirement is safe, rejection metrics by reason and site, and reason codes that let a remote integrator self-diagnose.

for a principal

Treat every published rule as a multi-year commitment and cost its migration up front. The judgment call is how much version overlap you are prepared to operate against how fast you need to move, with no ability to force an upgrade at the far end.

## A signing scheme is an interface, not an implementation detail When the callers are your own services, a signing change is a coordinated deploy. When the callers are several hundred unattended instruments, each behind a customer's network, running firmware shipped by a third party, every element of the scheme becomes a long-lived public commitment. The design question stops being "is this secure" and becomes **"what can I change later, and how"**. The useful discipline: for every rule you publish, write down the migration that changes it. If you cannot describe one that does not require every site to update on the same evening, the rule is wrong. ## What the contract has to name - **A scheme version in the request.** The covered-bytes list will change — a new required header, a different digest algorithm, a corrected encoding rule. With a version identifier the verifier picks the right canonicalisation and both generations run side by side; without one, the change is a cutover. - **A key identifier in every request.** The verifier must select the secret from the request rather than from the caller's identity alone. This is what makes every other overlap possible. - **Purpose-scoped secrets.** A secret that signs result uploads should not also authenticate a configuration fetch or a firmware callback. The benefit is blast radius: revoking the uploader's secret after a site's on-premise server is rebuilt stops uploads, not everything else. - **A delivery promise.** Instruments retry when an acknowledgement is lost, so the same correctly signed bytes will legitimately arrive twice. Say so, name the field you deduplicate on, and state what the second attempt returns. Silence here produces duplicate records that are nobody's bug and everybody's problem. - **Rejection semantics.** Which status, which reason codes, and what a caller should do with each one. - **Test vectors.** One key, one request, the intermediate canonical string and the expected tag. A vendor integrating your scheme debugs against that file at 2 a.m. in another timezone, or opens a ticket you answer in three days. ## Rotation without a flag day The only rotation that works at this scale is **two secrets valid at once**, chosen by the key identifier the request carries. The sequence is unremarkable once the scheme allows it: 1. Issue the second secret and start accepting both. Nothing at any site changes. 2. Roll the new secret out to sites at whatever pace the field allows — weeks is normal. 3. Watch per-secret request counts until the old identifier goes quiet, and chase the stragglers by name rather than by hope. 4. Withdraw the old secret, remembering that a request signed just before withdrawal can still arrive inside the skew window. The measurement in step 3 is the part teams skip, and it is the part that makes step 4 a decision rather than a gamble. Note what this is *not*: it is the rotation of the scheme's signing secret, distinct from the lifecycle of a long-lived bearer credential, which is a different node's subject. ## Changing the covered bytes is the breaking change | Change | Breaks existing callers? | How it ships | |---|---|---| | New signing secret for a caller | No | Accept both, roll, retire | | New optional unsigned header | No | Document it as unauthenticated | | New **required** signed header | Yes — every canonical string changes | New scheme version, dual acceptance, announced retirement | | Wider skew window | No | Announce; note the extra replay and revocation exposure | | Narrower skew window | Yes, for drifting clocks | Measure the tail first, quarantine, then narrow | The fourth and fifth rows catch people: a window is part of the contract as surely as the canonical string, and narrowing one is an availability change for the callers who were relying on the slack. ## Rejections have to be debuggable from the other end Every failure mode here is remote. A stale timestamp, an unknown key identifier, an unsupported scheme version and a bad tag are all authentication failures — `401` with `WWW-Authenticate`, never `403`, which would claim you know who is calling and are refusing them. But they demand different actions from the caller: fix your clock, re-provision, upgrade firmware, and "you are being tampered with or misconfigured" respectively. So return a machine-readable reason and, for a skew rejection, your own server time. Be careful about what distinguishing reasons gives away: an unknown key identifier and a bad tag are worth collapsing into one answer, because separating them turns your endpoint into an oracle for guessing valid identifiers, while separating a clock problem from a signing problem leaks nothing a caller does not already know. Finally, instrument the scheme from the provider side: rejection rate by reason, by site and by firmware version. A scheme you cannot see failing is one whose next breaking change you will discover from a customer.

  • How do you know it is safe to withdraw the older signing secret?
    Count accepted requests per key identifier and per site, and only withdraw when the old identifier has been at zero for longer than the longest legitimate gap between uploads — a seasonal instrument may be silent for a month and reappear. Withdrawal is then a decision with evidence behind it rather than a date someone picked.
  • Should an integrator be able to test the scheme without a real secret and a live endpoint?
    Yes. Publish test vectors — a fixed key, a fixed request, the intermediate canonical string and the expected tag — plus a verification sandbox that echoes the canonical string it rebuilt. Almost every integration failure is a canonicalisation mismatch, and both let the integrator find the mismatching line themselves.
  • Why collapse an unknown key identifier and a bad signature into the same rejection?
    Because distinguishing them lets anyone probe which key identifiers exist, one request at a time, without holding a secret. The caller gains nothing from the distinction either: in both cases its credential does not work and it must re-provision. A clock-skew reason is different, since the caller can act on it and it reveals nothing.
  • What does the published contract owe callers about signing errors you introduce yourself?
    A statement of what happens when your own verify path changes — for example if you correct an encoding rule that was wrong in the document. That is a new scheme version, not a silent fix, because callers implemented what you published and their canonical strings are correct against it.

Publishing a plug shape that four hundred appliances are already wired for. You can add a second socket beside the old one and let installations move across over a year, but you cannot change the old one's pin spacing on a Tuesday night.

saying these in an interview costs you the question

  • Rotating the signing secret with a cutover date and an email
  • Adding a header to the required signed set as a patch release
  • One signing secret shared across every endpoint and purpose
  • Returning a bare 401 with no machine-readable reason
  • Assuming a correctly signed request only ever arrives once
  • Publishing the scheme without test vectors or a sandbox