skip to content

In production, a system using pre-signed S3 URLs for client uploads starts seeing a spike in 403 SignatureDoesNotMatch and expired-token errors from clients on flaky mobile networks, even though the URLs were generated correctly seconds before use. What are two distinct root causes worth investigating, and how do they differ?

level: seniorimportance: should knowfreq 45%

answer

  1. skew vs mutation, two different bugs
  2. signature covers headers, not just the URL
  3. compare signed timestamp vs receive time
  4. multipart per-part URLs for resumable retries

basics

~20 s

Either the client's clock is off so the link looks expired too early, or the upload got retried in a way that changed something about the request (like its size), so it no longer matches what was originally signed. Both look like the same error but need different fixes.

solid answer

~40 s

First cause: clock skew. The client's or an intermediate proxy's clock is off from the signing service's clock, so a URL that's actually still within its validity window gets rejected as expired by the time-bound signature check. Second cause: request mutation on retry. Mobile networks retry uploads, and if the retry changes any signed element, such as a header, byte range, or chunking, the signature no longer matches the actual request being sent. The fix for skew is NTP-synced clients and reasonable clock-skew tolerance; the fix for mutation is re-signing on retry rather than replaying a stale signed request, or switching to per-part presigned URLs for multipart, resumable uploads.

go deeper

for a junior

Should recognize that both symptoms look like 'the link didn't work' but may have different underlying causes worth asking about.

for a middle

Should be able to name clock skew as one cause and describe checking timestamps in logs to diagnose it.

for a senior

Should independently identify both clock skew and retry-induced signature mismatch, and propose distinct fixes for each, including multipart per-part URLs.

for a principal

Should also flag the CDN cache-key failure mode, weigh skew-tolerance policy trade-offs, and design signing infrastructure that supports safe re-signing on retry across the fleet.

## How the signature is validated Pre-signed URLs are validated by recomputing a signature over a canonical form of the request, namely: - the verb - resource path - specific headers - query parameters - a timestamp and comparing it against the signature embedded in the URL, subject to an expiration window. That mechanism means the signature is fragile to two very different classes of problems that both surface to the client as an authentication failure, such as 403 SignatureDoesNotMatch or an expired-token rejection, even though their causes and fixes are unrelated. ## First cause — clock skew The first cause is **clock skew**. Signature validity is time-bound: the storage service checks the request's timestamp, such as AWS's `X-Amz-Date` header, against its own clock, allowing only a bounded window, roughly 15 minutes of skew tolerance on AWS by default, before rejecting the request as expired or not yet valid. If a mobile device's system clock drifts, which is common on devices with unreliable NTP sync, airplane-mode toggling, or manually set clocks, a URL that is well within its intended validity period by the issuing server's clock can appear expired or premature by the client's clock, causing rejections that have nothing to do with the URL's actual freshness. This shows up disproportionately on mobile because desktop operating systems sync time more reliably, and mobile networks and devices are more likely to have skewed or delayed clock sync. ## Second cause — request mutation on retry The second, distinct cause is **request mutation on retry**. The HMAC signature covers more than just the URL string; it typically covers specific headers like `Content-Length`, `Content-MD5`, or `Content-Type` if they were included in the signing process, and, for chunked or multipart transfers, characteristics of the body itself. Mobile networks are flaky, so client HTTP libraries and OS-level upload managers often retry failed requests automatically. If that retry logic doesn't resend byte-for-byte the same request that was originally signed, for example: - resuming from a different byte offset, - changing Content-Length because of a partially sent body, - or switching chunked-encoding strategy, the signature computed by the storage service on the actual retried request no longer matches the signature embedded in the URL, and the request is rejected. Critically, simply reissuing the identical original request would work; it's specifically a mismatch introduced by smart retry or resume logic that breaks it. ## Telling them apart in production Distinguishing the two in production requires different diagnostics. - **For clock skew**, compare the request's signed timestamp against server-side receive time in access logs; a consistent multi-minute offset points to skew, and the fix is ensuring client devices sync time via NTP and, if you control the signing window, allowing a slightly more generous tolerance rather than the tightest possible expiry. - **For request mutation**, compare the exact bytes and headers of the retried request against what was originally signed; this shows up as intermittent failures correlated with connection drops mid-transfer rather than any particular time-of-day pattern, and the fix is architectural: don't retry an old signed request with different parameters, either re-fetch a freshly signed URL from your app server before each retry attempt, or move to a signed-URL scheme designed for resumability, such as issuing a presigned URL per part in an S3 multipart upload, so an interrupted part can be retried in isolation without invalidating the whole transfer's signature. ## A third, related failure mode A third, related failure mode worth knowing about, even though it wasn't one of the two asked for, is a CDN or proxy sitting in front of a presigned GET URL caching a response keyed only on the path while ignoring the query string where the signature and expiry live; this can cause a client to receive a cached response for a signature that's already expired, or, more dangerously, one client's cached response served for a completely different request if the cache doesn't fully differentiate by query parameters, silently breaking the isolation the pattern was meant to provide. ## The fix that makes it worse The instinctive but wrong fix engineers reach for, namely just making the expiry longer so retries have more room, treats the symptom rather than the cause and actively makes things worse: it doesn't fix skew or mutation mismatches, since those fail regardless of how long the window is, while it does widen the exposure window if the URL leaks.

  • Why can't you just increase the expiry window to work around retry-related signature mismatches?
    Because expiry only controls how long a correctly-formed request stays acceptable; it does nothing about a retried request whose signed elements, like Content-Length or byte range, no longer match what was originally signed. Widening expiry also increases the exposure window if the URL leaks, so it trades one problem for a worse one without fixing the original bug.
  • How would a caching CDN in front of a presigned GET URL cause a subtle bug?
    If the CDN caches based on path while ignoring the query string, where the signature and expiry live, it may serve a cached response for a signature that has already expired, or serve one client's cached response to a different request entirely, breaking both freshness and isolation guarantees.
  • What's the operational fix teams use to avoid re-signing races when a mobile client needs to retry a large upload over an unreliable connection?
    Switch to S3 multipart upload with a presigned URL issued per part, or a presigned initiate plus presigned upload-part URLs, so a failed part can be retried independently without invalidating the whole transfer's signature, or adopt a resumable protocol fronted by short-lived per-chunk tokens.

Like a concert ticket with a printed time window and seat number: if the venue's clock is wrong, a still-valid ticket gets rejected at the door as too early or expired, and if you tear the ticket and tape it back together slightly differently before showing it again, the barcode no longer scans as the original, even though it's basically the same ticket.

saying these in an interview costs you the question

  • Assumes the error is just 'URL expired' without considering skew vs mutation vs CDN caching
  • Suggests making URLs valid for days to fix retries
  • Doesn't know the signature covers headers and body characteristics, not just the path
  • No mention of clock sync or multipart uploads as remedies

context