skip to content

A batch job writing thousands of objects per second under one Amazon S3 key prefix starts getting HTTP 503 SlowDown responses. What is being throttled, and what do you change?

level: seniorimportance: should knowfreq 40%

answer

  1. the ceiling is not per bucket
  2. key layout is capacity planning
  3. 503 SlowDown is retryable, not fatal
  4. backoff needs jitter across a fleet
  5. repartitioning is automatic but not instant

basics

~20 s

S3 scales request rate per key prefix, not per bucket: roughly 3,500 write and 5,500 read requests per second each. Spread keys across many prefixes to multiply that ceiling, and retry 503 SlowDown with exponential backoff plus jitter while S3 repartitions.

solid answer

~50 s

S3 publishes a per-prefix request rate — at least 3,500 PUT/COPY/POST/DELETE and 5,500 GET/HEAD requests per second — and the important word is *per prefix*, not per bucket. A bucket has no documented aggregate ceiling, so concentrating a high-rate job under one prefix caps you at one prefix's worth of throughput while the rest of the key space sits idle. Two changes follow. First, make the client well behaved: 503 `SlowDown` is retryable, so back off exponentially with jitter rather than hammering, which the AWS SDKs do if you let them, and consider the adaptive retry mode. Second, redesign the key layout so writes fan out across many prefixes — a hashed or high-cardinality leading path segment instead of one hot date directory. S3 repartitions automatically as load rises, but that takes time, so a burst into a cold prefix throttles before it adapts.

go deeper

for a junior

Know that S3's request-rate limits apply per key prefix rather than per bucket, and that 503 SlowDown means retry later, not that the request was invalid.

for a middle

Quote the floors — about 3,500 writes and 5,500 reads per second per prefix — and explain how key layout determines how many prefixes your traffic actually touches.

for a senior

Diagnose it end to end: backoff with jitter and adaptive retries on the client, key fan-out as the structural fix, and the cold-prefix ramp that explains why the burst throttled but the rerun did not.

for a principal

Own the tradeoff you are buying — hashed prefixes cost you listability and date-filtered lifecycle rules, and batching small writes into fewer objects may beat both throughput tuning and request cost.

## What the limit actually is Amazon S3 documents a request rate of **at least 3,500 PUT/COPY/POST/DELETE and 5,500 GET/HEAD requests per second per partitioned prefix** in a bucket. Two things about that sentence carry all the weight: - **"per prefix", not per bucket.** There is no published aggregate request-rate limit for a bucket. Capacity comes from how many prefixes your traffic touches. - **"at least".** These are floors S3 commits to, not caps. S3 raises effective capacity by splitting the key space across more partitions as sustained load grows. ## What counts as a prefix A prefix is any leading portion of the key up to a partition boundary S3 chooses — not "the part before the first slash", and not a folder, because S3 has no folders. `logs/2026/08/21/host-7/app.log` contains several candidate prefixes. What matters is that keys sharing a long common leading string tend to land in the same partition, and keys that diverge early do not. So the classic hot-spot is a design where every writer produces keys with an identical long leading path — one date directory, one tenant, one sequential counter — and all of them therefore contend for one partition's budget. ## Why 503 SlowDown appears, and why it disappears later S3 partitions the key space automatically and repartitions as sustained request rates increase. That adaptation is not instantaneous: it takes time for S3 to observe the load and split. The consequence is a specific and very recognizable failure shape — a job that ramps from zero to a very high rate against a *new* prefix gets 503 `SlowDown` in its first minutes, and the same job run again later against the now-warm prefix runs clean. Candidates who have not seen this conclude their code is broken; candidates who have, ramp their concurrency up gradually. 503 `SlowDown` is explicitly a **retryable** error. The correct client response is exponential backoff **with jitter** — jitter matters because a fleet of workers that all back off by the same deterministic amount re-converges and thunders again in lockstep. The AWS SDKs implement this; the mistake is usually a hand-written client, or a wrapper that treats any 5xx as fatal, or a retry loop with no ceiling that makes the overload worse. botocore and several SDKs also offer an **adaptive** retry mode, which adds client-side rate limiting on top of backoff so the client throttles itself when it sees sustained 503s. ```bash # Let the SDK/CLI back off and self-throttle instead of hammering. aws configure set default.retry_mode adaptive aws configure set default.max_attempts 10 ``` ## Redesigning the key layout The durable fix is to make the traffic touch more prefixes. Concretely: - Put a **high-cardinality component early** in the key — a short hash of the record id, a shard number, a tenant id — rather than after a long shared path. - `2026/08/21/<hash>/<id>` concentrates the day's traffic; `<hash>/2026/08/21/<id>` spreads it. - Understand the cost: fanning out by hash makes prefix listing and lifecycle rules that filter by date harder, because your date is no longer the leading component. This is a genuine tradeoff between write throughput and queryability, and saying so out loud is what distinguishes a senior answer. A historical note worth getting right: before 2018, S3 required random hash prefixes for *any* high request rate, and a lot of folklore survives from that era. Today the per-prefix rates are much higher and S3 partitions automatically, so you do **not** need to obfuscate keys as a matter of course. You spread across prefixes when your measured rate against one prefix approaches the documented floor — not reflexively. ## The rest of the checklist - **Reduce the request count, not just spread it.** Thousands of small objects per second is often a batching problem: aggregating many small records into fewer larger objects removes the throttling entirely and cuts request cost. - **A single hot object cannot be fixed with prefixes.** If the pressure is thousands of GETs per second on *one* key, more prefixes do nothing — that is a caching problem, and a different design conversation. - **Measure before redesigning.** CloudWatch request metrics for S3 (opt-in, per bucket or per filter) show request rates and 5xx counts, so you can see whether the throttling is broad or concentrated. ## The shape of a good answer Name the per-prefix nature of the limit, explain that a bucket has no published aggregate ceiling, describe backoff with jitter as the immediate client fix, describe key fan-out as the structural fix, mention that S3 repartitions but not instantly, and acknowledge the tradeoff that hashed prefixes cost you listability. That covers the mechanism, the mitigation and the judgment.

  • Your job ramps to full rate immediately and throttles, but the same job runs clean an hour later. What is happening?
    S3 repartitions the key space as sustained load grows, and that adaptation takes time. A cold prefix hit with an instant burst throttles until S3 splits it; the warmed prefix then absorbs the same rate. The practical mitigation is to ramp concurrency up gradually rather than starting at maximum.
  • Would adding prefixes help if the load is thousands of GETs per second against one single object?
    No. Prefix fan-out multiplies capacity across distinct keys, but one object lives in one place, so its requests all land the same way. That is a caching or content-distribution problem rather than a key-layout problem, and it is solved in front of S3 rather than inside it.
  • What does the adaptive retry mode in the AWS SDKs add over standard retries?
    Standard mode retries retryable errors such as 503 SlowDown with exponential backoff and jitter. Adaptive mode adds a client-side rate limiter that measures throttling responses and slows the client's own request rate, so a fleet degrades gracefully instead of retrying its way deeper into the overload.

saying these in an interview costs you the question

  • Thinks the request-rate limit is per bucket
  • Says you must always randomize S3 key prefixes
  • Treats 503 as fatal instead of retrying with backoff
  • Retries immediately with no jitter across many workers
  • Opens a support ticket to raise a bucket request quota

context