How do you handle 429 RESOURCE_EXHAUSTED from the Gemini API in production?
answer
- 429 means capacity, not correctness
- which window — minute or day?
- backoff needs jitter and a cap
- some codes must never be retried
- limit yourself before the API does
basics
~20 s429 RESOURCE_EXHAUSTED means a quota was hit — requests per minute, tokens per minute, or requests per day. Retry per-minute breaches with exponential backoff and jitter and a capped attempt count; a daily breach will not clear by retrying, so shed or queue the work.
solid answer
~50 sA 429 from `generateContent` is `RESOURCE_EXHAUSTED`: some quota dimension for your project and model is spent. The first job is deciding *which* dimension, because they recover on very different timescales — requests-per-minute and tokens-per-minute windows clear in seconds, while a requests-per-day cap does not clear until the daily reset, so a backoff loop against it just burns attempts and money. Retry the per-minute cases with exponential backoff plus jitter, a hard attempt cap, and an overall deadline; treat a daily exhaustion as load-shedding or queueing, and alert. Pair that with client-side control so you stop generating 429s in the first place: a concurrency limiter or token bucket sized under your real limits, and `client.models.count_tokens` to check big prompts against the token budget before sending. In the SDK these surface as `google.genai.errors.APIError` (with `ClientError` for 4xx and `ServerError` for 5xx), and you branch on `.code`.
code
python · 16 linesimport random, time
from google import genai
from google.genai import errors
RETRYABLE = {429, 500, 503, 504}
def generate_with_retry(client, *, model, contents, attempts=5):
for attempt in range(attempts):
try:
return client.models.generate_content(model=model, contents=contents)
except errors.APIError as exc:
if exc.code not in RETRYABLE or attempt == attempts - 1:
raise
delay = min(32.0, 2 ** attempt) * random.random() # full jitter
time.sleep(delay)
raise RuntimeError("unreachable")go deeper
Know that 429 means a rate or quota limit was hit rather than a bad request, and that the standard response is to wait and retry rather than to change the prompt.
Explain exponential backoff with jitter and a capped attempt count, and be able to say which status codes are retryable — 429, 500, 503, 504 — versus 400, 403 and 404, which never are.
Demonstrate that quota has dimensions with different recovery times, that a daily exhaustion must be shed or queued rather than retried, and that client-side concurrency limits plus token counting prevent most 429s from happening.
Own the capacity design: quota separation between interactive and bulk traffic, admission control and degradation policy, cost of retried generations, and which quota metrics become the service-level indicators you plan headroom against.
## What the error actually says Google APIs pair an HTTP status with a canonical status name. HTTP 429 carries `RESOURCE_EXHAUSTED`, and for `generateContent` it means one of your quota dimensions is spent. It is not a statement about the prompt, the model's health, or your key's validity — it is pure capacity accounting. ## The dimensions and why they matter Quota on the Gemini API is enforced per project, and typically **per model**, across several dimensions at once: - **RPM** — requests per minute. - **TPM** — tokens per minute (input tokens dominate here, and a single huge document can consume the whole window on its own). - **RPD** — requests per day. They have different recovery behaviour, and conflating them is the mistake this question is designed to expose: - Blow through RPM and the window rolls forward continuously; a few seconds of backoff genuinely fixes it. - Blow through TPM and the same is true, but the fix is often to *send less* rather than to wait — one oversized prompt can re-trigger it immediately on retry. - Blow through RPD and no amount of backoff helps until the daily reset. A retry loop here is a machine that converts your compute into 429s. Because quota is per model, one model being exhausted does not mean another is, which is what makes a fallback to a smaller or different model a viable degradation strategy where quality allows. ## Retry mechanics When you do retry, retry properly: - **Exponential backoff with jitter.** Full jitter — sleeping a random interval up to the current backoff ceiling — prevents a fleet from re-synchronising into a thundering herd. Fixed-interval retry across many workers reproduces the burst that caused the 429. - **A cap on attempts and an overall deadline.** Three to five attempts, bounded by the caller's latency budget. A request nobody is waiting for any more should not still be retrying. - **Honour server timing hints.** Google error payloads may carry retry timing in the error details; when present, prefer it over your own schedule. - **Idempotency is free here.** `generateContent` is stateless, so a retry is safe from a correctness standpoint — but it is a *new billed generation*, not a resumption, so retries cost real money. ## Which errors are retryable at all Branch on the status, not on the presence of an exception: - **429 RESOURCE_EXHAUSTED** — retryable, subject to the dimension caveat above. - **500 INTERNAL / 503 UNAVAILABLE** — transient server side; retry with backoff. - **504 DEADLINE_EXCEEDED** — retryable, but a repeat usually means the request is too large or the generation too long; streaming or a smaller task is the real fix. - **400 INVALID_ARGUMENT** — never retry. The request is malformed and will fail identically forever. - **403 PERMISSION_DENIED** — never retry. The key, project, or API enablement is wrong. - **404 NOT_FOUND** — never retry. Usually a mistyped or retired model id. In `google-genai` these arrive as `google.genai.errors.APIError`, specialised into `ClientError` for 4xx and `ServerError` for 5xx, and the numeric status is on `.code`. Catching bare `Exception` and retrying everything is how a permanent 400 turns into a five-times-slower permanent 400. ## Prevention beats reaction A production system should rarely see 429 at all, because it governs itself: - **Limit concurrency** at the client with a semaphore or a token bucket sized under your actual RPM, so the queue forms in your process where you control it, instead of at the API where it costs you errors. - **Budget tokens before sending.** `client.models.count_tokens(model=..., contents=...)` gives the real count for a prompt; use it to reject or split oversized inputs before they eat the TPM window. - **Separate the traffic classes.** Interactive user requests and bulk backfills competing for one quota pool means the backfill starves the users. Give them different projects/keys, different priorities, or move bulk work to the offline batch path rather than the online endpoint. - **Instrument.** Track 429 rate by dimension and by model, not just a global error count — the response to "RPM spikes at the top of the hour" is scheduling, while the response to "RPD exhausted at 3pm daily" is a quota increase or a workload cut. ## The shape of a good answer Name the error, name the dimensions, explain that backoff only helps some of them, distinguish retryable from permanent statuses, and finish on prevention — rate limiting and admission control — rather than on the retry loop. Interviewers ask this because a candidate who only knows "retry with backoff" will build a system that hammers a daily quota all afternoon.
- Why is full jitter preferred over plain exponential backoff across a fleet?Because plain doubling makes every worker that failed at the same moment retry at the same moment, reproducing the burst that caused the 429 and stretching the outage. Sleeping a random interval up to the current ceiling spreads the retries across the window, so recovering capacity is absorbed smoothly instead of being re-saturated on each wave.
- Which Gemini API errors should never be retried, and why?400 INVALID_ARGUMENT, 403 PERMISSION_DENIED and 404 NOT_FOUND. All three describe a request or configuration that is wrong in a way time cannot fix — malformed body, missing permission or API enablement, mistyped or retired model id. Retrying them adds latency and log noise while guaranteeing the same failure, and it hides the real defect from whoever is on call.
- Your interactive traffic starts failing because a nightly backfill exhausted the quota. How do you fix it structurally?Separate the pools rather than tuning retries. Give bulk work its own project or key so it cannot consume the interactive allowance, run it through the offline batch path instead of the online endpoint, and put admission control in front of it with a concurrency cap. Prioritisation inside a single shared quota is the fallback, not the goal.
- How do you tell a tokens-per-minute breach from a requests-per-minute one?By correlating with your own instrumentation: a TPM breach shows a normal or low request rate alongside unusually large prompts, while an RPM breach shows a request-count spike with ordinary prompt sizes. Recording prompt token counts per call — from usage_metadata or count_tokens — is what makes the distinction visible; without it, both look like an undifferentiated 429 rate.
saying these in an interview costs you the question
- Retrying a daily-quota 429 in a tight loop until it clears
- Catching every exception and retrying it identically
- Fixed-interval retries across a whole worker fleet
- Assuming 429 means the prompt or model is wrong
- Treating quota as a single number rather than several dimensions