Your GitHub Models prototype starts returning HTTP 429 under load — how do you respond?
answer
- quota gate, not an auth gate
- several ceilings at once
- account-scoped, shared with CI
- backoff with jitter, bound concurrency
- retries buy headroom, not capacity
basics
~20 sTreat 429 as quota, not a bug: GitHub Models throttles per account on requests per minute and per day, tokens per request and concurrency, with tighter limits on larger models. Back off using the wait the error states, and move sustained load to a paid endpoint.
solid answer
~50 sFirst establish which limit you hit. The free allowance is expressed as several independent ceilings — requests per minute, requests per day, input and output tokens per request, and concurrent requests — and they are tighter for high-tier models than for small ones. They are scoped to the **account**, so other experiments and CI jobs share your budget. Operationally: honour the wait period the 429 response describes rather than retrying immediately, add exponential backoff with jitter, cap concurrency client-side so you throttle yourself before the service does, and consider a smaller catalog model where quality allows. But the honest senior answer is that retry tuning only buys headroom — a workload that reliably saturates a prototyping quota belongs on paid usage or a dedicated provider endpoint with capacity you control. Instrument it: log the model, the limit type and the wait so you can see whether you are near a ceiling before users do.
go deeper
Know that 429 means rate limited, not unauthorised, and that the correct reflex is to wait the period the response gives you rather than retry straight away.
Explain the separate ceilings — requests per minute and per day, tokens per request, concurrency — that they vary by model tier, and how backoff with jitter differs from naive retry.
Show operational judgment: bound concurrency client-side, instrument 429s by model and limit type, trim prompt bloat, and recognise when the workload has structurally outgrown a prototyping quota.
Own the capacity policy: who may generate load against a shared account quota, the threshold at which a workload must move to paid or provider capacity, and the degradation behaviour the product accepts.
## Read the signal correctly HTTP 429 means you were authenticated and refused for quota. Rotating tokens, re-checking permissions or switching SDKs are all wasted moves. The first question is *which* ceiling you hit, because GitHub Models enforces several at once: - **Requests per minute** — burst shape. - **Requests per day** — sustained volume. - **Tokens per request** (input and output caps) — a single oversized prompt fails regardless of pacing. - **Concurrent requests** — parallel fan-out, which is what usually surprises people who added a thread pool. Limits differ by model class: small models get generous allowances, larger flagship and reasoning models much smaller ones. And they are **account-scoped**, not per-application, so your load test, a colleague's notebook and a CI evaluation job all draw from the same budget. The account's plan tier also matters — a paid developer plan raises the allowances that a free account gets. ## The immediate fix 1. **Do not hot-retry.** An immediate retry on 429 makes the situation worse and can extend the penalty window. Read the wait period the response reports and sleep at least that long. 2. **Exponential backoff with jitter.** Uncoordinated clients that all back off by the same amount reconverge into a thundering herd; randomised delay spreads them. 3. **Bound concurrency client-side.** A semaphore sized below the service's concurrency ceiling converts a spray of 429s into orderly queuing you control. 4. **Shed or queue rather than fail.** For batch work, a bounded queue with a worker pool is strictly better than parallel bursts; for interactive work, decide deliberately whether to degrade (cached answer, smaller model) or surface an error. 5. **Right-size the request.** Token-per-request limits are hit by prompt bloat — long retrieved context, whole files pasted in, an unbounded conversation history. Trimming history and retrieved chunks often removes the failure entirely. ## The structural fix Retry logic is a shock absorber, not capacity. If 429s are routine rather than exceptional, the workload has outgrown the free tier and you have three real options: - **Enable paid usage.** GitHub Models supports billed per-token usage beyond the free allowance once it is turned on for the account, which raises the ceiling without changing a line of code. - **Move to a provider or cloud endpoint.** A vendor's own API or a cloud-hosted deployment of the same model family gives you purchasable quota, quota increases you can negotiate, and regional placement. Because the request shape matches, this is a base URL, credential and model-identifier change. - **Reduce demand.** Cache repeated prompts, batch work into off-peak windows, route easy requests to a smaller model and reserve the flagship for hard ones, and stop calling the model for things a deterministic function can answer. ## What separates a senior answer - **Isolation thinking.** Because quota is account-scoped, a load test can starve a demo happening at the same time. Decide who is allowed to generate load, and keep continuous CI evaluation on a small model or a schedule that does not collide with human use. - **Observability.** Log every 429 with model id, limit type, wait duration and caller. Alert on a rising rate *before* it becomes an outage. Track token usage per request so you see prompt bloat trending up. - **Graceful degradation as a product decision.** Falling back to a smaller model changes answer quality; that is a decision for the product, not something to bury in a retry handler. - **Explicit exit criteria.** Write down the request rate at which this workload must leave the free tier, and validate the migration path early rather than during an incident. ## Anti-patterns to name Minting extra tokens or extra accounts to multiply quota is both ineffective and a terms problem. Infinite retry loops without a ceiling turn throttling into a hang. Catching the error and returning an empty answer hides a capacity problem from everyone who could fix it.
- Why is bounding concurrency client-side better than just retrying on 429?Retrying reacts after the service has already spent effort rejecting you, and uncoordinated retries synchronise into bursts. A semaphore sized below the concurrency ceiling shapes traffic before it leaves your process, so latency stays predictable, work queues instead of failing, and you keep headroom for other consumers on the same account.
- How does the account scope of these limits change how you run load tests?It makes load testing a shared-resource event. Your test competes with colleagues' notebooks, CI evaluation jobs and any demo running at the same time, and can exhaust the daily allowance for everyone. Schedule it, announce it, cap it, or point the test at a paid endpoint — that is also the endpoint whose capacity you actually need to validate.
- When is falling back to a smaller catalog model the right response to throttling?When the task tolerates it and you have measured that it does. Routing easy requests to a small model and reserving the flagship for hard ones is legitimate capacity design, but the quality change is a product decision, not a detail to hide in a retry handler. Evaluate both models on real prompts first, and log which model served each request.
- What is the fastest way to reduce token-per-request failures?Shrink the prompt. Those failures are usually unbounded conversation history, oversized retrieved context, or whole files pasted into the message. Trim history to a window, cap retrieved chunks, and summarise older turns. Measure input tokens per request and treat a rising trend as a defect, because it raises both failure rate and cost.
saying these in an interview costs you the question
- Retries immediately on 429 with no backoff
- Regenerates the token, thinking 429 is an auth failure
- Creates extra accounts or tokens to multiply free quota
- Assumes rate limits are per application rather than per account
- Treats retry logic as a substitute for buying capacity