skip to content

OpenRouter starts returning 402 on some calls and 429 on others — what differs?

level: seniorimportance: should knowfreq 45%

answer

  1. One is money, one is pacing
  2. Retrying will never fix one of them
  3. Free slugs carry their own caps
  4. Daily cap tiers on lifetime purchases
  5. 503 is not a rate limit

basics

~20 s

402 means the OpenRouter credit balance cannot cover the call — no amount of retrying helps, only a top-up. 429 means a rate limit was hit, on OpenRouter's free-tier caps or an upstream provider, and it clears with time.

solid answer

~50 s

They are different failures with different fixes. **402** is insufficient credits: the account balance is exhausted or negative, so the request was rejected on commercial grounds. A retry loop will hammer the API forever and never succeed; the only remedies are topping up, enabling auto top-up, or moving the workload to a bring-your-own-key route. Alarm on it as a hard, page-worthy condition. **429** is a rate limit, and on OpenRouter it comes from two different places: the platform's own caps on free `:free` model variants — a per-minute request cap plus a daily cap that is far lower until you have purchased a threshold amount of credits — or an upstream provider's limit surfaced through the gateway on a paid route. It clears with time, so it deserves backoff and, on a paid route, possibly a different provider or model. Do not conflate either with **503**, which means no provider matching your routing constraints was available, or **502**, which is an upstream provider error.

code

python · 18 lines
python
import random, time, requests

class CreditsExhausted(Exception):
    pass

def call_with_policy(session, url, payload, headers, budget=4):
    for attempt in range(budget):
        r = session.post(url, json=payload, headers=headers, timeout=60)
        if r.status_code == 402:
            raise CreditsExhausted("top up OpenRouter credits; retrying cannot help")
        if r.status_code == 503:
            raise RuntimeError("no provider matched routing constraints")
        if r.status_code in (429, 502):
            time.sleep((2 ** attempt) + random.random())
            continue
        r.raise_for_status()
        return r.json()
    raise RuntimeError("retry budget exhausted")

go deeper

for a junior

Recall that 402 means you are out of OpenRouter credits and 429 means you are being rate limited, and that only one of the two is worth retrying.

for a middle

Explain the two sources of 429 — the platform's free-variant caps and an upstream provider's limit — and describe a retry policy that branches on status code rather than message text.

for a senior

Show the operational design: runway-based balance alarms, auto top-up, bounded backoff with jitter, and a clear separation of 402, 429, 502 and 503 in both handling and dashboards.

for a principal

Own the exposure created by one shared prepaid balance: per-key spending limits, degraded modes when credits run out, and the decision about which workloads move onto provider contracts so a single drained pool cannot take every product down.

## Two failures that look alike in a log line Both show up as red in a dashboard and both stop traffic, but the operational response is opposite. One is a money problem you cannot retry your way out of; the other is a pacing problem you can. Getting them backwards produces either a retry storm against an empty wallet or a needless page for something that would have cleared in sixty seconds. ## 402 — insufficient credits OpenRouter is prepaid. When the balance cannot cover a call, the request fails with HTTP 402. The important properties: - **It is not transient.** Retrying with exponential backoff is exactly wrong: the condition changes only when money arrives. - **It is total.** Every model, every provider, every route through the account fails simultaneously, because they all draw on one balance. This is the flip side of the single-balance convenience. - **It is predictable.** The credits endpoint reports purchased credits and lifetime usage, so remaining balance is knowable at any time. A burn-rate alarm — hours of runway remaining, not just dollars left — gives you days of warning. Mitigations, roughly in order of maturity: auto top-up so the balance refills without a human; a monitored low-balance alarm with a real on-call route; per-key limits so one runaway workload cannot drain the shared pool; and, for large steady workloads, bring-your-own-key routes so the bulk of inference bills to a provider contract instead of credits. ## 429 — rate limited HTTP 429 says slow down. On OpenRouter there are two distinct sources. **Platform caps on free model variants.** Model slugs suffixed `:free` are served at no token cost and are correspondingly limited: a per-minute request cap, plus a daily request cap. The daily cap is deliberately tiered — small for accounts that have never purchased credits, substantially larger once lifetime purchases pass a threshold (documented as ten credits at the time of writing). This trips people who prototype on `:free`, ship it, and discover the ceiling in production. Free variants are for evaluation; a product should not depend on them. **Upstream provider limits.** On a paid route, the limit belongs to the provider actually serving the model. Because your traffic shares that provider's capacity with other gateway users, a 429 can arrive even when your own volume is modest. When your routing configuration permits alternatives, the gateway can send the request to another provider serving the same model rather than failing it — which is part of why routing preferences and cost behaviour are entangled in practice. Handling: back off with jitter, cap the retry budget, and make the caller's timeout shorter than the retry chain so a queued request cannot outlive the user waiting for it. ## Distinguishing them in code Branch on status, not on the message string. A pragmatic policy: - **402** → stop retrying, fail the request, fire a critical alert, and if a degraded mode exists, switch to it. - **429** → retry with exponential backoff and jitter, up to a small budget; if the route is free-tier, do not retry into a daily cap that resets on a clock, surface the limit instead. - **503** → no provider satisfied the routing constraints; loosening a preference or allowing another provider is the fix, not waiting. - **502** → the upstream provider errored; one retry is reasonable, and repeated 502s from the same provider argue for routing away from it. - **408** → timeout; treat like any slow-call failure. ## The reporting angle Because 402 is account-wide, it is the single most damaging failure mode of the one-balance design and deserves the same seriousness as an expiring TLS certificate: a dated, monitored, owned thing. Because 429 is route-specific, it belongs in per-model dashboards where a rising rate is an early signal that a provider is saturating and your routing needs to change before customers notice.

  • Your service is on a free :free model slug and fails every afternoon but works each morning. What is happening?
    You are hitting the daily request cap on free variants, which resets on a clock rather than sliding. Morning traffic consumes the day's allowance and the rest of the day returns 429. Backoff cannot fix a cap that resets at a fixed time. The real fix is to move the workload onto a paid slug; free variants are for evaluation, not production dependence.
  • How do you get advance warning of a 402 rather than discovering it in production?
    Poll the credits endpoint for purchased credits and lifetime usage, compute remaining balance, and alarm on runway — hours left at the current burn rate — not on the raw dollar figure, since burn rate changes with traffic. Pair that with auto top-up so the common case never needs a human, and keep the alarm as the backstop for a payment failure.
  • Why can a 429 arrive on a paid route even when your own request rate is low?
    Because the limit belongs to the upstream provider serving that model, and its capacity is shared across everyone reaching it through the gateway. Your own volume is only part of the picture. Where your routing configuration allows alternative providers for the same model, the request can be served elsewhere instead of failing, which is why a saturating provider shows up as a routing question as much as a client-pacing one.

saying these in an interview costs you the question

  • Retrying a 402 with exponential backoff
  • Treating 402 and 429 as one generic rate-limit failure
  • Assuming free-tier caps are per model rather than account-wide
  • Reading 503 as a rate limit and waiting it out
  • Depending on :free slugs for production traffic

context