What is the difference between GitHub's primary and secondary rate limits, and how should a client react?
answer
- two throttles, not one
- a quota versus a behaviour rule
- headers tell you one of them
- concurrency and content creation punished separately
- exponential backoff with jitter
basics
~20 sPrimary limits are the published hourly request quota for your credential, reported in x-ratelimit-* response headers. Secondary limits are anti-abuse throttles on burst, concurrency and content creation, returned as 403 or 429. Back off on both; never retry straight away.
solid answer
~60 sThe **primary** limit is a per-hour budget attached to the credential: unauthenticated requests get a small per-IP allowance, an authenticated personal access token a much larger hourly one, and a GitHub App installation gets its own budget that scales with the installation. Every REST response carries `x-ratelimit-limit`, `x-ratelimit-remaining`, `x-ratelimit-used`, `x-ratelimit-reset` (a UTC epoch second) and `x-ratelimit-resource`, so a client can steer by watching `remaining` and pausing until `reset` rather than sprinting into a wall. Separate buckets exist for search and for GraphQL, and `GET /rate_limit` reports them all without consuming quota. **Secondary** limits are different: they are abuse-prevention rules on *behaviour* — too many concurrent requests, too much burst against one endpoint, too many content-creating requests in a short window — and they are not exposed as a countdown you can watch. They surface as a 403 or 429 whose body says you exceeded a secondary rate limit, sometimes with `retry-after`. The correct client honours `retry-after` when present, otherwise backs off exponentially with jitter, keeps concurrency low, and never retries a secondary-limit rejection immediately.
code
console · 18 lines$ curl -sD - -o /dev/null -H "Authorization: Bearer $TOKEN" \
https://api.github.com/repos/acme/widgets/issues?per_page=100
HTTP/2 200
x-ratelimit-limit: 5000
x-ratelimit-remaining: 4987
x-ratelimit-used: 13
x-ratelimit-reset: 1755612000
x-ratelimit-resource: core
etag: W/"3f1c..."
$ curl -sD - -o body.json -X POST -H "Authorization: Bearer $TOKEN" \
https://api.github.com/repos/acme/widgets/issues/7/comments -d '{"body":"ping"}'
HTTP/2 403
x-ratelimit-remaining: 4102
retry-after: 60
$ cat body.json
{"message":"You have exceeded a secondary rate limit. Please wait a few minutes before you try again."}go deeper
Know that GitHub's API limits how many requests you may make per hour and that responses carry headers telling you how much is left. Waiting rather than retrying instantly is the key habit.
Explain the x-ratelimit-* headers and the reset timestamp, name the separate buckets for core, search and GraphQL, and describe conditional requests with ETags returning 304 without spending quota.
Demonstrate the diagnosis: distinguish primary from secondary on a 403, honour retry-after, back off with jitter, cap concurrency, pace writes, and instrument remaining quota so exhaustion is predicted rather than discovered.
Own the design so limits are rarely reached — credential isolation between automation and human tooling, webhook-first architecture, caching layers, and a shared client library that enforces the backoff policy for every team.
## Two different mechanisms GitHub throttles in two ways and interviewers want to hear that you know they are not the same thing. **Primary rate limits** are a quota: X requests per hour for this credential, on this resource class. They are predictable, published, and fully observable from response headers. Hitting one is a planning failure. **Secondary rate limits** are abuse protections: they trigger on *shape* of traffic rather than total volume — how many requests you have in flight at once, how hard you hammer a single endpoint in a short window, how quickly you create content (issues, comments, pull requests). They are deliberately not exposed as a precise counter you can optimise against, because that would let an abuser sit exactly at the line. Hitting one is a behaviour failure. ## Reading the primary limit Every REST response includes: - `x-ratelimit-limit` — the size of the bucket. - `x-ratelimit-remaining` — how much is left. - `x-ratelimit-used` — how much has been consumed. - `x-ratelimit-reset` — a Unix timestamp in UTC when the window resets. - `x-ratelimit-resource` — which bucket this request drew from, for example `core`, `search`, or `graphql`. When you exhaust it you get a 403 or 429 with `x-ratelimit-remaining: 0`, and the only correct response is to wait until `x-ratelimit-reset`. `GET /rate_limit` returns every bucket's state and does not itself count against the limit, which makes it the right thing to call at startup and for dashboards. Rough shape of the buckets, all subject to change: unauthenticated requests are limited per source IP and are tiny; an authenticated user token gets a much larger hourly allowance; the Search API has its own far smaller per-minute allowance because each query is expensive; GraphQL is metered in points rather than requests. GitHub App installations get their own budget, which scales with the size of the installation — one reason to move heavy automation off personal tokens. **Conditional requests are the underused lever.** Store the `ETag` from a response and send it back as `If-None-Match`. A `304 Not Modified` reply does not count against the primary rate limit, which makes polling loops dramatically cheaper without changing their logic. ## GraphQL is metered differently The GraphQL API charges a **point cost** per query, calculated from how many nodes the query could return, against an hourly point budget. You can ask the API what a query costs by including the `rateLimit` field in the query itself, which returns `limit`, `cost`, `remaining`, and `resetAt`. There is also a ceiling on how many nodes a single call may request. The practical consequence: a badly shaped GraphQL query with large `first:` values can be far more expensive than the same data fetched over several REST calls, so "GraphQL to save rate limit" is only true if you actually measure the cost. ## Reacting correctly 1. **Distinguish the two.** On a 403/429, check `x-ratelimit-remaining`. Zero means primary — sleep until `x-ratelimit-reset`. Non-zero, with a body mentioning a secondary rate limit, means behavioural throttling. 2. **Honour `retry-after` when present.** It is authoritative; do not shorten it. 3. **Otherwise back off exponentially with jitter.** Without jitter, a fleet of workers all resume in lockstep and re-trigger the limit immediately. 4. **Cap concurrency.** Secondary limits punish parallelism hard; a small fixed worker pool with a shared token bucket usually outperforms a large one that keeps getting rejected. 5. **Serialise writes.** Content-creating requests (comments, issues, pull requests) are throttled more aggressively than reads. A migration that posts thousands of comments must pace itself deliberately, not fan out. 6. **Never retry-storm.** Immediately re-issuing a secondary-limit rejection extends the penalty. ## Designing so you rarely hit them - Prefer webhooks over polling; events cost nothing against your quota. - Use conditional requests for the polling you keep. - Ask for `per_page=100` so a full listing costs the fewest requests. - Cache aggressively at your own layer; most repository metadata changes rarely. - Give heavy automation its own identity — a GitHub App installation — so it does not share a budget with interactive tooling and so its usage is separately observable. - Instrument `x-ratelimit-remaining` as a gauge and alert on it trending toward zero, rather than discovering exhaustion as a wave of 403s at 3am. ## The failure mode to describe The classic incident is a batch job that fans out one request per repository across a large organisation, with high concurrency and no backoff. It exhausts the primary limit within minutes, every worker starts retrying, the retries trip the secondary limit, and unrelated tooling sharing the same token goes down with it. The fix is the whole list above, but the structural one is credential isolation: automation should never share a rate-limit bucket with the tools humans depend on.
- How do you tell a primary limit from a secondary limit when both can return 403?Look at x-ratelimit-remaining on the rejected response. Zero means the hourly quota is spent and you must wait until x-ratelimit-reset. A non-zero remaining alongside a body saying you exceeded a secondary rate limit means behavioural throttling, where retry-after, if present, governs and otherwise you back off exponentially and reduce concurrency.
- A nightly job across 2,000 repositories exhausts the limit every run. What do you change first?Cut requests, not just pace them: use per_page=100, store ETags and send If-None-Match so unchanged resources return 304 without spending quota, and replace whatever the job is detecting with webhook events where one exists. Then move the job onto its own GitHub App installation so it stops sharing a budget with interactive tooling.
- Does moving to GraphQL automatically reduce rate-limit pressure?No. GraphQL is metered in points based on how many nodes a query could return, not in requests, so a query with large page sizes across nested connections can cost more than the equivalent REST calls. Include the rateLimit field in the query to see the actual cost, and shape the query to that number rather than assuming a saving.
- Why does jitter matter in the backoff?Because without it every worker that was rejected at the same moment retries at the same moment, recreating the burst that tripped the limit and extending the outage. Randomising each worker's delay spreads the resumption, which is the difference between recovering in one backoff cycle and oscillating through several.
saying these in an interview costs you the question
- Assumes any 403 from GitHub means bad credentials
- Retries immediately after a rate-limit rejection
- Thinks more parallel workers finish the job faster
- Believes GraphQL has no rate limit
- Ignores ETags and re-fetches unchanged resources