skip to content

How does the DeepSeek API behave under load, given it publishes no rate limits?

level: seniorimportance: should knowfreq 44%

answer

  1. no published per-key limit
  2. pressure shows up as waiting
  3. filler bytes, not error codes
  4. 503 and 402 mean different things

basics

~20 s

DeepSeek does not enforce a per-key request or token limit. Under heavy traffic it queues your request and holds the HTTP connection open, sending filler — blank lines when not streaming, SSE comment lines when streaming — so client timeouts, not 429s, are the usual failure mode.

solid answer

~50 s

DeepSeek's documented position is that it does not constrain a user's rate; it tries to serve every request. The consequence is that pressure shows up as latency rather than rejection. When the platform is busy, your HTTP request stays connected while it waits, and the server emits keep-alive filler so intermediaries do not drop it: empty lines for a non-streaming call, and SSE comment lines beginning with a colon for a streaming call. Both are inert — a conforming SSE client ignores comments, and the blank lines precede the real JSON body. So the client-side risks are an aggressive read timeout that kills a request which would have succeeded, and a parser that chokes on the filler. Errors still exist and mean different things: 429 rate limit reached, 503 server overloaded — both worth retrying with backoff — while 402 insufficient balance is a prepaid-credit problem that no retry will fix.

go deeper

for a junior

Know that DeepSeek may simply take a long time instead of rejecting you, and that a very short client timeout is the most likely cause of failures you see. Do not assume slow means broken.

for a middle

Explain the keep-alive filler concretely: blank lines for non-streaming, SSE comment lines starting with a colon for streaming, both ignored by a conforming client. Then separate 429 and 503 from 402.

for a senior

Demonstrate the production judgment — bound concurrency on your side because the provider will not, tune read timeouts to measured p99, make retries idempotent, and route 402 to a human alert rather than into the retry loop.

for a principal

Own the risk that a provider with no published limits gives you no contractual capacity. Decide what evidence justifies depending on it for a critical path, what the degradation story is when latency triples, and whether a second provider is warranted.

## The unusual policy Most providers publish a rate-limit table — requests per minute, tokens per minute, tiers you graduate through — and reject you with 429 when you exceed it. DeepSeek's documented stance is different: it does not constrain a user's rate and will try to serve every request. There is no per-key RPM or TPM number to look up, which means capacity planning against DeepSeek cannot be done by reading a limits page. It has to be done by measuring latency under your own load. ## What back-pressure looks like on the wire Because the platform absorbs rather than rejects, congestion surfaces as waiting. When DeepSeek's servers are under heavy traffic your request may take a long time to produce a response, and during that wait the HTTP connection is deliberately kept alive rather than closed. The server sends filler bytes whose only job is to stop proxies, load balancers and client libraries from treating the silence as a dead connection: - **Non-streaming requests** receive empty lines while they wait. When the answer is ready, the real JSON body follows. A tolerant JSON client skips the leading whitespace without noticing. - **Streaming requests** receive SSE comment lines — lines beginning with a colon. The Server-Sent Events specification defines a line starting with `:` as a comment that the client must ignore, so a conforming SSE parser drops them silently and only surfaces the real `data:` events. The operational lesson is that a hand-rolled parser is where this bites. Code that assumes the first bytes off the socket are JSON, or that treats any non-`data:` line in an SSE stream as a protocol violation, will report failures that are actually the server behaving as designed. ## Timeouts become the real limit If the provider will not reject you, your own timeout becomes the rejection mechanism. A five- or ten-second read timeout — perfectly reasonable against a fast endpoint — will abort DeepSeek calls that were merely queued, and if the client then retries, it adds load to a system that is already congested, producing a retry storm against a provider that was never going to shed the traffic for you. Sound practice is a generous read timeout tuned to observed p99 rather than to a guess, a bounded overall deadline so a request cannot hang forever, and idempotency at the application layer so an abandoned call that the server eventually completed does not cause duplicate side effects. Streaming helps here for a second reason beyond perceived latency: once the first real token arrives you have evidence the request is progressing, which lets you distinguish "queued" from "stuck". ## The error codes still matter "No rate limits" does not mean no errors. DeepSeek's documented status codes include several you must handle distinctly: - **429, rate limit reached** — you are sending requests too quickly. Still possible despite the general policy; back off and pace. - **503, server overloaded** — the platform is under high traffic. Wait and retry. - **500, server error** — a transient fault; retry after a short pause. - **402, insufficient balance** — your account has run out of prepaid credit. This one is categorically different: it is not transient, backoff will never clear it, and retrying just burns your own capacity. It needs an alert to a human and a top-up. - **401** for a bad key, **400** for a malformed body and **422** for invalid parameters are client-side and must not be retried at all. A retry policy that lumps every non-200 into "retry with backoff" is the classic failure here: it turns a billing outage into a silent, long-running loop that nobody notices until the feature has been down for an hour. ## Designing around it Since the provider will not push back, you must bound concurrency yourself. Cap in-flight requests at the application or fleet level so that a traffic spike converts into your own queue — where you can observe it, shed low-priority work, and degrade gracefully — rather than into thousands of connections all waiting on DeepSeek at once, each holding a socket and a thread. Monitor time-to-first-token separately from total latency; under congestion the first diverges while the second stays roughly proportional, which is a clean early signal. Finally, remember that the prepaid model interacts with all of this. Because 402 arrives with no warning at the API layer, balance should be monitored as an availability concern, not as a finance one: the account balance endpoint gives you remaining credit, and alerting on it well above zero prevents a full-feature outage caused by an empty wallet.

  • Why is 402 the one status code you must never put behind exponential backoff?
    Because it is not transient. 402 means the prepaid balance is exhausted, and no amount of waiting restores it — only a top-up does. A retry loop on 402 turns a billing event into an invisible outage that keeps consuming your own workers while every call fails. Treat it as a page-a-human condition and monitor remaining credit as an availability metric, not as a finance report.
  • If DeepSeek does not rate-limit you, why bound concurrency on your side at all?
    Because unbounded fan-out hurts you, not the provider. Every queued call holds a socket, a connection-pool slot and often a thread; a spike converts your service into thousands of waiting requests with no capacity left for anything else. A concurrency cap moves the queue inside your system, where you can measure it, prioritise, shed low-value work and fail fast instead of stalling every caller.
  • How do you tell a queued request apart from a stuck one?
    Watch time-to-first-token rather than total latency, which means preferring streaming for anything user-facing. Under congestion the first real SSE `data:` event is delayed while keep-alive comment lines keep arriving, so a stream that is still receiving filler is alive and waiting. A connection with no bytes at all for an extended period is a genuinely stuck one and is the right case for an overall deadline to abort.

saying these in an interview costs you the question

  • Assumes a slow DeepSeek call means the request was dropped
  • Treats keep-alive blank lines as a malformed response
  • Retries a 402 insufficient balance with exponential backoff
  • Sets an aggressive read timeout and concludes the API is unreliable
  • Claims DeepSeek publishes a per-key tokens-per-minute limit

context