Why would a Kubernetes client-go program log "client-side throttling, not priority and fairness", and how does that differ from an HTTP 429?
answer
- two throttles, two places
- token bucket inside the process
- five per second, burst of ten
- server rejection carries Retry-After
- the log line names which one
basics
~20 sThat delay comes from client-go's own token-bucket rate limiter (5 QPS, burst 10 unless you change them), which holds requests before they leave the process. An HTTP 429 means the API server's Priority and Fairness filter rejected a request that did arrive.
solid answer
~40 sEvery client-go `rest.Config` has `QPS` and `Burst` fields. If you leave them at zero, client-go uses 5 requests per second with a burst of 10, enforced by a token bucket inside the process. When a request waits a long time for a token, client-go logs the wait with the reason `client-side throttling, not priority and fairness`. The request never reached the server, so the fix is on the client: make fewer calls, or raise `QPS`/`Burst` on purpose. An HTTP `429 Too Many Requests` is different. kube-apiserver's API Priority and Fairness filter classified the request, found its priority level full, and rejected it with a `Retry-After` header. client-go retries those responses on its own, waits as `Retry-After` says, and stops after 10 retries by default.
code
go · 17 linespackage main
import (
"k8s.io/client-go/kubernetes"
"k8s.io/client-go/rest"
)
func newClient() (*kubernetes.Clientset, error) {
cfg, err := rest.InClusterConfig()
if err != nil {
return nil, err
}
// Zero values fall back to QPS 5 and Burst 10.
cfg.QPS = 20
cfg.Burst = 40
return kubernetes.NewForConfig(cfg)
}go deeper
Remember the two places a Kubernetes call can be held: client-go's own limiter (5 QPS, burst 10 by default) and the API server's 429 with Retry-After.
Explain the token bucket behind QPS and Burst, why the log names the reason, and that client-go retries 429 responses on its own, honouring Retry-After.
Show that raising client limits only moves the queue into the server. Check APF rejection metrics before and after tuning, and cut call volume with caching before adding tokens.
Frame client QPS as a shared-capacity policy. Set defaults and a review step for operators that raise their limits, so one team's tuning does not become another team's 429s.
## Two throttles on one request path A request from a Kubernetes client can be held up at two separate points before it does any work. The first is **inside the client process**. client-go is the Go client library behind kubectl, kube-controller-manager and most operators, and it rate-limits its own outgoing calls. The second is **inside kube-apiserver**. **API Priority and Fairness (APF)** classifies every request that arrives and caps how many run at once in each priority level. From outside, both look like "my calls are slow". They fail in different places, though, and you fix them in different places. ## The client-side limiter in client-go - Every `rest.Config` has a `QPS` field (float) and a `Burst` field (integer). - If both are left at zero, client-go uses `DefaultQPS` = **5** and `DefaultBurst` = **10**. - The limiter is a **token bucket**. The bucket holds up to `Burst` tokens and refills at `QPS` tokens per second. Each request takes one token, and a request that finds the bucket empty waits. - A clientset built from one config shares one limiter across all its typed clients. Every resource the program touches draws from the same bucket. - When a wait is long, client-go logs "Waited before sending request" with the reason `client-side throttling, not priority and fairness`. That wording is there so nobody blames the server. - Built-in components set their own values. The kubelet config has `kubeAPIQPS` (default 50) and `kubeAPIBurst` (default 100), and kube-controller-manager takes `--kube-api-qps` and `--kube-api-burst`. The request never left the process, so nothing shows up in the API server's metrics. The only evidence is the client's own log and its latency. ## The server-side answer: HTTP 429 APF acts on requests that **did** arrive: - It matches the request to a **FlowSchema**. The FlowSchema names a **PriorityLevelConfiguration**, which allows only a set number of requests to run at once. These slots are called **seats**. - If the level has no free seat, the request waits in a queue. The server answers **`429 Too Many Requests`** in three cases: the queue is full, the wait runs too long, or the level is set to reject instead of queue. - The response carries a **`Retry-After`** header. The API server computes it for each priority level. It starts at 1 second and goes up while that level keeps dropping requests. - The response also carries `X-Kubernetes-PF-FlowSchema-UID` and `X-Kubernetes-PF-PriorityLevel-UID`. These identify the schema and level that handled the request by UID, and on purpose not by name. - client-go treats a 429 with `Retry-After` as retryable. It sleeps and retries, up to **10** times by default, and `Request.MaxRetries` changes that number. ## Telling them apart | Signal | Client-side throttling | APF rejection | |---|---|---| | Where the request waits | In the client's token bucket | In a kube-apiserver queue, then rejected | | HTTP status seen | None, because the request has not been sent | `429` with `Retry-After` | | Evidence | Client log reason `client-side throttling, not priority and fairness` | `apiserver_flowcontrol_rejected_requests_total` rises, and the response has PF headers | | Who fixes it | The client's author | The cluster operator, and often the client's author too | ## What to do about each 1. **Client-side wait.** First ask whether that many calls are needed. A controller that re-reads the same objects on every reconcile should read from a local cache instead. Batching and backoff often remove the need for more tokens. 2. If the calls are needed, raise `QPS` and `Burst` **on purpose and by a moderate amount**. Afterwards, watch the server's APF metrics, because the extra load now lands on a priority level other clients share. 3. **Server-side 429.** Honour `Retry-After`, which client-go already does. Repeated 429s mean the client should send less, not retry harder. 4. If the rejected client is legitimate and important, the fix is in the APF configuration: a better FlowSchema, or more shares for its level. That is the cluster operator's decision. Both limits exist for the same reason: kube-apiserver is shared. A higher client limit does not create server capacity. It only moves the queue from the client's process into the server, where other teams' requests are waiting too. ## A worked example In a document-OCR pipeline, one team's dispatcher creates one Job per scanned document. After a burst of 3,170 uploads, its log fills with waits tagged `client-side throttling, not priority and fairness`, each about 2.4 seconds. The API server's rejected-requests counter stays flat. Together, those two signals say the dispatcher still runs on client-go's defaults of 5 QPS and burst 10, so the backlog is inside the pod: 3,170 creates at 5 per second take more than ten minutes. Setting `QPS` to 20 and `Burst` to 40 removes most of the wait. The team then checks that its own flow has not become the one that causes 429s for everyone else.
- If client-go retries 429s automatically, why can raising QPS and Burst make things worse?A higher client limit moves the waiting from the client's process into kube-apiserver. There, the extra requests share a priority level with other clients. Once that level is full, everyone in it queues longer and starts getting 429s, including the client that raised its limit. Its automatic retries then add even more load. Raise the limit only when the calls are really needed, and watch the server's APF rejection metrics afterwards.
- How can a client tell from the HTTP response which APF objects rejected it?An APF response carries two headers, `X-Kubernetes-PF-FlowSchema-UID` and `X-Kubernetes-PF-PriorityLevel-UID`. They give UIDs, not names, so the names a cluster admin chose are not exposed. Someone who can read the FlowSchema and PriorityLevelConfiguration objects can match those UIDs to names. `kubectl -v=8` prints response headers, so the values are easy to capture.
Client-side throttling is like a ticket machine that lets only a few customers out of the lobby each minute. A 429 is the shop itself turning people away at the counter because every till is busy.
saying these in an interview costs you the question
- The client-side throttling log means the API server rejected the request.
- Setting client-go QPS and Burst very high is always safe because the server protects itself.
- An HTTP 429 from the API server means the client lacks RBAC permission.
- client-go gives up on the first 429, so every application must write its own retry loop.
- client-go's QPS limit is enforced cluster-wide for each ServiceAccount.