A Kubernetes search-autocomplete service shows p99 doubling and 24% of CFS periods throttled during an 18-hour certificate-expiry window of TLS re-handshakes; how do you choose between resizing thread pools, raising the CPU limit, or removing it?
answer
- cause before capacity
- thread count versus budget
- Downward API rounds up
- requests protect neighbours, not limits
- Guaranteed QoS needs both limits
basics
~20 sFirst match the service's parallelism to its CPU limit, since most throttling below the limit is over-parallelism. Raise the limit when real demand exceeds it. Remove it only with an accurate request and an accepted loss of Guaranteed QoS.
solid answer
~50 sI start with the cheapest fix that removes the cause. If the runtime sized its pools from the node's core count, I feed it the limit instead, for example through a Downward API `resourceFieldRef` on `limits.cpu`, which rounds up to whole cores. If throttling persists because the handshake burst genuinely needs more CPU, I raise the limit, and usually the request with it so the scheduler accounts for it. Removing the CPU limit removes the CFS quota entirely; the request still sets the container's share under contention, so neighbours stay protected as long as requests are honest. The costs: the pod drops from Guaranteed to Burstable, loses static CPU manager pinning, a namespace LimitRange may re-inject a default limit, and a regulated cluster's policy may forbid it. I keep the memory limit either way, because memory cannot be throttled.
code
yaml · 22 linesapiVersion: v1
kind: Pod
metadata:
name: autocomplete-api
spec:
containers:
- name: api
image: registry.example.internal/search/autocomplete:3.14.2
env:
- name: WORKER_THREADS
valueFrom:
resourceFieldRef:
containerName: api
resource: limits.cpu
divisor: "1"
resources:
requests:
cpu: 600m
memory: 1536Mi
limits:
cpu: 1500m
memory: 2Gigo deeper
Remember the three levers: fewer threads, a higher limit, or no limit. Know that the memory limit stays in every case.
Explain how the Downward API exposes limits.cpu, why it rounds up, and why a request still sets CPU share when there is no limit.
Show the order of the fixes and re-measure after each one. List the concrete costs of removing a limit: QoS class, CPU pinning, LimitRange defaults and policy.
Frame the cluster-wide policy on CPU limits: honest requests enforced, limits optional for latency-sensitive tiers, and how cost attribution changes when bursts are free.
## The incident A search-autocomplete service runs on a 16-node regulated-workload cluster with `limits.cpu: 1500m` and `requests.cpu: 600m`. During an 18-hour certificate-expiry window, clients reconnect repeatedly and every reconnection costs a **TLS handshake**, which is CPU-heavy. The dashboards show p99 doubled, **24.3% of CFS periods throttled** (2,187 of 9,000 in 15 minutes), `kubectl top pod` around 410m, and no restarts. The CPU limit is clearly binding in bursts. There are three levers, and they are not equivalent. ## Lever 1: match parallelism to the limit Throttling far below the limit is usually **over-parallelism**: a runtime sized its worker pool from the node's core count, so dozens of threads drain a 150ms-per-100ms quota in a few milliseconds. Fix the cause first: - check what the runtime reports as its CPU count **from inside the pod**; - pass the limit to the application explicitly. The **Downward API** can expose `limits.cpu` as an environment variable through `resourceFieldRef`; with `divisor: "1"` the value is **rounded up** to whole cores, so 1500m becomes `2`; - set the runtime's own parallelism setting from that variable. This keeps the pod spec unchanged and often removes most of the throttling. It fails when the service genuinely needs more CPU than the budget allows. ## Lever 2: raise the limit If the counters still show throttling after the thread count matches, demand really exceeds 1.5 cores in bursts. Raise the limit, and review the **request** at the same time: the scheduler places pods by requests, so a limit far above the request means the extra CPU exists only when neighbours are idle. On a busy node the service can then be slow for a different reason: it gets only its request's share. ## Lever 3: remove the CPU limit Without `limits.cpu`, the kubelet sets **no CFS quota**: the container can use any idle CPU on the node, and its **request** still sets its weighted share when CPU is contended. Many teams run latency-sensitive services this way. The trade-offs are concrete: | Consequence | Why it matters | |---|---| | QoS class drops from **Guaranteed** to **Burstable** | Guaranteed needs requests equal to limits for CPU and memory in every container | | No exclusive cores from the kubelet's `static` CPU manager policy | Pinning requires a Guaranteed pod with an integer CPU request | | A namespace **LimitRange** may inject a default limit | The limit comes back at admission and throttling returns | | A Downward API `limits.cpu` now resolves to **node allocatable** | The runtime sees the node's cores again and over-parallelises | | Cost attribution and noisy-neighbour arguments change | Bursts use CPU that other teams' pods are not using | | A regulated cluster's admission policy may require limits | The change may simply be rejected | Removing the limit is safe only when **requests are honest**: they are what protects the neighbours. ## A decision sequence 1. **Confirm** throttling from the CFS counters, lined up with p99. 2. **Fix parallelism** to match the limit, redeploy, and re-measure over the same window. 3. If throttling remains, **raise the request and limit together** to what the burst needs. 4. **Remove the limit** only if the service needs short bursts well above its steady use, the requests are accurate, pinning is not required, and policy allows it. 5. **Keep the memory limit** in every case: CPU overuse only delays work, while unbounded memory growth can push the whole node into memory pressure and hurt other tenants. ## What not to do - Do not scale replicas as the only fix: every new replica has the same thread-to-quota mismatch. - Do not raise the limit without re-checking the thread count, because on bigger nodes the pool grows again. - Do not remove the limit and set the request to near zero: the pod then gets almost no CPU under contention. - Do not judge the fix by `kubectl top pod`; judge it by the throttling ratio and p99. - Do not treat the certificate window as a one-off: any event that makes every client reconnect at once, such as a rollout of the service's callers, produces the same burst, so the fix should hold for the next one too. Measure each change over a window that includes a burst, since a quiet hour will show zero throttling whatever the settings are.
- After removing the CPU limit, the team's WORKER_THREADS variable jumped to 48. Why?A Downward API `resourceFieldRef` on `limits.cpu` falls back to the node's allocatable CPU when the container sets no CPU limit. On a 48-core node the variable becomes about 48, so the runtime over-parallelises again. Without a limit, derive the thread count from `requests.cpu` or set it explicitly.
- Why is it considered safe to drop CPU limits but not memory limits?CPU is compressible: when it is short, the kernel shares it by the containers' requests, and the only cost is delay. Memory is not: a container that grows without a limit can push the node into memory pressure, so the kubelet evicts pods or the kernel kills processes, possibly in other teams' pods. Keeping the memory limit contains that damage to the service itself.
saying these in an interview costs you the question
- Scaling out replicas fixes throttling caused by oversized thread pools
- Without a CPU limit a container has no protection or priority at all
- Removing the CPU limit keeps the pod in the Guaranteed QoS class
- A higher CPU limit always means the container gets more CPU on busy nodes
- Memory limits can be dropped as safely as CPU limits