How do you tell apart kubectl's 'connection refused', 'x509: certificate has expired' and 'TooManyRequests' errors, and where does each one point?
answer
- which layer answered
- TCP vs TLS vs HTTP
- server cert vs client cert
- 401 means authentication, not TLS
- 429 comes from a live server
basics
~20 sConnection refused means nothing is listening, so the API server is down or the address is wrong. An x509 expiry means the TLS handshake failed, usually on an expired serving certificate. TooManyRequests (HTTP 429) means a live API server is rejecting load.
solid answer
~50 sEach error comes from a different layer, so read which layer failed. `The connection to the server … was refused` is TCP: nothing accepts connections at that address. Either kube-apiserver is down or restarting, or the kubeconfig points at the wrong host or port. `x509: certificate has expired or is not yet valid` is TLS: the server answered, but kubectl rejected its serving certificate. The usual causes are an expired certificate or a badly wrong clock. An expired *client* certificate in the kubeconfig looks different. The API server checks client certificates after the handshake, so kubectl prints `error: You must be logged in to the server (Unauthorized)`. `Error from server (TooManyRequests)` is HTTP 429 from a healthy API server that is protecting itself, usually through API Priority and Fairness. The fix there is to reduce the load, not to restart anything.
code
bash · 4 lineskubectl config view --minify -o jsonpath='{.clusters[0].cluster.server}'
kubectl get --raw='/readyz?verbose' -v=6
openssl s_client -connect 10.40.17.9:6443 </dev/null 2>/dev/null \
| openssl x509 -noout -enddatego deeper
Remember that the three messages come from three layers: nothing listening, a certificate rejected, or a server that is up but busy.
Explain the request path, TCP then TLS then authentication then flow control, and why an expired client certificate shows up as a 401 rather than an x509 error.
Show you would confirm with evidence: the kubeconfig server address, openssl dates on the serving certificate, readyz output, and the clients driving a 429.
Argue for monitoring that catches each class early: certificate-expiry alerts, API server availability probes, and dashboards of rejected requests by flow.
## Read the error by layer A `kubectl` call goes through several layers in order: DNS and TCP to the API server address, then the TLS handshake, then authentication and authorization, then admission and storage. An error message tells you **how far the request got**, and that is most of the diagnosis. | kubectl output | Layer | What it means | Where to look next | |---|---|---|---| | `The connection to the server <host:port> was refused - did you specify the right host or port?` | TCP | Nothing is accepting connections at that address | kubeconfig `server:` field, load balancer, the kube-apiserver process | | `Unable to connect to the server: … i/o timeout` | Network | Packets are dropped or not routed | Firewall, security groups, VPN, load balancer health | | `… x509: certificate has expired or is not yet valid` | TLS | kubectl rejected the **server's** certificate | Serving certificate expiry, node or laptop clock | | `… x509: certificate signed by unknown authority` | TLS | Wrong CA in the kubeconfig | `certificate-authority-data` in the kubeconfig | | `error: You must be logged in to the server (Unauthorized)` | Authentication | HTTP 401, credentials rejected | Expired client certificate or token | | `Error from server (Forbidden)` | Authorization | HTTP 403, RBAC denied the request | Role bindings | | `Error from server (TooManyRequests)` | Flow control | HTTP 429, a live server is shedding load | Request volume, noisy clients | ## Connection refused `connection refused` is a TCP reset: the host is reachable but no process is listening on the port. Common causes: 1. **kube-apiserver is down or restarting.** On a kubeadm cluster it is a static pod, so a bad manifest edit or a failing dependency keeps it in a restart loop. If the load balancer in front has no healthy backends, kubectl sees refused connections or timeouts. 2. **The wrong address.** A stale kubeconfig context, or the `localhost:8080` fallback kubectl uses when it finds no kubeconfig at all, gives the same message. kubectl can do nothing more here. The next step is on the control-plane node itself. ## x509 errors The TCP connection worked and the server presented a certificate, but **kubectl** rejected it. `certificate has expired or is not yet valid` usually means: - the API server's **serving certificate** has expired. kubeadm issues leaf certificates valid for one year, so a cluster that has not been upgraded or renewed for a year hits this; or - a **clock** is badly wrong, on the machine running kubectl or on the node. The **client** certificate in the kubeconfig fails differently. kube-apiserver asks for a client certificate during the handshake but only **checks** it in its authentication step. An expired client certificate is therefore rejected with **401**, and kubectl prints `You must be logged in to the server (Unauthorized)`. Checking and renewing certificates is covered by the cluster-PKI topic; for triage, it is enough to know which side's certificate the message is about. ## TooManyRequests (429) This is the most reassuring of the three: the API server is up, TLS works and you are authenticated. The server is **refusing work** because a priority level is saturated. On current Kubernetes that is **API Priority and Fairness**, which queues requests and rejects them with 429 when the queues are full. Things to know: - client-go (and so kubectl) **retries** a 429 that carries a `Retry-After` header, up to 10 times by default. The first symptom is often a slow kubectl, and the error appears later. - The cause is usually **load**: a controller stuck in a hot loop, many clients doing full LIST calls, or thousands of pods from a batch job all updating status at once. - Restarting the API server does not help and loses its in-flight work. Find the heavy client instead. Tuning the flow-control configuration is a separate topic. ## A quick decision list - Refused or timeout: go to the network and the API server process. - x509: check certificate dates and clocks. - 401 or 403: check credentials and RBAC. - 429: check who is generating the load. ```bash # Show the exact failing layer with full HTTP detail kubectl get --raw='/readyz?verbose' -v=6 ```
- kubectl hangs for a long time and then reports TooManyRequests. Why the delay?client-go retries a 429 that carries a Retry-After header, up to 10 times by default, sleeping between attempts. kubectl only reports the error after those retries run out. The slowness is the first symptom of flow control rejecting requests, so start by looking for the client or controller that is flooding the API server.
- The kubectl error is 'x509: certificate signed by unknown authority' rather than an expiry. What changed?The server's certificate is valid, but kubectl doesn't trust the CA that signed it. The kubeconfig's certificate-authority-data doesn't match the cluster's CA. Typical causes are a kubeconfig from a rebuilt cluster, a context pointing at the wrong cluster, or a proxy or load balancer terminating TLS with its own certificate.
saying these in an interview costs you the question
- Connection refused means the kubeconfig certificate expired
- An expired kubeconfig client certificate always shows up as an x509 error
- A 429 means the API server is down and needs a restart
- All kubectl errors mean the cluster's workloads are failing
- An x509 expiry error can be fixed by passing a different namespace