A gRPC server starts closing a caller's connections with the debug data too_many_pings — what did the caller's keepalive settings do wrong?
answer
- how a caller notices a silent peer
- it is a transport frame, not a call
- the receiving side keeps score
- idle connections are the usual trigger
- both sides must agree on the interval
basics
~20 sKeepalive pings were sent more often than the server tolerates, typically on connections with no active calls. The server counted strikes and, past its limit, sent a GOAWAY carrying ENHANCE_YOUR_CALM and the debug data too_many_pings.
solid answer
~30 sA gRPC client's keepalive sends an HTTP/2 `PING` every `GRPC_ARG_KEEPALIVE_TIME_MS` and declares the connection dead if no acknowledgement arrives within `GRPC_ARG_KEEPALIVE_TIMEOUT_MS` (default 20000 ms). `GRPC_ARG_KEEPALIVE_PERMIT_WITHOUT_CALLS` decides whether it pings while no call is in flight. The server enforces limits from the other side: `GRPC_ARG_HTTP2_MIN_RECV_PING_INTERVAL_WITHOUT_DATA_MS` (default 300000 ms) sets how often a caller may ping an idle connection, `GRPC_ARG_HTTP2_MAX_PINGS_WITHOUT_DATA` (default 2) caps pings with no data between them, and each violation is a strike. Past `GRPC_ARG_HTTP2_MAX_PING_STRIKES` (default 2) the server sends a `GOAWAY` with error `ENHANCE_YOUR_CALM` and debug data `too_many_pings`. Lowering the client interval without raising the server's tolerance produces exactly this.
code
http · 7 lines# caller -> server, on a connection with no active calls
PING (opaque-data, ACK=0)
PING (opaque-data, ACK=0)
PING (opaque-data, ACK=0)
# server -> caller, once the strike limit is exceeded
GOAWAY last-stream-id=<last accepted> error=ENHANCE_YOUR_CALM debug-data="too_many_pings"go deeper
Know that gRPC has a transport-level keepalive and that it is how a client notices a peer that disappeared without saying so.
Distinguish the ping interval from the acknowledgement timeout, and from a call's deadline. Explain why pinging an idle connection is treated differently.
Read the failure as a settings mismatch between two sides: name the strike mechanism and say which value you would change, on which side, and why.
Keepalive numbers are a cross-team contract. Decide fleet-wide defaults with the teams operating the servers, because a client-side value alone can only ever be half of one.
## Why a balancer needs keepalive at all A subchannel in `READY` is only a belief. If the worker behind it loses power, is partitioned away, or has its connection silently dropped by a stateful middlebox, nothing arrives to say so: no reset, no close, no error. The transport sits there looking healthy, the picker keeps handing it a share of every caller's calls, and each of those calls waits for its deadline. gRPC's answer is a transport-level keepalive built on the HTTP/2 `PING` frame. The client sends a ping, the peer must acknowledge it, and an unacknowledged ping is how a caller learns the peer is gone. The subchannel then leaves `READY`, the picker drops it from rotation, and the channel re-resolves. Keepalive is therefore not a nicety — it is the input that makes a client-side balancer's view of its endpoints true. ## The client's three settings - `GRPC_ARG_KEEPALIVE_TIME_MS` — how often to send a ping on a connection. - `GRPC_ARG_KEEPALIVE_TIMEOUT_MS` — how long to wait for the acknowledgement before declaring the connection dead. The default is **20000 ms**. - `GRPC_ARG_KEEPALIVE_PERMIT_WITHOUT_CALLS` — whether to ping at all when no call is in flight. This is the one that causes trouble, because a caller that pings an idle connection is generating traffic the server gets nothing for. Note which timeout this is. It is the wait for a ping acknowledgement, and it has nothing to do with the deadline a caller sets on a call: a call can exceed its deadline on a perfectly healthy connection, and a connection can fail keepalive while every call on it is well inside its own deadline. ## The server's three settings, and the strikes Pings are not free: each one is a frame the server must read and answer, and a large fleet of callers pinging aggressively is a denial-of-service pattern whether or not anyone intended it. So the server polices them: - `GRPC_ARG_HTTP2_MIN_RECV_PING_INTERVAL_WITHOUT_DATA_MS` — the minimum interval it will accept between pings on a connection carrying no data. The default is **300000 ms**, which is five minutes and is far longer than most people's client interval. - `GRPC_ARG_HTTP2_MAX_PINGS_WITHOUT_DATA` — how many pings it will accept with no data between them. The default is **2**. - `GRPC_ARG_HTTP2_MAX_PING_STRIKES` — how many violations it tolerates before acting. The default is **2**. A ping that breaks one of the first two rules is a **strike**. When the strikes exceed the limit, the server sends a `GOAWAY` frame with the error code `ENHANCE_YOUR_CALM` and the debug data `too_many_pings`, and closes the connection. ## Reading the failure The sequence that produces the report in the question is almost always this: 1. Someone lowers the client keepalive interval to detect dead workers faster. 2. `GRPC_ARG_KEEPALIVE_PERMIT_WITHOUT_CALLS` is on, so idle connections are pinged too. 3. The server's idle-ping interval is left at its five-minute default. 4. Pings on idle connections accumulate strikes; the server sends the connection-closing frame. 5. The client reconnects, resumes pinging at the same rate, and is closed again — a loop that looks like a flapping network and is entirely self-inflicted. The fix is a matching pair of settings, not a client-side change alone: raise the server's tolerance to at least the client's interval, or raise the client's interval to the server's tolerance. Because both sides are involved, a caller that cannot change the server it calls has only one of those options. ## What the connection-closing frame does and does not do `GOAWAY` stops new streams on that connection while the streams already running are allowed to finish. It is not a reset of individual calls, so the immediate blast radius is smaller than it looks: in-flight fingerprint matches complete, and only new calls have to wait for the replacement connection. What does hurt is the loop — a whole fleet reconnecting every few seconds spends its time re-establishing transports and re-resolving instead of matching fingerprints. ## Tuning it honestly - Set the keepalive interval from how long you are willing to keep routing calls to a dead endpoint, not from a wish for fast detection in the abstract. - Leave pinging on idle connections off unless the connections really do sit idle for long stretches and you really do need them proven. - Treat the ack timeout as a network-latency budget; the 20000 ms default is deliberately generous because a false positive tears down a working connection. - Agree the numbers with whoever operates the server. This is one of the few gRPC settings where a client-side value is meaningless without the server's matching one.
- Why does keepalive matter to a client-side balancer at all?It is how a caller notices a peer that vanished without closing anything. Until a ping goes unacknowledged the subchannel stays READY, so the picker keeps routing a share of every caller's calls to an endpoint that will never answer, and each of those calls burns its full deadline.
- Which client setting decides whether pings are sent on an idle connection?GRPC_ARG_KEEPALIVE_PERMIT_WITHOUT_CALLS. With it off, pings stop when no call is in flight, which is what a server's defaults expect. Turning it on is what usually starts the strike count, because idle-connection pings are exactly what the server's minimum interval polices.
- Do calls already running die when that connection-closing frame arrives?No. The frame stops new streams from being started on the connection while the streams already in progress run to completion. The damage from a strike loop is the repeated reconnect and re-resolution cost, not aborted work.
Knocking on a neighbour's door every thirty seconds to confirm someone is home. Do it a few times too often and they stop answering the door at all — which is precisely the information you were not trying to obtain.
saying these in an interview costs you the question
- Thinks keepalive pings are free, so a shorter interval is always safer
- Believes a server cannot object to how often a caller pings
- Confuses the keepalive acknowledgement timeout with a call's deadline
- Assumes the connection-closing frame aborts calls already in flight
- Thinks pinging on an idle connection is on by default