skip to content

A gRPC client makes a unary call and sets no deadline. What does the server assume, and why is that dangerous?

level: juniorimportance: must knowfreq 72%

answer

  1. unbounded is the default, not bounded
  2. nothing on the wire means no bound
  3. the server infers the limit from the request
  4. grpc-timeout is absent, not zero
  5. DEADLINE_EXCEEDED (4) is the bounded outcome

basics

~20 s

With no deadline the call carries no grpc-timeout request field, and the gRPC specification tells the server to assume an infinite timeout. One stalled dependency then holds caller threads and memory until something else breaks.

solid answer

~40 s

A deadline is a caller decision. When the calling code sets one, the client library converts it into the `grpc-timeout` request field, and the server knows exactly how long the answer is still wanted. When nobody sets one, that field is simply absent, and the gRPC specification says the server should assume an infinite timeout — the handler may run as long as it likes and the caller waits. That is how one slow dependency becomes an outage: every waiting call holds a request slot and the memory behind it, and the pressure spreads back through whoever called you. With a deadline, the call instead fails fast and predictably with `DEADLINE_EXCEEDED (4)`, and the server learns it can stop working.

code

http · 16 lines
http
# no deadline - nothing bounds this call
:method POST
:scheme https
:path /pharmacy.v1.Eligibility/CheckPrescription
:authority eligibility.svc.internal
te: trailers
content-type: application/grpc+proto

# the same call, bounded at 900 milliseconds
:method POST
:scheme https
:path /pharmacy.v1.Eligibility/CheckPrescription
:authority eligibility.svc.internal
te: trailers
content-type: application/grpc+proto
grpc-timeout: 900m

go deeper

for a junior

Know that gRPC has no default deadline: if your code does not set one, none is sent and the server assumes it may take as long as it needs. Setting one is a single call-site decision.

for a middle

Explain the mechanism: a deadline becomes the grpc-timeout request field, the server reads it, and expiry surfaces as DEADLINE_EXCEEDED (4) on the caller while the server is told to stop work.

for a senior

Show the failure shape. Describe how unbounded calls accumulate request slots and memory under a stalled dependency, and how you would choose the bound from the caller's own latency requirement rather than the dependency's average.

for a principal

Treat 'every outbound call carries a bound' as a property of the platform rather than a habit of individual teams, and be able to say who owns the numbers and how a violation is detected before production finds it.

## The bound is a client decision, and it travels A **deadline** on a gRPC call is an instant, chosen by the caller, after which the answer has no value. The generated client turns that instant into a duration — the time still remaining when the request is written — and sends it as one request metadata field, **`grpc-timeout`**. The server reads that field and knows precisely how long the caller is prepared to wait. Nothing in the protocol fills that field in for you. If the calling code never sets a deadline, the client library sends no `grpc-timeout` field at all, and the gRPC specification tells the receiving server to **assume an infinite timeout**. There is no protocol-level default, no implicit thirty seconds, nothing. ## What an unbounded call costs Picture an eligibility check at a pharmacy counter. A prescription is scanned, a person is standing there, and behind the one call the counter makes sit a formulary lookup, a plan-coverage lookup and a prior-authorisation lookup. If the answer takes more than about a second it is worthless — the pharmacist has already started doing something else. Now the prior-authorisation dependency stalls. With no deadline anywhere: - the counter's call never returns, so the software at the counter keeps waiting on a response nobody will use; - the eligibility server keeps its handler alive, holding the request, the partially-built response and the open call it made downstream; - each new scan adds another stuck call, so concurrency climbs with arrival rate and is never relieved by completion; - memory and request slots run out, and calls that had nothing to do with prior authorisation start failing too. That is why interviewers ask this of juniors: the omission is invisible in the code and catastrophic in production. Nothing in the call site shows that a bound is missing, because a missing bound looks exactly like normal code. ## Deadline set versus deadline absent | | deadline set | no deadline set | |---|---|---| | who decides | the calling code | nobody | | on the wire | `grpc-timeout` in the request metadata | the field is absent | | server's assumption | finish within this duration | an infinite timeout | | what the caller gets on expiry | `DEADLINE_EXCEEDED (4)` | whatever eventually returns, or nothing | | failure shape | fast, uniform, attributable | slow, unbounded, contagious | ## What DEADLINE_EXCEEDED means, and to whom `DEADLINE_EXCEEDED` is status code **4** in gRPC's status enum. When the bound expires before a response arrives, the caller's call ends with that status — the client library does not keep waiting in the hope that something turns up. The important half is the other side. Because the bound travelled in `grpc-timeout`, the server knows it too, and a server framework will normally signal the handler that the call is over so it can stop doing work nobody wants. Two consequences follow: 1. A deadline is **not** merely a local give-up timer on the client. It is a fact both ends share, which is exactly what a plain client-side wait is not. 2. Work already committed before the deadline expired is **not** undone by it. If the handler had already written a record, the deadline does not roll that back; the status tells you the answer did not arrive in time, not that nothing happened. ## Choosing the value The honest number comes from the caller's own requirement, not from the dependency's average latency. At the counter the requirement is 'about a second', so the bound is set from that and the call is written to fail cleanly when it cannot be met. A bound generous enough that it never fires is a bound that does nothing; a bound below the service's normal response time turns healthy calls into failures. The useful range is between those two, and it is a product question before it is a technical one. A last point that catches people out: a deadline is **per call**. Setting one on a client object does not retroactively bound calls already in flight, and it is not a property of the underlying long-lived connection. Every call carries its own `grpc-timeout` field, or none.

  • Does the gRPC server learn that the caller's deadline expired, or only the client?
    Both. The bound travelled in the `grpc-timeout` request field, so the server knows the same instant and a server framework normally signals the handler that the call is over. That is the difference between a shared deadline and a purely local give-up timer.
  • If a gRPC call ends with DEADLINE_EXCEEDED (4), can you conclude the server did nothing?
    No. The status says the answer did not arrive within the bound, not that no work happened. A handler may have completed a write moments before the deadline expired. Whether that matters is a property of the operation, not of the status code.
  • Is a deadline a property of the gRPC client object or of the individual call?
    Of the individual call. Each request carries its own `grpc-timeout` field, computed from the caller's deadline at the moment the request is written. Calls already in flight are unaffected by a bound set afterwards, and the long-lived connection underneath has nothing to do with it.

A deadline is the moment the person at the counter gives up and walks out. Anything the system does after that instant costs money and helps nobody.

saying these in an interview costs you the question

  • Thinks the client library applies a sensible default deadline
  • Says an absent grpc-timeout field makes the call fail immediately
  • Confuses a per-call deadline with a transport keepalive timer
  • Assumes the server aborts long-running calls on its own
  • Believes a deadline only affects the client's own waiting