skip to content

In a gRPC bidirectional streaming call, must the server wait for the client's last message before replying?

level: middleimportance: nice to knowfreq 32%

answer

  1. two sequences, one call
  2. neither waits on the other
  3. order within a direction only
  4. no request-to-response pairing
  5. lockstep assumption deadlocks both sides

basics

~20 s

No. The two directions of a bidirectional gRPC call are independent: the server may answer after the first message, interleave replies with incoming messages, or wait for the whole sequence. The application protocol decides, not gRPC.

solid answer

~40 s

The two message sequences in a bidirectional call run **independently**. The server may send its first response after reading one request message, interleave freely while requests keep arriving, or read everything and answer at the end — all three are legal for the same method. gRPC guarantees ordering *within* each direction: messages arrive in the order that side sent them. It guarantees nothing *across* directions, so there is no request-to-response pairing unless the application builds one by putting a correlation field in the messages. The trap is assuming lockstep. If the client waits for a reply before sending its next message while the server reads the whole sequence before replying, both sides wait forever and the call only ends when its deadline expires.

code

pseudocode · 9 lines
pseudocode
// one bidirectional call, six hours, no lockstep

t0   client -> Sample(row 1)
t1   client -> Sample(row 2)
t2   server -> Correction(moisture offset)   // answered before the client finished
t3   client -> Sample(row 3)
t4   client half-closes                      // request sequence over
t5   server -> Correction(yield scale)       // response direction still open
t6   server -> status                        // one status ends the call

go deeper

for a junior

Know that both sides can send whenever they like in a bidirectional call, and that this is what makes it different from the other three kinds.

for a middle

Explain what is guaranteed — ordering within each direction — and what is not: any pairing between a response and the request that prompted it. Say who supplies the missing pairing.

for a senior

Show the deadlock: a client waiting for a reply and a server waiting for end-of-sequence will sit there until the deadline. Then say that the cure is writing the turn-taking rules down, not a protocol feature.

for a principal

The point to make is that a bidirectional call ships an empty application protocol. Deciding who speaks when, and how replies are correlated, is a contract your teams must own as deliberately as the message schema.

A bidirectional streaming call gives each side its own sequence of messages, and those two sequences are **independent**. Nothing in gRPC pairs a response with the request that prompted it, and nothing forces one side to wait for the other. That freedom is the entire reason the kind exists — and it is also where the design mistakes live. ## What independence actually allows For one and the same bidirectional method, all of these are legal: - The server reads one message and immediately sends five responses. - The server sends a response before it has read anything at all — for instance an initial configuration message as soon as the call opens. - The server reads the client's entire sequence, waits for the half-close, and only then sends anything, behaving exactly like a client-streaming method. - The two sides interleave in any pattern at all, for hours, in both directions at once. The method kind fixes only *that* each side may send a sequence. **What that sequence means, and when each side speaks, is the application's protocol** layered on top. ## What is guaranteed and what is not | Property | Guaranteed? | |---|---| | Messages in one direction arrive in the order that side sent them | Yes | | A response corresponds to a particular request message | No — build it yourself | | One response per request message | No | | A total order across both directions | No | | The server has read message *n* when it sends response *n* | No | The second row is the one that catches people. If the client sends ten queries on the request sequence and ten answers come back, nothing in the protocol says the third answer belongs to the third query. If the answers can arrive out of order — and in a genuinely concurrent server they can — the application has to carry a correlation field inside its own message types and match on it. Falling back on arrival order is an assumption, not a guarantee. ## After a half-close The independence extends to the end of the call. The client may finish its request sequence with a half-close while the server keeps sending responses for as long as it wants; the response direction is untouched by the client's half-close. This is the common "upload everything, then keep listening" pattern, and it only works because the two directions end separately. The server's own way of finishing is to end the call with its single status, and that ends the whole call, both directions at once. ## The deadlock this creates The damaging misreading is that a bidirectional call is a strict question-and-answer alternation. Two sides written under that assumption can wedge each other permanently: 1. The client sends a message and then waits for a response before sending the next one. 2. The server was written to read the client's *entire* sequence before it replies to anything. 3. The server waits for a request sequence that will never end, because the client is waiting for a reply that will never come. Neither program has crashed, neither has an error to log, and the call ends only when its deadline expires. The fix is not a protocol feature; it is deciding, explicitly, which side speaks when, and writing both ends to that decision. Concretely, a bidirectional protocol is worth writing down in a sentence or two before any code: "the client sends samples continuously; the server sends a correction whenever it has one; neither side waits for the other." ## When the independence is the point The harvester is a good example. The machine sends a yield sample every few seconds for hours. The back office sends a moisture-calibration correction whenever its model produces one — perhaps three times in six hours, perhaps not at all. There is no correspondence between the two sequences and no reason to invent one. A client-streaming call could not carry the corrections at all, since the server gets exactly one message and only after the client is finished. Two separate calls would work, but then the correction feed and the upload have separate lifetimes, separate failures and no shared outcome. One bidirectional call gives both flows one lifetime and one ending.

  • If nothing pairs responses with requests, how do you match an answer to the query that caused it?
    Put a correlation field in your own message types — a request identifier the client sets and the server echoes — and match on that field rather than on arrival order. gRPC orders each direction independently and makes no promise across them, so arrival order is only reliable if the server itself guarantees it, which is an application decision.
  • Can the server send a message before it has read any request message?
    Yes. The response direction opens with the call, so the server can send immediately — a common use is an initial settings or acknowledgement message that tells the client the session is ready. Nothing requires a request message to have arrived first.

saying these in an interview costs you the question

  • Assumes the two directions alternate in strict lockstep
  • Expects exactly one response per request message
  • Thinks the server must read everything before it may reply
  • Believes messages from both directions share one global order
  • Matches replies to requests by arrival order with no correlation field