skip to content

When should a gRPC method server-stream a large result set instead of returning it in paged unary calls?

level: seniorimportance: should knowfreq 52%

answer

  1. how it is produced, not how big
  2. does the result exist yet
  3. unbounded feeds have no last page
  4. a cursor is a position you can return to
  5. one long call dies with its server

basics

~20 s

Stream when the results are produced over time, are large or unbounded, or the caller should see the first one before the last exists. Page with unary calls when the caller needs to resume from a known position or hold no long-lived call.

solid answer

~50 s

Server-streaming earns its place when the result is **produced incrementally**, is large enough that materialising it whole is wasteful, or is genuinely unbounded — a live feed has no last page. The caller then sees the first result as soon as it exists rather than after the last one does. Paged unary calls stay the honest answer when the caller needs to **resume from a known cursor** after losing its place, when each page is independently useful or repeatable, or when nothing on the path should hold a call open for minutes. The judgment interviewers are listening for is that a stream is one long-lived call with one outcome: it dies with the server instance holding it, has no position to resume from unless your messages carry one, and cannot be re-entered in the middle.

go deeper

for a junior

Know the two options exist and the basic rule: stream results that arrive over time, page results that already exist and are browsed.

for a middle

Justify the choice from how the result is produced rather than from its size, and say what a streaming call gives the caller that a page does not: the first result before the last one exists.

for a senior

Name the costs both ways — no resumption point in a stream unless your messages carry one, a handler occupied for the call's whole life, and every call in flight lost on a restart — and show the same data can justify either design.

for a principal

The call worth owning is whether long-lived calls belong in this system at all, given how often it deploys and how callers recover. That answer sets a convention many services will follow.

"Just stream it" and "just page it" are both real designs, and the difference between an engineer who has operated them and one who has not is that the first can say when each is wrong. The choice is not about the size of the result; it is about how the result is produced and what the caller has to be able to do afterwards. ## When server-streaming earns its place - **The results are produced over time.** A yield map computed row by row, or an analysis that emits findings as it discovers them, has nothing to page: page two does not exist yet when page one is requested. Streaming lets the first result reach the caller the moment it exists. - **The sequence is unbounded.** A live feed of corrections while a machine is cutting has no last page and therefore no page count, no offset and no terminal condition to code against. - **Materialising the whole result is the expensive part.** If assembling the complete answer means holding it all at once on the server, a stream lets the server emit and forget. - **The caller consumes in order and only once.** A batch job writing every row into a file wants exactly a sequence; a cursor would be ceremony around a loop it is going to run to completion anyway. ## When paged unary calls are the honest answer - **The caller must be able to resume from a position.** Each page is an independent call identified by a cursor, so a caller that lost its place can ask for the same position again. A broken stream has no position at all unless the application put one inside its own messages — the protocol does not give you one. - **The caller wants arbitrary access.** "Page seven", "the newest twenty", "skip to this key" are natural for a paged call and unnatural for a sequence you can only read forwards. - **Nothing on the path should hold a call open.** A long call occupies a server-side handler for its whole life and ends when that instance goes away, so a routine restart kills every call in flight. Short calls simply continue against whatever instance answers next. - **The result is small and already in hand.** A few hundred records that exist the moment the request arrives are one unary response. Streaming them adds a sequence to consume and an ending to handle for no benefit at all. | Question to ask | Server-streaming | Paged unary | |---|---|---| | Does the result exist all at once? | no — emitted as produced | yes | | Can the caller ask to resume at a position? | only if your messages carry one | yes, by cursor | | How long is a single call? | as long as the result takes | short | | What does a server restart cost? | every call in flight ends | the next page goes elsewhere | | Is a total count knowable? | often not | usually | ## The harvester, both ways When the office asks for a finished field's yield map — two hundred thousand rows that exist already, viewed in a grid where the operator jumps around — paged unary calls fit: each page is a short call, the viewer can re-request a page it lost, and a redeploy is invisible. When the same map is being *computed* row by row from a run that finished ten seconds ago, and the office wants to watch it fill, server-streaming fits: there are no pages yet, and the first rows are useful before the last exist. So the same data supports both answers, and the deciding question is not "how much" but **"does it exist yet, and what must the caller be able to do after it loses its place?"** ## Two mistakes worth naming 1. **Streaming as a way to avoid designing a cursor.** A stream makes the pagination problem disappear from the API and reappear in the failure path. When the call breaks at row 140,000, the caller starts again at row one unless the messages carried a position. That position is a cursor; you did not avoid it, you just did not design it. 2. **Paging a live feed.** Polling page after page of a sequence that is still growing produces duplicates, gaps at the boundary, and a per-call cost paid forever. If there is no last page, a paged call is answering a question the data cannot support. ## Saying it in an interview The compact version: *stream when the result is produced over time or has no end; page when the caller needs positions it can return to.* Then add the cost — a stream is one long-lived call with one outcome, tied to one server instance for its life — and you have shown the judgment rather than the preference.

  • A server-streaming call breaks after 140,000 of 200,000 rows. Where does the caller restart?
    At the beginning, unless the application put a position inside its own messages — a row key or sequence number the caller remembers and can send in a new request. The call itself offers no resumption point: it is one call, it ended, and a new call starts from whatever the request says.
  • Is a large result on its own a reason to stream?
    No. A large result that exists all at once and is consumed in order can be a stream, a single response or a set of pages — size alone does not decide. What decides is whether the result is produced incrementally, whether it has an end, and whether the caller needs positions it can return to.
  • What does a rolling restart cost each design?
    A streaming call ends when the instance running it goes away, so every call in flight fails and each caller must start again. Paged unary calls are short, so an in-progress page either completes or is re-issued, and the next page is simply served by another instance. That difference is often what settles a borderline choice.

saying these in an interview costs you the question

  • Says streaming always beats paging for a large result
  • Assumes a broken stream resumes where it stopped
  • Uses a stream to avoid designing a cursor
  • Pages a feed that has no last page
  • Streams a small result that already exists in full
  • Ignores that a long call dies with the instance serving it