skip to content

What are the default, committed, buffered and pending streams in the BigQuery Storage Write API?

level: middleimportance: should knowfreq 55%

answer

  1. visibility and duplication are the two axes
  2. one stream needs no creation at all
  3. offsets are what buy exactly-once
  4. one type reproduces load-job atomicity without files

basics

~20 s

They differ in when rows become visible and what guarantee you get. The default stream is at-least-once and always available; committed streams make rows visible on append and support exactly-once via offsets; buffered streams hold rows until you flush; pending streams hold everything until one atomic commit.

solid answer

~50 s

The BigQuery Storage Write API exposes one implicit stream and three stream types you create yourself. The **default stream** (`_default`) needs no creation, accepts writes from many concurrent clients, and gives at-least-once semantics — no offsets, so a retried append can duplicate. A **committed** stream you create with `CreateWriteStream`; rows are readable as soon as they are appended, and because you can supply an offset per `AppendRows` request, the service rejects duplicate offsets and you get exactly-once. A **buffered** stream writes rows that stay invisible until you advance the flush cursor with `FlushRows` — row-level visibility control, mostly used by connectors. A **pending** stream buffers everything invisibly until you call `BatchCommitWriteStreams`, at which point all rows appear at once: batch semantics without files. Choose default for high-throughput at-least-once, committed for exactly-once streaming, pending for atomic batch.

code

text · 4 lines
text
default   : append -> visible immediately        (at-least-once, no offsets)
committed : append(offset) -> visible immediately (exactly-once via offsets)
buffered  : append -> invisible -> FlushRows(offset) -> visible
pending   : append -> invisible -> Finalize -> BatchCommitWriteStreams -> all visible

go deeper

for a junior

Know that the BigQuery Storage Write API has a default stream you can write to immediately, and that other stream types exist for stronger guarantees.

for a middle

Be able to place all four on the visibility/guarantee axes and name the calls involved — CreateWriteStream, AppendRows with an offset, FlushRows, BatchCommitWriteStreams.

for a senior

Show you would pick a stream type from the pipeline's duplication tolerance and throughput target, and that you know exactly-once costs per-stream bookkeeping and throughput, not just a flag.

for a principal

Frame it as a guarantee budget: decide organisationally where exactly-once is genuinely required versus where at-least-once plus downstream deduplication is cheaper to run and easier to operate.

## Why there are four The Storage Write API unified BigQuery's batch and streaming ingestion behind one gRPC service. To cover both, it had to expose different *visibility* and *guarantee* models over the same append mechanism. Each stream type is one point in that space. The common machinery is the same everywhere: you open a connection, send `AppendRows` requests carrying rows serialised as protocol buffers, and the service acknowledges each request. What differs is when a reader of the destination table can see those rows, and whether a retried append can duplicate them. ## The default stream Every table has an implicit stream named `_default`. You do not create or finalize it, many writers can use it at once, and rows are queryable essentially as soon as they are acknowledged. It does **not** accept offsets, and therefore gives **at-least-once** semantics: if your client retries an append whose acknowledgement was lost, the rows land twice. This is the right default for high-volume telemetry where a small duplicate rate is tolerable or is cleaned up downstream, and it is the simplest thing to operate — no stream lifecycle, no offset bookkeeping. ## Committed streams You create one with `CreateWriteStream(type=COMMITTED)`. Rows are visible to readers as soon as they are committed to the stream, so freshness matches the default stream. The difference is that each `AppendRows` request may carry an **offset** — the row position this batch starts at. The service tracks the stream's current offset and rejects an append whose offset has already been written, so a retry after an ambiguous failure is a no-op rather than a duplicate. That is how **exactly-once** streaming is achieved. The price is bookkeeping: one stream per writer, offsets tracked by the writer, and `FinalizeWriteStream` when the writer is done. Throughput per stream is also lower than piling everything into the shared default stream, so exactly-once is a real cost, not a free flag. ## Buffered streams A buffered stream accepts appends that are **not yet visible**. Visibility advances only when you call `FlushRows` with an offset — everything up to that offset becomes readable. This gives row-level control over when consumers see data, which is what connectors and frameworks want when they need to align BigQuery visibility with their own checkpointing. Most application code never needs it; it exists for integrations. ## Pending streams A pending stream buffers all appended rows invisibly. Nothing is readable until you `FinalizeWriteStream` and then `BatchCommitWriteStreams` — one call that can commit **several** streams together, atomically. At that moment every row from every committed stream becomes visible at once. That is exactly the semantics of a load job, achieved without writing files: parallel workers each own a pending stream, all of them finish, one commit publishes the whole batch, and a failure anywhere means nothing was published. It is the right choice when the data is already in your process (so writing it to Cloud Storage first would be pure overhead) but you still want batch atomicity. ## Choosing - Tolerate duplicates, want the least machinery and the most throughput → **default stream**, plus downstream deduplication if needed. - Must not duplicate, need rows visible continuously → **committed** stream with offsets. - Need to control visibility at a checkpoint boundary from a framework → **buffered**. - Want a batch to appear all at once or not at all → **pending**. ## What is the same across all four Billing is by bytes ingested in every case — stream type changes semantics, not cost. Rows land first in a write-optimised buffer and are reorganised into columnar storage in the background; queries read both, so freshness is not affected by that conversion. Schema is negotiated per connection from the protocol-buffer descriptor you supply, so a schema change on the table generally means reopening the connection with an updated descriptor. ## The legacy path for contrast Before the Storage Write API, streaming meant the REST method `tabledata.insertAll`, which sends JSON rows and offers only **best-effort** deduplication via an `insertId` per row within a short window. There is no exactly-once mode, no pending semantics and no atomic commit. New work should use the Storage Write API; recognising `insertAll` in an existing pipeline and knowing why it is being replaced is the interview-relevant part.

  • Why does the default stream not support offsets?
    Because it is shared: many concurrent writers append to the same implicit `_default` stream, so there is no single writer whose position an offset could describe. That sharing is what makes it simple and high-throughput, and it is exactly why it can only offer at-least-once. Exactly-once requires an application-created stream that one writer owns.
  • How would several parallel workers publish one batch atomically with the Storage Write API?
    Each worker creates its own pending stream and appends its share of the rows, then finalizes its stream. When all workers report success, the coordinator calls `BatchCommitWriteStreams` listing every stream; all rows become visible in one step. If any worker fails, the coordinator simply never commits, and nothing was ever visible.

saying these in an interview costs you the question

  • Believing the default stream gives exactly-once delivery
  • Thinking pending-stream rows are queryable before the commit
  • Assuming stream type changes what you are billed
  • Treating insertId deduplication as an exactly-once guarantee
  • Creating a committed stream but never supplying offsets, then expecting no duplicates

context