skip to content

If an Airbyte sync fails mid-stream, does the next attempt resume or start over?

level: seniorimportance: should knowfreq 52%

answer

  1. progress is remembered, not the row count
  2. the memory only advances behind accepted data
  3. some connectors only remember at the very end
  4. finished streams keep their own progress
  5. expect the overlap to arrive twice

basics

~20 s

It depends on whether progress was checkpointed. Airbyte only commits state for records the destination has accepted, so a connector that emits state during the read resumes from the last committed point; one that emits state only at the end restarts the stream.

solid answer

~50 s

Airbyte's resume point is the stream's committed **state** — the cursor value it is safe to restart from. Two conditions must both hold for a failure to be cheap. First, the source connector must checkpoint *during* the read rather than only at the end; many database and API connectors emit state after each slice or page, others emit once at completion. Second, the destination must have committed the records preceding that checkpoint, because Airbyte only advances state behind data the destination acknowledged. That ordering is what makes the pipeline at-least-once rather than at-most-once. With modern per-stream state, streams that finished keep their progress even if a later stream fails, so a retry does not redo them. Full-refresh streams generally restart from the beginning. Expect duplicates on resume: records committed before the failure and re-read after it appear twice — harmless under a deduped write mode, permanent under plain append.

code

json · 7 lines
json
{
  "type": "STREAM",
  "stream": {
    "stream_descriptor": { "name": "orders", "namespace": "public" },
    "stream_state": { "updated_at": "2026-08-20T04:11:07Z" }
  }
}

go deeper

for a junior

Recall that Airbyte remembers a per-stream resume point and that a failed sync retries from it rather than always starting over. Knowing that duplicates can result is enough at this level.

for a middle

Explain the ordering: the source emits state, the platform holds it until the destination confirms the records, then it is persisted. That ordering is what makes the pipeline at-least-once.

for a senior

Diagnose the real case: repeated attempts with identical counts mean the connector is not checkpointing, and the fix is narrower slices or a different connector, not more retries. Tie duplicate handling back to the chosen write mode.

for a principal

Own the operating envelope: which streams are allowed to be large enough that a failed sync is expensive, what backfills are windowed rather than one heroic run, and what the source system is permitted to absorb when state is cleared.

## What state is State is Airbyte's memory of how far a stream got. For a cursor-based incremental stream it is essentially the cursor value up to which data is known to be safely in the destination; for other connectors it can be a page token, a slice boundary, or a log position. It is stored by the platform per connection, and it is the only thing that makes the next sync incremental instead of a full re-read. ## When state is committed The crucial rule is that state advances **behind** the data. The source emits records and, at points of its choosing, a state value meaning 'everything before this is emitted'. The platform holds that value until the destination confirms it has committed the corresponding records; only then is state persisted. If the process dies between those two moments, the state stays where it was and the records are re-read next time. That is a deliberate at-least-once design: Airbyte would rather deliver a record twice than lose it. ## Whether the connector checkpoints at all This is the part candidates miss. Some connectors emit state repeatedly during a read — after each slice of time, each page, each batch of rows — so a failure at 80% costs you 20%. Others emit state only once, when the stream finishes; a failure at 99% costs you everything, and the attempt after it starts from the same place as the one before. When a huge stream keeps failing and never advances, this is usually why. The fix is connector-side (slice the read more finely, use a connector version that checkpoints, or reduce the stream to a narrower time window) rather than something you can configure away in the connection UI. ## Per-stream versus global state Modern Airbyte keeps state per stream: each stream carries its own cursor value, so if stream A completes and stream B fails, A's progress survives and the retry only redoes B. Older global-state connections stored one blob for the whole connection, and a failure could throw away progress across all of them. When you are debugging a connection that seems to redo finished work, the state format is a legitimate thing to check. ## What resuming means for the destination Resuming does not undo what was already written. Records committed before the failure are in the raw table, and the re-read after the failure appends them again. Under `Incremental | Append + Deduped` this is invisible — the primary key collapses the versions and the newest wins — which is a strong reason to prefer that mode for large, failure-prone streams. Under `Incremental | Append` the duplicates are permanent and every downstream consumer inherits them. Under a full-refresh overwrite stream the question is moot in a different way: the destination swaps in the new table only when the read completed, so a failed sync leaves the previous table intact and the next attempt re-reads everything. ## Retries and attempts A failing sync is retried automatically a bounded number of times before the sync as a whole is marked failed, and each attempt begins from the last committed state. This is why a well-checkpointed large stream can crawl forward across several attempts and eventually finish, while a poorly checkpointed one loops on the same first slice forever. When you see repeated attempts with identical record counts, the sync is not making progress and more retries will not help. ## Clearing state deliberately Sometimes you *want* to lose the resume point: after changing a cursor or primary key, after fixing a source bug that corrupted historical rows, or after a schema change that requires a rebuild. Clearing a connection's data and state (labelled *Reset* in older releases and *Clear* in current ones) wipes the stored state so the next sync reads the stream from the beginning; newer releases also offer a refresh that re-reads everything while keeping existing records in place. Treat it as an operational lever with a real cost — a full re-read of the source — not a routine fix. ## What to say in an interview Name the three variables: does the connector checkpoint mid-stream, has the destination committed the data behind that checkpoint, and is the state per stream. Then name the consequence: resume is cheap and duplicates are expected, so the write mode you chose determines whether those duplicates are a problem.

  • Why does Airbyte commit state only after the destination has accepted the records?
    Because committing earlier would make loss possible: if the process died between advancing the cursor and writing the data, the records would never be re-read. Ordering state behind confirmed writes makes the pipeline at-least-once — duplicates on resume, never a silent hole.
  • A 40-million-row stream fails at 90% and every retry restarts from zero. What is happening and what do you do?
    The connector is emitting state only at the end of the stream, so nothing is committed to resume from. Options: use a connector version that checkpoints per slice, split the stream into narrower bounded windows you sync sequentially, or switch that source to a mode that checkpoints. Adding retries changes nothing.
  • How do duplicates from a resumed sync interact with the stream's write mode?
    Under Incremental | Append + Deduped they vanish: the primary key collapses the re-read versions and the latest cursor wins. Under Incremental | Append they persist permanently in the destination. Under Full Refresh | Overwrite the question does not arise, because a failed sync never swaps its table in.
  • When would you deliberately clear a connection's state?
    After changing the cursor or primary key, after fixing a source-side bug that corrupted already-synced history, or when a schema change requires rebuilding the stream. It forces the next sync to read from the beginning, so treat it as a planned full re-read with a real cost to the source.

saying these in an interview costs you the question

  • Assumes every connector checkpoints progress mid-sync
  • Thinks state advances as soon as the source emits records
  • Says a resumed sync never re-delivers records
  • Believes more retries fix a stream that never checkpoints
  • Expects a failed full-refresh sync to leave the table half replaced

context