When would you choose Full Refresh | Overwrite over incremental for an Airbyte stream?
answer
- ask about the source before asking about the size
- no trustworthy column to read by
- deletes cannot be expressed by reading forward
- one mode repairs the past, the other never revisits it
- cheap most nights, honest once a week
basics
~20 sChoose overwrite when the stream is small enough to re-read and correctness matters more than cost: no trustworthy cursor, rows mutated without bumping one, or hard deletes that must disappear. Overwrite self-heals every run; incremental never revisits what it already passed.
solid answer
~50 sThe deciding questions are about the source, not the volume alone. Reach for **Full Refresh | Overwrite** when there is no cursor you trust, when rows change without advancing one, when hard deletes must be reflected, or when the stream is small enough that a full re-read is cheaper than the operational cost of getting incremental right. Its real advantage is that it is self-healing: every sync discards prior state, so yesterday's drift, a bad backfill and a mis-set cursor are all corrected without intervention. Reach for **Incremental | Append + Deduped** when the stream is large, the source has a reliable modification cursor, and a full read would strain an API's rate limits or a production database. A hybrid is often the honest answer: incremental on the schedule, plus a periodic full refresh — weekly, monthly — to re-establish truth. State the tradeoff explicitly as cost against correctness, and note that overwrite keeps no history.
code
text · 7 linesconnection: prod-postgres -> warehouse
countries (900 rows, no updated_at) Full Refresh | Overwrite
price_list (12k rows, hard deletes) Full Refresh | Overwrite
orders (400M rows, db-set updated_at) Incremental | Append + Deduped
page_events (append-only, monotonic id) Incremental | Append
feature_flags (small, audit history wanted) Full Refresh | Appendgo deeper
Recall the basic contrast: overwrite re-reads everything and replaces the table, incremental reads only what the cursor advertises. Small tables and missing cursors point to overwrite.
Explain the mechanics behind the choice: why a full refresh self-heals, why an incremental stream never revisits what it passed, and why deletes are invisible to cursor-based reads.
Show the production reasoning — source load, API limits, sync duration versus schedule — and propose the hybrid of incremental plus a periodic full refresh, with a way to detect drift between them.
Own the estate-level policy: which streams may claim mirror-of-source semantics, what full-refresh load the source systems are budgeted to absorb, and when a stream must be escalated to CDC instead of being refreshed harder.
## Frame it as source properties, not table size Candidates usually answer 'overwrite for small tables, incremental for big ones'. Size matters, but it is the third question, not the first. The first two are about whether incremental extraction can even be correct for this source. **Does a trustworthy cursor exist?** The stream needs a column that advances on every write path, is non-nullable, and is not bypassable by scripts or a second writer. Many API endpoints expose nothing suitable; many database tables have an `updated_at` maintained by one code path out of three. Without such a column, an incremental stream is quietly lossy, and no amount of tuning fixes it. **Do deletes matter?** Cursor-based extraction cannot express a hard delete: the row is gone, so it carries no cursor value to be read. If the destination must not serve rows that no longer exist at the source — entitlements, prices, active-customer lists — either the source emits deletions as records (as CDC-enabled database sources do) or the stream is refreshed in full. **Only then: what does a full read cost?** Cost lands in three places — the source system (query load on a production database, quota and rate limits on an API, wall-clock time), the warehouse (the load and any typing work), and the schedule (a stream that takes six hours cannot sync hourly). ## What overwrite buys you Self-healing is the property worth naming. Because each sync discards the previous contents, every class of historical corruption — a bad cursor, a mangled backfill, a source-side fix applied to old rows, a period of silent loss nobody noticed — is repaired automatically on the next run. Incremental streams have the opposite property: everything the cursor passed over is wrong forever until someone deliberately intervenes. Overwrite also needs no primary key and no cursor, which removes the two configuration choices that cause the most silent damage. Destinations implement overwrite as a load followed by a swap, so a failed sync leaves the previous table intact rather than a half-loaded one — but consumers should still expect the table's contents to change wholesale at sync time, and long-running queries can straddle a swap. ## What overwrite costs you You pay the full extraction on every run, which is exactly the load a busy source cannot absorb hourly. You keep no history: the destination knows only the current snapshot, so any question about what a row looked like last Tuesday is unanswerable unless something else recorded it. And you cannot express intra-day change on a stream that only refreshes overnight. ## The middle options `Full Refresh | Append` keeps the snapshots — one full copy per sync, distinguishable by the extraction timestamp. Good for auditing a small reference table's evolution, and a storage mistake for anything large. `Incremental | Append` is right for genuinely append-only sources such as event logs, where nothing is ever updated and deletes do not occur. `Incremental | Append + Deduped` is the workhorse for large mutable tables with a good cursor, and the safest incremental choice because replays and boundary re-reads collapse harmlessly. ## The hybrid worth proposing Run the stream incrementally on its normal schedule and refresh it fully on a slower cadence, using a scheduled clear-and-resync or a full-refresh run. This gives cheap freshness most of the time and a periodic re-establishment of truth that bounds how long any drift can persist. It is the answer that shows production experience, and the follow-up an interviewer will ask is what the full re-read costs the source — have a number-shaped answer about duration and load, even if the actual numbers are yours to measure. ## Per stream, not per connection Remember that the choice is per stream. A single connection to a production database can overwrite a dozen small reference tables nightly while running the two enormous transactional tables incrementally. Applying one mode to every stream in a connection is a smell, not a standard. ## Summarising the decision No trustworthy cursor, or deletes that must propagate: overwrite while the volume allows it, and escalate to a CDC-capable source when it stops allowing it. Reliable cursor and volume that hurts: incremental with deduplication. Either way, say out loud what you are trading — freshness and source load against correctness and history — because that reasoning is what the question is testing.
- The table is too large to overwrite but has no reliable cursor. What now?Stop trying to solve it in the connection. Either the source gains a database-maintained modification column or soft-delete flag, or you move that stream to a CDC-capable source connector so changes and deletes arrive from the log instead of a polled column. A windowed partial refresh of recent partitions is a stopgap, not a fix.
- How do you decide how often to run the periodic full refresh in a hybrid setup?Bound it by how long you can tolerate undetected drift and by what the source can absorb. Measure the full read's duration and load, schedule it in a quiet window, and tighten the cadence if reconciliation checks keep finding differences. It is a risk-versus-load decision you should be able to defend, not a default.
- What breaks for consumers when a stream switches from overwrite to deduped incremental?Deletes stop propagating, so stale rows accumulate; and the table starts reflecting only what the cursor advertises. Consumers who relied on the table being an exact mirror need to know. It is a contract change worth announcing, plus a reconciliation check to prove the new mode is keeping up.
Incremental is patching a map from change notices; overwrite is resurveying the whole territory. Patching is cheap until a change notice was never issued, and only the resurvey finds what quietly disappeared.
saying these in an interview costs you the question
- Decides purely on row count without checking cursor quality
- Assumes incremental is always the more professional choice
- Forgets that overwrite is what makes deletes disappear
- Applies one sync mode to every stream in a connection
- Claims overwrite preserves history of previous syncs