In Census, how does the Mirror sync behavior differ from Update or Create?
answer
- one behavior can remove records
- think rsync with the delete flag
- absence in the model becomes an action
- upsert never takes anything away
- a shrinking model is a deletion event
basics
~20 sUpdate or Create only ever inserts or updates records present in the model. Mirror also removes or archives destination records that have dropped out of the model, so the destination ends up matching the model exactly — including its absences.
solid answer
~50 sBoth behaviors write the rows your model contains, but they disagree about rows it no longer contains. **Update or Create** is an upsert: a row present in the model is inserted if the destination has no matching record and updated if it does, and a row that disappears from the model is simply left alone downstream. **Mirror** makes the destination a faithful reflection of the model, so a record that leaves the model is deleted or archived on the other side, depending on what the destination supports. Mirror is what you want for a suppression list or an ad audience, where "no longer qualifies" must actually mean removal. It is dangerous on records humans depend on: a filter bug or a failed upstream job that shrinks the model will quietly delete rows in the CRM.
code
text · 8 linesmodel rows at run 1: [email protected], [email protected], [email protected]
model rows at run 2: [email protected], [email protected] -- bob dropped out
Update or Create -> alice updated, carol updated,
bob left untouched in the destination
Mirror -> alice updated, carol updated,
bob deleted or archived in the destinationgo deeper
Remember the headline: one behavior only adds and updates, the other also takes away. Know that removing a row from the source model is what triggers downstream deletion under the second one.
Explain all three cases a behavior must answer — unseen row, matched row, vanished record — and give a destination where propagating absence is the point, such as an ad audience or a suppression list.
Demonstrate the operational instinct: a shrinking model reads as a delete instruction, so gate the sync on volume tests, trigger it from the transformation job, and prefer a status field over destructive semantics where one exists.
Own the blast-radius framing — choose sync semantics by what a wrong run costs in each destination, and set an organizational default that destructive behaviors require a review and a quality gate.
## The one difference that matters Every reverse-ETL sync answers three questions: what to do with a model row the destination has never seen, what to do with one it already has, and what to do about a destination record the model no longer mentions. Census's sync behaviors differ almost entirely on the third question. **Update or Create** is a plain upsert. New row, matched by the sync's unique identifier, becomes an insert; already-present row becomes an update of the mapped fields. A record that vanishes from the model is untouched — the destination keeps whatever it last received, forever, until something else changes it. **Mirror** adds the removal leg. The destination is made to reflect the model's membership: rows that left the model cause the corresponding record to be deleted or archived downstream, subject to what the destination's API actually supports. Some destinations hard-delete, some archive or set an inactive flag, some (like ad-platform audiences) simply drop the member from the list. The other behaviors are the same question answered more narrowly. **Update Only** refuses to insert, so it can enrich records that already exist but never creates new ones. **Create Only** inserts new records and never touches existing ones. Append and send shapes exist for destinations that model events rather than mutable records — every row becomes a new event, and the notion of "the same record" does not apply. ## When Mirror is the right answer Membership semantics. A suppression list where a customer who unsubscribed must actually stop being messaged. An ad-platform custom audience where a churned account should stop being retargeted, both because it wastes budget and because continuing to target them can be a compliance problem. A "currently in trial" segment driving in-app messaging. In all of these, the absence of a row is meaningful data, and an upsert-only sync leaks stale membership forever. ## When Mirror is the wrong answer Anything where the destination record has a life beyond your model. Salesforce Contacts carry activity history, notes, open opportunities and relationships. If your model's `where` clause changes — someone adds a filter, an upstream table rebuilds partially, a join fans out or drops rows — Mirror interprets the shrunken model as an instruction to delete, and it will faithfully execute. The damage is not obviously reversible; a deleted CRM record takes its history with it. ## The upstream-failure coupling This is the point interviewers are usually probing. Mirror couples the correctness of your transformation layer to the survival of records in a business-critical system. A dbt model that runs on partial source data and produces ten thousand rows instead of a hundred thousand looks *fine* to the warehouse — it just has fewer rows. To a Mirror sync it looks like ninety thousand deletions. The mitigations are ordinary data-quality discipline applied at the right place: a row-count or volume test on the model that fails the job before the sync ever fires; triggering the sync from the transformation job's success rather than from a clock; and preferring archive-style removal over hard delete when the destination offers both. ## Recovering from a bad Mirror run Honest answer: it depends on the destination. Salesforce has a recycle bin with a retention window; an ad audience can be rebuilt by simply re-running the sync once the model is fixed, because membership is cheap and stateless; a hard-deleted record in a tool with no soft delete is gone. This asymmetry is exactly why the choice of behavior should follow the destination's blast radius, not just the semantics you want. ## Getting the semantics without the risk Often you do not need deletion at all — you need a status field. Rather than mirroring a `active_customers` model and deleting whoever leaves it, sync the full population with Update or Create and map a `status` column that flips to `Churned`. Downstream automation filters on the field. Nothing is destroyed, the history stays, and a broken model produces a wrong flag rather than a missing record. Reach for Mirror when the destination genuinely has no such field — ad audiences being the clearest case. ## How to answer Name the difference in one sentence (Mirror propagates absence, upsert does not), give one destination where absence must propagate, then volunteer the failure mode: a shrinking model becomes a deletion event, so gate the sync on data-quality tests and prefer a status field where one exists.
- What is the risk of pointing a Mirror sync at Salesforce Contacts?Any change that shrinks the model — a new filter, a partially rebuilt upstream table, a join that drops rows — reads as an instruction to delete those contacts, taking their activity history and relationships with them. Gate the sync behind row-count tests on the model, trigger it from the transformation job's success rather than a clock, and prefer archiving over hard delete where the destination supports it.
- When would you sync a status column with an upsert instead of using Mirror?Whenever the destination has a field that can express 'no longer qualifies'. Sync the full population with Update or Create and flip a status field to Churned or Inactive; downstream automation filters on it. Nothing is destroyed, history survives, and a broken model yields a wrong flag rather than a missing record. Mirror earns its risk mainly for ad audiences and suppression lists that have no such field.
- What does the Update Only behavior buy you that an upsert does not?It enriches records that already exist and refuses to create new ones. That matters when the destination's record population is owned by another system or by a sales team, and you only want the warehouse to decorate existing records with scores and attributes. It also keeps a bad model from flooding the CRM with thousands of junk records.
Update or Create is copying files into a folder; Mirror is rsync with --delete, which also removes anything on the far side that is no longer in the source.
saying these in an interview costs you the question
- Assuming every sync behavior is just an upsert with a different name
- Thinking a row leaving the model always removes it downstream
- Using Mirror on CRM records that carry their own history
- Believing deletions from a mirrored sync are always recoverable
- Not connecting a partially built upstream model to mass deletion