Why does Census need a writable schema in your warehouse, and what does that cost?
answer
- it needs somewhere to remember
- diffing requires the previous run's answer
- bookkeeping tables live in your own warehouse
- cadence multiplies warehouse compute
- reset the state and everything resends
basics
~20 sIt stores bookkeeping tables recording what each sync already sent, so it can diff the model each run and push only changed rows. The cost is warehouse compute and storage per sync, and a state reset that resends every row into a rate-limited destination.
solid answer
~50 sDiff-based syncing needs memory of the previous run, and that memory lives in your warehouse rather than the vendor's — Census asks for a dedicated schema it can write bookkeeping tables into. Each run it queries your model, compares it against that recorded state, and emits only the new and changed records. Two consequences follow. First, every sync is warehouse work: a query plus a comparison, billed as your compute, which is why a five-minute schedule on a large model costs real money even when nothing changed. Second, that state is load-bearing. If it is reset — the schema is dropped by a cleanup script, the sync is rebuilt, or a full resync is deliberately requested — the next run sees the entire model as new and pushes every row, which is exactly how a well-behaved integration exhausts a destination's daily API allowance in one morning.
code
text · 4 linesnormal run : model 4,120,000 rows -> diff 830 changed -> 830 API upserts
state reset : model 4,120,000 rows -> diff ALL rows -> 4,120,000 API upserts
destination throttles, sync stalls,
unrelated integrations sharing the org's API quota fail toogo deeper
Know that the sync remembers what it already sent so it can send only what changed, and that this memory is stored in the warehouse rather than in the vendor's systems.
Explain the run loop — query the model, compare against recorded state, emit inserts and updates — and connect sync cadence and mapping width to warehouse compute cost.
Be able to narrate the full-resend incident: how state is lost or reset, why every row then looks new, and how the resulting burst exhausts a shared destination API allowance and breaks unrelated integrations.
Own the cost and blast-radius model — set cadence from how often models actually change, protect the state schema in warehouse governance, and decide organizationally who may trigger a full resync against a shared CRM.
## Why any diffing sync needs state A reverse-ETL sync's job is to make a destination agree with a model. Doing that naively means sending every row every run, which destinations will not tolerate: their APIs are rate-limited, often metered, and frequently slower than the warehouse by orders of magnitude. So the sync must answer "what changed since last time?", and answering it requires remembering last time. That memory is the bookkeeping state. Census keeps it in your warehouse, in a schema you grant it write access to, rather than copying your data out to vendor-side storage. The privacy argument is the appealing one — your customer records stay inside your own account and only the rows actually being activated leave it. The operational argument is the one interviews care about: this makes sync state a table in a database your team administers, with everything that implies. ## What each run actually does Roughly: run the model's query, materialize or scan the result, compare it against the recorded state for this sync, compute the set of inserts and updates (and, under a mirror-style behavior, removals), send those to the destination's API, then record the new state including which records the destination rejected. The comparison is warehouse work. On a wide model with millions of rows it is not free, and it happens on every scheduled run whether or not anything changed. ## The cost profile Think of it in three parts. **Warehouse compute** — a scan and a comparison per sync, so cadence multiplies cost. A sync every fifteen minutes over a hundred-million-row model is a standing query workload you have to budget for, and the standard mitigation is to align the schedule to how often the model can actually change rather than to how fresh someone wishes it were. **Warehouse storage** — modest, but real, and it grows with the number and width of syncs. **Destination API consumption** — governed by how many rows the diff produced, which is normally small and occasionally enormous. A useful design instinct: make the model narrow. Only map the fields you truly want the warehouse to own, and only select those columns. Every extra column is another value that can change and trigger a row into the diff, so wide models produce chattier syncs. ## The failure mode: a full resend The operational story worth telling is the morning where a normally-quiet sync pushes the whole model. Causes cluster into a few shapes. The bookkeeping state was lost — a warehouse cleanup script dropped an unfamiliar schema, an environment was recloned, a sync was recreated from scratch. A full resync was triggered on purpose, sometimes to repair a suspected drift, without anyone checking the destination's remaining quota. Or the model's values genuinely all changed at once, because a transformation was refactored, a rounding or timezone conversion shifted every value, or a new column was mapped into the sync. The effect is the same either way: every row looks new, the destination throttles, the sync stretches for hours or fails partway, and other integrations sharing that org's API allowance start failing too. The blast radius extends past your pipeline, which is what makes it a senior question. ## Operating around it Treat the bookkeeping schema as pipeline infrastructure: exclude it from warehouse cleanup automation, from clone-and-scrub scripts and from cost-driven table pruning, and make sure the people who administer the warehouse know what it is. Understand your destination's rate limits and remaining allowance before triggering a deliberate full resync, and prefer to run one out of business hours. Refactor transformation logic knowing that a change touching every row is also a change pushing every row, and stage such changes rather than shipping several at once. Watch the per-run record count, not just success and failure — a run that sent a hundred times its usual volume is an incident even when it eventually completes. ## How to answer Start with the mechanism (diffing needs prior state, and the state lives in your warehouse), give the privacy rationale briefly, then spend most of the answer on consequences: warehouse compute scales with cadence and model width; the state is load-bearing infrastructure; losing or resetting it produces a full resend that collides with the destination's rate limit and hurts other integrations sharing it. That progression — mechanism, cost, failure mode, mitigation — is what distinguishes someone who has operated the tool from someone who has read the marketing page.
- What should you check before deliberately triggering a full resync?The destination's rate limit and how much of the current window is already consumed by other integrations, then the model's row count against a normal run's diff size. Schedule it outside business hours if the destination is a shared CRM, because a multi-million-row push can exhaust an org-wide API allowance and take unrelated automations down with it.
- Why does mapping extra columns into a sync make it more expensive?Every mapped column is another value whose change can pull a row into the diff. A wide sync produces far more changed rows than a narrow one over the same model, so it consumes more destination API calls and more warehouse comparison work. Map only the fields the warehouse genuinely owns and select only those columns in the model.
- What is the argument for keeping sync state in the customer's warehouse rather than vendor-side?Customer records stay inside the account and only rows actually being activated leave it, which simplifies the privacy and compliance conversation. The tradeoff is that the state becomes infrastructure your team can accidentally destroy — a cleanup script dropping an unfamiliar schema is enough to trigger a full resend.
saying these in an interview costs you the question
- Thinking every sync sends the whole model every run
- Treating the bookkeeping schema as scratch data safe to drop
- Ignoring that a wide mapping produces a chattier sync
- Assuming a higher sync frequency is free because nothing changed
- Requesting a full resync without checking the destination's rate limit