skip to content

In Hightouch, what is a model and how does a sync push it to a destination?

level: juniorimportance: must knowfreq 62%

answer

  1. It pushes out of the warehouse, not in
  2. A query result plus a schedule
  3. One column has to identify each row
  4. Diff against last run, send the delta
  5. Model plus primary key plus field mapping

basics

~20 s

A Hightouch model is a warehouse query, table or dbt model with a unique primary key. A sync binds that model to one destination object, maps its columns onto destination fields, and writes the rows into a SaaS tool on a schedule.

solid answer

~50 s

In Hightouch a **model** defines the rows you want to activate: a SQL query against your warehouse, a table selector, or a dbt model. Every model must declare a **primary key** column that uniquely identifies a record, because that key is how Hightouch tracks a row across runs and how it matches the row to a record in the destination. A **sync** binds one model to one destination object — CRM contacts, a messaging-platform user profile, an ad-platform customer list — and maps model columns onto destination fields, including the identifier used for matching. Syncs run on a schedule (interval or cron, or triggered when the upstream transformation job finishes), and by default Hightouch compares the current query result against the previous run so only added, changed and removed rows are sent rather than the whole model.

code

sql · 10 lines
sql
-- Hightouch model: one row per customer, keyed by customer_id
select
  c.customer_id,            -- declared as the model primary key
  c.email,
  c.lifecycle_stage,
  s.total_spend_usd,
  s.last_order_at
from analytics.dim_customer c
join analytics.fct_customer_spend s using (customer_id)
where c.email is not null

go deeper

for a junior

Be ready to state the direction — warehouse out to SaaS tools — and name the three pieces: a model, its primary key, and a field mapping to one destination object.

for a middle

Explain why the primary key is required: it is both the identity used to diff against the previous run and the value mapped to the destination's external identifier for matching.

for a senior

Show that you keep the definition in the warehouse and that you know a run sends a diff, not the whole model, so you can reason about API call volume before enabling a sync.

for a principal

Own the governance angle: one set of warehouse models feeding many destinations, with mapping changes reviewed like code, versus every team inventing its own segment definition inside a vendor UI.

## What Hightouch is Hightouch is a reverse-ETL, or "data activation", tool. It reads rows that already exist in your data warehouse and writes them into operational SaaS systems — CRMs, marketing and messaging platforms, ad networks, support desks, spreadsheets. The direction is the thing to fix in your head first: a managed EL connector pulls data *into* the warehouse, while Hightouch pushes modeled data *out of* it. It is warehouse-native, meaning the record definitions stay in your warehouse and Hightouch queries them; it is not a place where you re-model your customer data. ## The model: a query is the unit of activation A model is the definition of the rows you want to send. You can define one as a SQL query against the connected warehouse, by selecting an existing table, or by pointing at a dbt model (Hightouch also integrates with other modeling tools). The result is a rowset with one row per record — one row per customer, per account, per subscription — and whatever columns the destination will need. The practical consequence is that all the joins, filters and business logic live in the warehouse, where they are version-controlled, testable and reusable, rather than being buried in a sync tool's UI. If "active trial account" needs a new definition, you change the model, and every sync built on it changes with it. ## The primary key and why it is mandatory Every model must declare a primary key column, and that column must be unique and **stable** across runs. It does two jobs: 1. **Identity across runs.** Hightouch remembers the previous run's result and compares it to the current one. The primary key is how it decides that row `101` is the same row as last time and merely changed, rather than a brand-new record. 2. **Matching in the destination.** The key (or a column derived from it) is mapped to the destination's external identifier field so an update lands on the existing record instead of creating a duplicate. If the key's value changes for the same logical entity — say you re-key a model from a surrogate `user_id` to `email` — every old row looks removed and every new row looks added, and the sync churns the entire destination. Duplicated key values are just as damaging: two rows claiming the same identity produce inconsistent writes to one destination record. ## The mapping A sync connects one model to one destination object and maps model columns to destination fields. Alongside the field mapping you configure how records are matched — typically an external-ID field on the destination object holding the model's primary key. Columns you do not map are simply not sent; adding a column to the model does not silently start populating a destination field. Destination-side constraints (required fields, picklist values, email format, field types) are enforced by the destination's API, and violations come back as per-row rejections you can inspect in the run. ## Schedules and incrementality Syncs run on an interval or cron schedule, can be triggered when an upstream transformation job finishes, or can be kicked off through an API call. Between runs Hightouch keeps state describing the previous result, so a scheduled run computes a diff and sends only the added, changed and removed rows. This is what makes reverse ETL affordable against rate-limited SaaS APIs: a ten-million-row model with two thousand daily changes should cost roughly two thousand records' worth of API calls, not ten million. A **full resync** deliberately ignores that state and re-sends every row. You reach for it after changing the field mapping, after the destination was wiped or re-created, or when you suspect the state and the destination have drifted apart — and you reach for it knowing it is the expensive path. ## What one run actually does Query the model in the warehouse, diff it against the stored previous result, translate the changed rows into destination API calls in whatever batch shape that API supports, and record per-row outcomes: succeeded, rejected (with the destination's error), or queued for retry. The run summary — rows queried, changed, sent, rejected — is the first thing you read when a sync misbehaves. ## Where people go wrong The frequent beginner errors are: assuming Hightouch ingests *from* SaaS tools; defining a model without a genuinely unique key; including a volatile column such as a recomputed timestamp, which makes every row look changed on every run; and confusing "removed from the model" with "deleted in the source" — a row leaving the query's result set only means it no longer matches the definition.

  • Why must the model's primary key be stable and not just unique?
    Uniqueness stops two rows colliding on one destination record; stability is what lets Hightouch recognise a row across runs. If the key's value changes for the same entity, the old value looks removed and the new one looks added, so the sync churns the destination and, in a removal-capable mode, can delete records it is about to recreate.
  • When would you deliberately run a full resync instead of the normal incremental run?
    After changing the field mapping so previously unsent columns get populated, after the destination object was wiped or rebuilt, or when you suspect the stored state and the destination have drifted apart. Treat it as expensive: it re-sends every row and can exhaust a rate-limited destination API, so schedule it outside peak windows.
  • Where does the business logic for an audience or segment belong?
    In the warehouse model, not in the sync configuration. Keeping the definition as SQL or a dbt model makes it testable, reviewable and reusable across several destinations, and it means the same "active trial" definition drives the CRM, the messaging tool and the ad platform instead of three drifting copies.

The model is a saved search over your warehouse; the sync is a mail-merge that keeps a downstream address book matching that search.

saying these in an interview costs you the question

  • Says Hightouch pulls data from SaaS apps into the warehouse
  • Claims a model works without a declared primary key
  • Assumes every scheduled run re-sends the entire model
  • Thinks Hightouch stores a permanent copy of your customer data
  • Confuses a row leaving the model with a deleted source row

context