Would you adopt Hightouch or build warehouse-to-SaaS syncing in house, and why?
answer
- The SQL is the easy part
- Count the destinations, not the lines of code
- Ask who is allowed to change a mapping
- The models stay yours either way
- Hybrid is a legitimate answer
basics
~20 sBuy when destinations are many and their APIs churn, non-engineers must own mappings, and per-row error reporting matters. Build when one or two stable destinations, strict data-residency rules or a punishing cost curve justify permanently owning diffing, retries, rate limiting and observability.
solid answer
~50 sThe naive comparison is wrong because the SQL and the API call are the easy part — a competent engineer writes that in a week. What a tool like Hightouch actually sells is the long tail: run-to-run diff state so you send only changes, per-record error surfacing instead of a stack trace, destination-specific quirks (bulk endpoints, quota behaviour, picklist validation), backfills and full resyncs, retries that respect rate limits, a mapping UI non-engineers can safely change, and alerting. I would buy when the destination count is growing or churning, when marketing and sales own the field mappings, and when the team has no appetite for on-call on a marketing integration. I would build when there is one high-volume, stable destination whose cost under a volume-priced vendor dwarfs the engineering, or when policy forbids a third party querying the warehouse. A hybrid is common and defensible: buy for the long tail, build for the one destination that dominates volume.
go deeper
Understand that the query and the API call are only a fraction of the work; change detection, retries and error reporting are the rest.
Be able to list what a managed tool provides beyond the happy path — diff state, per-row errors, backfills, destination quirks, a mapping UI — and why each takes real engineering.
Argue from the operational reality: who is on call, how rejected rows become visible, and how destination API quota is shared and protected under either option.
Own the decision framework and its thresholds — destination churn, mapping ownership, volume against the pricing shape, residency, latency — plus the hybrid option and an explicit trigger to revisit.
## Frame the question properly Almost every candidate answers this by comparing "a SQL query plus an HTTP POST" to a vendor invoice, and that comparison always favours building. It is the wrong comparison. Write down what a production reverse-ETL path has to do, and the picture changes. ## What you are actually buying **Change detection with durable state.** Sending only what changed requires remembering the previous result for every model, keyed by a stable primary key, and surviving restarts, schema changes and partial failures. This is the piece teams consistently underestimate — and without it, every run is a full send against a rate-limited API. **Per-record error handling.** SaaS APIs reject individual records for individual reasons. You need per-row outcomes, the destination's message attached to the offending record, and a way for a human to see and fix them. A log line saying `HTTP 400 x 12,431` is not that. **Destination idiosyncrasy.** Every destination has its own matching semantics, bulk endpoint, quota model, required fields, picklist validation and pagination. Multiply that by the number of tools your go-to-market org uses, and by the rate those vendors change their APIs. This is the cost that never ends: not writing the first connector, but keeping fifteen of them working. **Operational surface.** Scheduling, retries that back off instead of amplifying a rate-limit problem, backfills and full resyncs, alerting on rejection rate, an audit trail of who changed a mapping, and access control so a marketer can edit a mapping without touching production code. **A safe interface for non-engineers.** This is often the real driver. If every new field on a CRM object requires a pull request, the data team becomes a ticket queue and the business routes around it with CSV exports. ## What building genuinely costs The first destination takes days. The second takes days. Then the requests arrive: could you also backfill, could you show which rows failed, could marketing change the mapping themselves, why did the CRM run out of API calls at 9am, can we sync every fifteen minutes instead of nightly. Each is reasonable and each is a feature you now own forever, on a system whose failure mode is quiet: the CRM shows stale values and everyone trusts them. Be honest about ongoing ownership too. Marketing integrations are rarely anyone's favourite on-call, and they break on someone else's release schedule. ## The decision inputs - **Number and churn of destinations.** One stable destination favours building; a growing catalogue favours buying, and the curve is steep. - **Who owns the mapping.** If the answer is not "engineers", the tool's UI and permissions model are the product you are paying for. - **Volume and the cost curve.** Vendor pricing in this category scales with synced volume and destination count; per-row costs that are trivial at a hundred thousand records can be significant at hundreds of millions. Model your own volume against the vendor's shape — do not assume either direction. - **Security and residency.** A hard requirement that no third party connects to the warehouse settles the question. So does the reverse: if the vendor is already an approved processor, the friction is low. - **Latency.** Near-real-time activation raises the engineering bar sharply — you need triggering, incremental change detection and back-pressure — which usually strengthens the buy case rather than weakening it. - **Team appetite.** A platform team with capacity and a strong internal integration framework can build; a four-person data team with a quarterly roadmap cannot absorb it without dropping something else. ## Hybrid is a real answer The strongest answer is often "both". Buy for the long tail of destinations where breadth and self-service matter, and build the single high-volume path — the ad platform receiving a hundred million rows, or the internal service that consumes the same data — where the volume-priced cost is worst and the API is stable enough to own. Do not let this become fifteen bespoke pipelines by accident; make it a deliberate, documented exception with a stated volume threshold. ## Lock-in is lower than people fear One argument in favour of buying that candidates often miss: the switching cost is small because the intellectual property stays in your warehouse. The record definitions are models — SQL or dbt — that you own. Moving vendors means recreating mappings and schedules, not migrating data or rewriting business logic. That asymmetry makes adopting a tool a much smaller commitment than adopting, say, a warehouse or a transformation framework, and it is the right thing to say when someone raises lock-in as a blocker. ## How to present it Do not answer with a preference. Answer with the inputs you would gather — destination count and churn, who owns mappings, volume against the pricing shape, security constraints, latency need, team capacity — say which way each pushes, and name the threshold at which you would revisit the decision.
- What is genuinely the hardest part to reproduce if you build it yourself?Durable run-to-run change detection plus per-record error handling at scale. Remembering the last result for every model, surviving restarts and schema changes, and attaching each destination rejection to the record that caused it is far more work than the extraction query or the API call, and it is what keeps API volume proportional to real change.
- How would you argue against a lock-in objection to adopting a reverse-ETL vendor?The business logic lives in your warehouse as models you own, so switching means recreating mappings and schedules rather than migrating data or rewriting definitions. That makes it one of the cheapest categories to reverse, unlike the warehouse or transformation layer, where the logic itself would move.
- When does a hybrid — buy for most destinations, build for one — stop being sensible?When the exceptions multiply. One documented high-volume path with a stated threshold is a deliberate tradeoff; five bespoke pipelines built ad hoc is the in-house platform you decided not to build, arrived at by accident and with no owner. Set the threshold explicitly and review it.
saying these in an interview costs you the question
- Compares only the SQL and the API call to the invoice
- Ignores who will own the mapping changes
- Assumes a vendor's per-row pricing shape without modelling volume
- Treats reverse ETL as a build-once project with no ongoing cost
- Claims lock-in is severe when the models stay in the warehouse