skip to content

A nightly Hightouch sync into a CRM hits API rate limits and rejects rows — how do you diagnose it?

level: seniorimportance: must knowfreq 50%

answer

  1. Read the run's per-row errors first
  2. Two different failures wear the same red badge
  3. Ask why so many rows looked changed
  4. One volatile column can dirty every row
  5. The quota belongs to the destination tenant

basics

~20 s

Read the sync run's row-level errors first; they usually name a destination validation failure or a rate-limit response. Then cut volume: find out why the diff marked so many rows changed, filter the model down, and batch or reschedule the writes.

solid answer

~50 s

Start with the run itself: Hightouch reports rows queried, rows changed and per-row rejections with the destination's error message, so you can separate a **validation** problem (bad email, missing required field, invalid picklist value) from a **quota** problem (the CRM refused further calls). Then ask why the run was so large. The usual culprit is a model column that changes every run — a recomputed score, `current_timestamp`, a freshly generated surrogate key — which makes the diff classify every row as changed; the fix is to remove or round that column. Next, reduce the payload honestly: filter the model to records the CRM actually needs, map fewer fields, and split one huge sync into targeted ones. Finally, respect the destination's API budget — schedule outside the window when other integrations are consuming it, and alert on rejection counts rather than discovering them the next morning.

code

text · 9 lines
text
Sync run summary
  rows queried   4,812,003
  rows changed   4,811,940      <-- almost everything "changed"
  rows rejected  2,140,882

Top rejection reasons
  REQUEST_LIMIT_EXCEEDED   api call limit reached for this org
  INVALID_EMAIL_ADDRESS    "n/a" is not a valid email address
  REQUIRED_FIELD_MISSING   LastName

go deeper

for a junior

Know that a sync run reports rows sent and rows rejected, and that the destination's error message on a rejected row is where troubleshooting starts.

for a middle

Explain how diffing keeps API volume proportional to change, and how a column that changes every run destroys that property by making every row look changed.

for a senior

Demonstrate the full diagnosis: separate quota errors from validation errors, trace the oversized diff to its cause, shrink the model, and stagger runs against a shared destination quota.

for a principal

Own the capacity conversation — who else spends this destination tenant's API budget, what the steady-state change volume will be as the model grows, and what the alerting contract is when rows are silently rejected.

## Why reverse ETL is rate-limit bound A warehouse can produce ten million rows in seconds; a SaaS API will accept some bounded number of records per unit time, often shared across every integration your company has pointed at that tenant. Reverse ETL is therefore never a throughput problem on the read side and almost always one on the write side. Every design decision — diffing, batching, field mapping, scheduling — exists to keep the number of destination API calls proportional to *what actually changed*, not to how big your model is. ## Step 1: read the run, not the dashboard A sync run gives you three numbers that answer most questions: rows queried, rows the diff considered changed, and rows the destination rejected. Alongside that, Hightouch surfaces per-row errors carrying the destination's own message. Sort those messages by frequency, because they split into two very different diagnoses. **Validation rejections** name a field: an invalid email format, a missing required field, a picklist value the CRM does not recognise, a string longer than the field allows. These are a data-quality problem in the model, and the fix belongs upstream — filter the offending rows out, or clean them in the warehouse so the destination never sees them. **Quota rejections** name the org or the API, not a field. They mean the CRM stopped accepting calls. No amount of retrying fixes that; you have to send less or send it at a different time. ## Step 2: find out why the run was so large If a model of ten million rows reports nine million changed on a nightly run, the diff is not lying — something in the row genuinely differs from last night. The most common cause is a **volatile column**: `current_timestamp as synced_at`, a score recomputed with a random tiebreak, a `days_since_signup` integer that increments daily, a timestamp with sub-second precision. Every such column makes the row's value differ, so every row is "changed" and every row becomes an API call. The fix is to make the model's output stable between runs: drop columns the destination does not use, round timestamps to the day if the destination only needs a date, and compute derived values at query time in the destination rather than shipping a daily-changing number. A second cause is an unstable primary key — if the key is derived from something that changes, every row looks removed and re-added. ## Step 3: reduce the payload honestly Three levers, in order of leverage: 1. **Filter the model.** Sales does not need every user who ever visited; sync the records the CRM actually acts on. A `where` clause is the cheapest optimisation in the pipeline. 2. **Map fewer fields.** Some destination APIs charge a call per record regardless of width, but narrower payloads reduce validation-rejection surface and make change detection less twitchy. 3. **Split the sync.** One sync for the high-churn attributes on a frequent schedule, another for slow-moving attributes daily. That way the expensive object is only touched when the fast-moving fields actually move. ## Step 4: respect the destination's budget The API quota belongs to the destination tenant, not to your sync. Other integrations, an ETL tool pulling data back out, and human users in the UI all draw from it. Practical measures: move the sync off the hour when everything else runs; stagger multiple syncs against the same tenant rather than firing them together; use the destination's bulk or batch endpoint where the connector supports it, since a batched call costs one unit for many records; and avoid running a full resync during business hours — a full resync ignores the stored diff state and re-sends every row, which is exactly the spike you are trying to avoid. ## Step 5: make rejections visible Rejected rows are silent data loss with a friendly UI. Treat them like a dead-letter queue: alert when the rejection rate for a run crosses a threshold, route the rejected records and their error messages somewhere reviewable, and fix the top error class rather than retrying blindly. A sync that has quietly rejected two percent of rows for a month has left a CRM full of stale records that everyone believes is fresh. ## Prevention Before enabling any new sync, estimate its steady-state daily change volume from the warehouse — count how many rows differ day over day — and compare it with the destination's documented limits. Then watch the first week of runs. A sync whose changed-row count is a stable small fraction of the model is healthy; one whose changed-row count equals the model size will fail the moment the model grows.

  • The model is unchanged but one run suddenly sent every row. What would you check?
    Whether someone triggered a full resync, whether the field mapping changed (which invalidates the comparison for the added fields), and whether the primary key's values shifted. All three make the previous run's state stop matching, so every row is classified as new or changed even though the business data did not move.
  • How do you keep rejected rows from becoming silent data loss?
    Treat them as a dead-letter stream: alert when a run's rejection rate crosses a threshold, land the rejected records with their destination error messages somewhere reviewable, and fix the largest error class at source. Blind retries just re-consume quota, because validation rejections will fail identically every time.
  • When is a full resync still the right call despite the cost?
    After changing the field mapping so newly mapped fields get populated on existing records, after the destination object was rebuilt or bulk-edited outside the sync, or when you have concrete evidence the stored state and the destination disagree. Schedule it in a low-traffic window and expect it to consume the destination's quota.

saying these in an interview costs you the question

  • Retries rejected rows without reading the error message
  • Blames the warehouse for a destination rate limit
  • Assumes rows rejected by the destination are retried forever
  • Adds a recomputed timestamp column and expects diffing to stay small
  • Runs a full resync as the first troubleshooting step

context