Your custom Airbyte source keeps failing on HTTP 429 — how do you make it back off?
answer
- Throttling is expected traffic on a backfill
- The API usually tells you how long to wait
- Bounded attempts, then a clean failure
- Never return the pages you happened to get
- Fewer, bigger requests beats better retries
basics
~20 sHandle 429 in the connector's retry layer: mark it retryable, derive the wait from the Retry-After header when the API sends one and fall back to exponential backoff, cap the attempts, and checkpoint state often so a final failure loses little work.
solid answer
~50 sIn the Python CDK an HTTP stream exposes retry hooks: decide whether a response is retryable, return how long to wait — reading `Retry-After` or a reset-timestamp header when the API provides one — and cap the attempts. In a declarative manifest the same job belongs to the requester's error handler: `response_filters` match the status and mark it retryable, and a backoff strategy supplies the delay, either exponential or taken from a response header. Two things separate a good answer. First, respect the vendor's own signal instead of guessing a sleep — blind exponential backoff on an API that tells you the reset time either waits too long or hammers it again too early. Second, never swallow the 429 and return partial data: that is silent data loss. Fail the sync after the retry budget, having checkpointed frequently, so the next run resumes near where it stopped.
code
yaml · 12 linesrequester:
type: HttpRequester
url_base: "https://api.example.com/v1"
path: "/events"
error_handler:
type: DefaultErrorHandler
response_filters:
- http_codes: [429]
action: RETRY
backoff_strategies:
- type: WaitTimeFromHeader
header: "Retry-After"go deeper
Know that HTTP 429 means the API is rate limiting you and that a connector should wait and retry rather than crash on the first one.
Explain where retry logic lives in the connector — the retry predicate and backoff hook in Python, or the requester's error handler and backoff strategies in a manifest — and why the retry count is bounded.
Show operational judgment: read the vendor's reset headers, throttle preemptively on remaining quota, fail rather than truncate, and tie checkpoint frequency to how much work a throttled failure discards.
Treat vendor quota as a capacity constraint you plan for: request budget per connection, backfill windows scheduled off-peak, page sizes and bulk endpoints chosen to cut request counts, and an explicit stance on how a connector reports quota exhaustion to its owners.
## 429 is normal traffic, not an exception Any connector that backfills a SaaS API will hit its rate limit; the limit exists precisely to stop what a backfill does. So rate limiting is part of a connector's steady-state behaviour, and "the sync fails on 429" means the connector was written as if the happy path were the only path. The design goal is to keep the sync alive at whatever throughput the vendor permits, and to fail loudly rather than quietly when it cannot. ## The Python CDK hooks A stream built on the CDK's HTTP base class decides retry behaviour in two places: a predicate that says whether a given response should be retried (the default already retries 429 and 5xx), and a method returning the number of seconds to wait before the next attempt, which receives the response so it can read headers. A maximum retry count bounds the loop. There is also a switch controlling whether a non-retryable error status raises or is passed through to your parsing code — useful for APIs that use 404 to mean "this sub-resource is empty", where raising would abort a stream that is actually fine. Newer CDK versions restructure this into an error-handler object with backoff strategies rather than method overrides, so check which shape your CDK version expects before writing overrides that are never called. ## The declarative equivalent A manifest requester takes an error handler. Its `response_filters` match on status codes, or on text in the body, and assign an action: retry, fail, ignore, or treat as rate limited. Its backoff strategies supply the delay: an exponential strategy with a factor, a constant strategy, or a strategy that reads a wait time out of a named response header — which is the one you want when the API sends `Retry-After` or a reset epoch. ## Respect the vendor's signal The difference between a connector that survives a backfill and one that does not is usually whether it reads the API's own headers. Many APIs return `Retry-After` in seconds, or a rate-limit reset expressed as an absolute timestamp, or remaining-quota headers you can use to slow down *before* being blocked. Using them beats guessing: exponential backoff either sleeps far longer than the actual window, turning a two-hour sync into an overnight one, or retries before the window resets and burns another attempt. Preemptive throttling — slowing when the remaining-quota header drops — is the more advanced version and turns a spiky connector into a well-behaved one. Also mind whether the limit is per second, per minute, per day or per credential. A daily quota cannot be waited out inside a sync; the correct response there is to fail cleanly after checkpointing and let the schedule pick it up, not to sleep for hours holding a worker. ## Failing well is part of the answer When the retry budget is exhausted, three things must be true. The failure must be surfaced as an error, not hidden — a connector that catches the 429 and returns the pages it managed is manufacturing silent data loss, and the destination will happily load an incomplete stream. The error should be classified so the platform and the user can tell a transient system problem from a configuration problem such as a plan-level quota that will never clear. And state must have been checkpointed as you went, so the next run resumes near where you stopped instead of replaying the whole backfill into the same wall. ## Reducing the pressure, not just absorbing it Retry logic is the last line. Ahead of it sit choices that determine how often you hit the limit at all: a larger page size means fewer requests for the same rows; narrower incremental windows mean you fetch less; avoiding one API call per parent record where a bulk endpoint exists can cut request counts by orders of magnitude; and concurrency, if the CDK version you use offers it, must be sized under the vendor's ceiling rather than at it. A connector that pages 25 rows at a time across a five-year backfill is rate-limited by its own design, and no backoff strategy repairs that. ## What interviewers listen for A weak answer is "add a sleep and retry". A strong one names the layer the retry belongs in, distinguishes reading the vendor's reset signal from blind exponential backoff, insists on failing rather than truncating, connects checkpoint frequency to the cost of a rate-limit failure, and mentions reducing request volume as the real fix for a chronically throttled backfill.
- Why is blind exponential backoff worse than reading the API's reset header?Exponential backoff guesses. It either sleeps far past the actual reset — turning a short sync into an overnight one — or retries before the window clears and wastes an attempt. When the vendor returns Retry-After or a reset timestamp, waiting exactly that long is both faster and gentler.
- What is wrong with catching the rate-limit error and returning the records already fetched?It converts a visible failure into silent data loss. The destination loads a partial stream, the sync reports success, and the checkpoint may advance past rows that were never emitted. Fail after the retry budget so the run is marked failed and the next one resumes from a checkpoint.
- How does checkpoint frequency change the cost of a rate-limit failure?If state is only emitted at stream end, a throttled failure at hour four discards four hours of requests and replays them into the same limit. Checkpointing every few thousand records, or per completed time slice, bounds the rework so each retry makes forward progress.
- A daily quota is exhausted mid-sync. Should the connector wait?No. Sleeping for hours holds a worker and a connection for no benefit, and the platform's own timeouts will likely kill it anyway. Checkpoint, fail with an error classified as a quota problem, and let the schedule resume it once the window rolls over.
saying these in an interview costs you the question
- Retrying 429 immediately with no delay
- Swallowing the error and returning partial records
- Ignoring Retry-After and always using exponential backoff
- Retrying forever with no attempt cap
- Treating a daily quota as something to sleep through