When does Fivetran stop being the right choice for a growing ingestion footprint?
answer
- it is a portfolio decision, not a doctrine
- a few tables usually dominate the bill
- catalog value falls as custom sources rise
- some constraints are not negotiable with a vendor
basics
~20 sWhen the consumption bill outgrows the engineering cost it replaces, when most sources are custom rather than catalog connectors, or when residency, network or latency constraints the hosted service cannot meet start driving the architecture.
solid answer
~50 sManaged EL is a trade of engineering time for a consumption bill, and the trade goes bad along three axes. **Volume**: a per-changed-row model prices well for many small SaaS sources and badly for a few enormous high-churn tables — at that point a self-hosted capture for those specific tables is often dramatically cheaper while everything else stays managed. **Coverage**: if a growing share of your sources are internal or exotic APIs with no catalog connector, you are paying for a catalog you barely use while still writing connector code. **Constraints**: data residency, private networking to sources the vendor cannot reach, sub-minute freshness, or a compliance posture that forbids data transiting a third party. The answer is rarely all-or-nothing — the mature position is a portfolio: managed for the long tail of commodity sources, self-hosted or hand-built for the handful that dominate cost or carry constraints. Decide with measured per-source cost, not with a philosophical preference.
go deeper
You are not expected to make this call, but know the basic trade: a managed service replaces connector maintenance work with a usage-based bill, and neither side is free.
Be able to name what insourcing re-acquires — API churn, state, backfill, retries, schema handling, on-call — and why a per-changed-row bill scales badly on a few very high-churn tables.
Show you would measure before arguing: per-source consumption and incident data, a narrow pilot on the costliest source, and a hybrid outcome rather than a wholesale migration.
Own the portfolio and the exit. Segment sources by cost and constraint, decide deliberately which stay managed, keep transformation logic vendor-neutral, and re-run the analysis as volumes and pricing move.
## Frame the question correctly "Build vs buy" invites a religious answer. The useful frame is a **portfolio decision per source**, revisited as volumes change. Most mature data platforms end up mixed: a managed service handling dozens of low-volume SaaS sources nobody wants to own, and a small number of high-volume or constrained pipelines handled differently. ## What you are buying, precisely A managed connector removes: source API churn, incremental state management, backfill logic, retry/resume, destination DDL for schema changes, and the on-call burden for all of it. That is real, recurring, unglamorous work with no competitive value. For a five-person data team with thirty SaaS sources, buying it is close to unarguable — thirty hand-written extractors is a full-time job that produces nothing a customer would pay for. ## Axis 1: volume and the cost curve Consumption priced per changed row scales with source churn. That curve is friendly for a marketing platform with tens of thousands of changed rows a month and unfriendly for a production database table with hundreds of millions. The asymmetry matters: a handful of tables typically dominate the bill. The correct response is surgical, not wholesale. Measure per-connector and per-table consumption. If three tables account for most of the spend, evaluate moving *those three* to self-hosted log-based capture and leave the other forty alone. Wholesale migration off a managed service to save money on a minority of the footprint usually costs more in engineering time than it saves, and it re-acquires all the maintenance you were paying to avoid. Also price the alternative honestly. Self-hosted ingestion is not free: compute, storage, on-call, upgrade toil, and the engineer-months to reach parity on retry, backfill and schema handling. Compare loaded cost to loaded cost, including the failure modes you will now own at 3am. ## Axis 2: coverage and the custom-source share The value of a connector catalog is proportional to how much of it you use. A team whose sources are all standard SaaS gets enormous leverage. A team whose sources are increasingly internal microservice APIs, partner feeds in bespoke formats, or legacy systems gets less: you are writing connector code anyway, now inside someone else's framework and lifecycle. Track the ratio of catalog-connector sources to custom ones over time — a rising custom share is the leading indicator that the economics are shifting. ## Axis 3: constraints the service cannot satisfy Some requirements are not negotiable with a vendor: - **Residency and sovereignty.** Data that must not leave a jurisdiction, or must not transit a third-party processor at all. - **Network reachability.** Sources inside a private network with no acceptable path for a hosted service, or a security posture that forbids one. - **Latency.** Managed EL is scheduled batch at heart. If a consumer genuinely needs seconds, a streaming capture pipeline is the right architecture regardless of cost. - **Semantics.** If you need exactly the change stream — every intermediate state, ordered — a loader that syncs current state on a schedule is the wrong tool no matter how well it is operated. When a constraint like this appears, the decision is architectural rather than financial, and cost arguments are a distraction. ## Axis 4: lock-in and exit cost The practical lock-in of a managed EL service is moderate but real: destination table shapes, bookkeeping columns, and every downstream model written against them. Reduce it deliberately — keep transformation logic in your own repository and warehouse, avoid embedding vendor-specific columns deep in business logic, and know what a migration would actually cost before you need to make the argument. A vendor relationship you cannot price your way out of is a governance problem, not just a cost one. ## How to actually run the decision 1. **Instrument first.** Per-source consumption, per-source freshness requirement, per-source incident count. Most teams argue this from anecdote because they never measured it. 2. **Segment the portfolio.** Sort sources by cost and by constraint. Look for the small set that dominates. 3. **Cost the alternative properly.** Engineer-months to parity, then ongoing operations. Include the on-call you re-acquire. 4. **Pilot narrowly.** Move one high-cost source and measure real total cost for a quarter before extrapolating. 5. **Re-evaluate on a schedule.** Volumes grow, pricing changes, connectors mature. A decision made at one scale is not a decision for all time. ## The answer that impresses It is not "buy, because engineering time is expensive" or "build, because the bill is outrageous". It is a candidate who says: here is how I would measure it, here is the shape of the cost curve, here is the small set of sources where the trade flips, and here is what I would deliberately keep managed even while insourcing the expensive tail — plus an honest account of what insourcing re-acquires. ## Version note Pricing models, plan structures and self-hosted alternatives all move; treat any specific comparison as time-bounded and re-derive it from measured usage rather than from a remembered rate card.
- How do you avoid the migration costing more than it saves?Move only the sources that dominate cost, and pilot one before extrapolating. Measure per-source consumption first — teams usually find a small minority of tables driving most of the bill. Keep the long tail managed: re-acquiring maintenance on forty low-volume connectors to save little is a net loss in engineering time.
- What do you re-acquire the day you insource a connector?Source API churn, incremental state and its edge cases, backfill, retry and resume, schema-change handling, destination DDL, monitoring and on-call. Plus infrastructure to run and upgrade. The engineer-months to reach parity are the visible cost; the recurring interrupt burden is the one teams systematically underestimate.
- Which requirements make the decision architectural rather than financial?Data residency or a prohibition on third-party processors, sources unreachable from a hosted service, sub-minute freshness needs, or a genuine requirement for the ordered change stream rather than periodic current state. When one of these applies, cost comparison is beside the point — the managed batch loader is simply the wrong shape.
- How do you limit lock-in while staying on a managed service?Keep transformation logic in your own repository running on your own warehouse, avoid threading vendor bookkeeping columns through business logic, and maintain a written, costed view of what migrating one source would take. The goal is being able to make a credible move, which is also what makes the commercial conversation a real negotiation.
saying these in an interview costs you the question
- Treats it as all-or-nothing rather than per source
- Compares vendor bill against zero engineering cost
- Ignores residency, network and latency constraints entirely
- Assumes self-hosting connectors is basically free once deployed
- Argues from preference without measuring per-source consumption