skip to content

Exit & Switching Costs

What makes leaving a provider expensive: proprietary interfaces, stored data, signed commitments, and the price of moving out. Asked because teams claim portability they have never costed.

on this pageshow

questions

18

Why would a team run a component itself on rented machines instead of using the provider's managed version?

level: juniorimportance: must knowfreq 62%

answer

  1. portability is a property of the seam
  2. rent capacity, not the capability
  3. same software under any landlord
  4. a move becomes a restore, not a rewrite
  5. patching, backups and on-call come back

basics

~20 s

Running the component yourself keeps the software identical wherever the machines are rented, so a move becomes a reinstall and a restore rather than a rewrite. You pay for that with the patching, backups and on-call the managed tier was doing.

solid answer

~40 s

A managed tier sells you the *capability*: the provider installs, patches, backs up and fails over the component, and you drive it through the provider's own control surface. Self-running sells you only machines. You install the same component yourself, so its configuration, its protocol, its on-disk format and its operational runbook are the same on anything that rents capacity. That is the portability argument - the seam sits **below** the platform, at the machine boundary, and a move means renting machines elsewhere and restoring the same software. The price is that everything the managed tier was doing silently - version upgrades, tested restores, failover, capacity headroom, security patching - returns to your team permanently, whether or not the move ever happens.

go deeper

for a junior

Recall the two rungs: a managed tier sells you the running capability, while renting machines sells you only capacity. The portability argument is simply that software you installed yourself is the same software anywhere.

for a middle

Explain where the seam sits. Self-running pushes the provider dependency down to machines, network and disks, so a move becomes a reinstall plus a restore instead of a rewrite of the surrounding integration.

for a senior

Show what the seam costs continuously. Patching, version upgrades, tested restores, failover and capacity headroom return to your team every month, whether or not the move that justified them ever happens.

for a principal

Frame it as buying an option and say what the option is worth. Fund the lower seam only where a plausible trigger exists and the team is already staffed to operate the component itself.

## Two ways to consume the same component A **managed tier** sells you a running capability. The provider installs the component, patches it, backs it up, replaces failed nodes and hands you an endpoint plus a **control surface** - a management API and a console - for creating, sizing, configuring and restoring it. What runs underneath may be a well-known component, the provider's own build of one, or an original design that merely looks familiar. You do not operate it, and mostly you cannot see it. **Self-running** the same component means renting only capacity. You take machines - the rung where the provider's responsibility stops at the hardware, the virtualisation layer, the physical network and the storage device - and you install, configure, monitor, patch and fail over the component yourself. The provider sells machines; the capability is yours to build and keep alive. ## Why the lower seam is the portable one Portability is a property of a **seam**: the line below which your system stops needing anything that only this provider has. Self-running drops that line to the machine boundary, and that has a specific, checkable consequence for a move. - The **configuration** is the component's own files or commands, not a provider console's fields, so it copies as-is. - The **protocol and client library** belong to the component, so application code is untouched by the move. - The **on-disk and backup format** is the component's own, so a move is a restore rather than an export, a transform and an import. - The **operational knowledge** - what its metrics mean, how failover behaves, which settings matter under load - travels with the team, because it was never knowledge about a provider. - What stays provider-specific is bounded: machine shapes and sizes, the private address range and its routing, the block storage attached to each machine, and the mechanism by which a machine obtains credentials. Compare the same move made from a managed tier. The component may be familiar, but everything built *around* it - how it is created, how it is sized, how it is backed up, who may call it, what it emits as telemetry, how it is upgraded - is the provider's design, and none of that design exists on the next platform. | What a move touches | Managed tier | Self-run on rented machines | |---|---|---| | Application code | rewritten where the interface is proprietary | unchanged: same component, same protocol | | Configuration | recreated in the new platform's control surface | the component's own files, copied | | Data | exported and imported through the platform's tooling | restored from the component's own backup format | | Operational knowledge | relearned for a new control surface | carried over intact | | Still provider-specific | essentially the whole surrounding design | machine shapes, network, disks, credential delivery | ## What the seam costs, continuously None of this is free, and the bill is not a one-off. Everything the managed tier was quietly doing comes back to you: 1. **Patching and version upgrades.** Security patches are routine; a major version upgrade of a component that owns data is a project, and it is now yours to plan, rehearse and run. 2. **Backups that have actually been restored.** Taking a backup is easy. A *tested* restore, with a known recovery point objective and recovery time objective, is the work. 3. **Failure handling.** Replacing a lost machine, promoting a standby and deciding when to fail over become decisions your on-call makes under pressure. 4. **Capacity.** Headroom, scaling steps and the storage growth curve are yours to watch, and running out is your outage. 5. **Security posture.** Network exposure, authentication settings, encryption configuration and audit output are configured by you rather than defaulted by the provider. That cost is charged every month whether or not you ever move. The honest framing is therefore that self-running **buys an option**, and the option carries a running premium. It is a good buy where the team is already staffed to operate that component, where a managed tier lacks a version, extension or setting you genuinely need, or where enough services depend on it that the alternative rewrite would be large. It is a poor buy where a small team adopts a data-owning component purely because a move is imaginable. ## What self-running does not do It does not remove the provider. The machines, the private network, the disks, the quotas that cap how many machines you may launch and the identity mechanism that hands those machines credentials are all still rented. The seam moved **down**, not away - which is exactly the point, because what remains above it is small, well understood and describable in a day rather than a quarter.

  • Does self-running the component remove every provider-specific dependency?
    No. You still depend on machine shapes, the private address range, block storage, the quotas that cap how much you can launch, and the mechanism that gives a machine its credentials. The seam is lower, not absent. What changes is the size of the move: from rewriting an integration to re-creating a bounded set of infrastructure and restoring data.
  • If the team never actually moves, was the self-run seam wasted?
    Not necessarily, but be honest about the accounting: you bought an option and paid an ongoing operations bill for it. That is a fair trade where the team is already staffed to run the component, or where the managed tier lacks a version or setting you need. Where neither holds, the managed tier is usually the better buy.
  • Which parts of the operating burden do teams most often underestimate?
    Tested restores and major version upgrades. Both are rare, both were being done quietly by the managed tier, and both surface only when they fail. Running a healthy component day to day is usually easy; it is the upgrade path and the restore nobody has rehearsed that consume the engineer-months nobody budgeted.

saying these in an interview costs you the question

  • Claims a managed tier is portable because the component underneath is a well-known one
  • Thinks self-running removes the provider dependency entirely, forgetting machines, network and disks
  • Assumes the portable seam is free once the component has been installed
  • Treats a backup taken by the platform's own tooling as restorable onto another provider
  • Says self-running is always cheaper because you only pay for machines
open as a page

A team says it is locked in to its cloud platform "because of the code" — what does lock-in actually mean, and which kinds are not code?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Lock-in is the cost of leaving, not an inability to leave. Code is one bill of four: data that is priced to move, operating knowledge the team has only here, and an unexpired term commitment are the other three.

open as a page

A migration off one platform is budgeted purely as engineer-months to rewrite code — which cost lines does that estimate miss?

level: middleimportance: must knowfreq 64%

basics

~20 s

A rewrite estimate covers one of four exit lines. The others are the outbound data charge plus the weeks the copy takes, both platforms billed through the overlap window, and the months still running on a term commitment.

open as a page

A managed service is wire-compatible with an open interface - what does that compatibility usually not cover?

level: middleimportance: must knowfreq 56%

basics

~20 s

Compatibility with an open interface normally covers the data path your application speaks, and stops at the edges: provisioning, sizing, authentication, backup and restore, telemetry, quotas and error behaviour all stay the provider's own design.

open as a page

Your board asks for multi-cloud after another company's outage, so which three distinct postures can that one word mean?

level: middleimportance: must knowfreq 62%

basics

~20 s

Multi-cloud covers three different architectures: the same workload live on two platforms at once, different workloads split across platforms with one home each, and a single live platform plus a documented exit plan. Each has a different bill and a different failure behaviour.

open as a page

You inherit a seven-year-old analytics estate on one cloud platform and are asked how locked in the company actually is — how do you answer with an inventory instead of a feeling?

level: seniorimportance: must knowfreq 57%

basics

~20 s

Size the four dimensions separately: proprietary interfaces written into code, data that would have to move, knowledge only this platform has taught the team, and months left on any term commitment. Then rank them, because one row usually dwarfs the rest.

open as a page

A payments ledger must run live on two cloud platforms at once, so what does that constraint take away from its design?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Designing for two live platforms shrinks the usable capability set to their intersection: only what both offer, in a shape both express. The richest managed tiers drop out, more components become yours to run, and the data layer has to settle on one writer or on conflict resolution.

open as a page

Why is stored volume alone a poor measure of data gravity when you price the data dimension of lock-in?

level: middleimportance: should knowfreq 47%

basics

~20 s

Data gravity is stored volume multiplied by what moving one unit costs — the per-gigabyte charge leaving the platform, the elapsed copy time, the effort of extracting it, and reconciling a dataset that keeps changing. Volume prices none of those.

open as a page

An exit plan allots one week to copy a 400 TB media archive off a platform — what makes that estimate wrong?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Sustained throughput sets the copy time: 400 TB at a steady 1 Gbit/s takes about 37 days. Archived objects must first be retrieved to a readable tier, and both the retrieval and the outbound gigabytes are charged.

open as a page

Rewriting the proprietary parts of a system is estimated in engineer-months — why is that number systematically low?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The estimate counts call sites, while the work is rebuilding what the component did behind them — retries, ordering, durability, scaling and the operational tooling around it. Those were the reasons it was adopted, and they are discovered during the port, not before.

open as a page

Your team wraps every provider call behind an in-house abstraction so services could move - what does owning it cost?

level: seniorimportance: should knowfreq 48%

basics

~20 s

An in-house abstraction becomes a product with its own maintainers, release process, bug queue and upgrade debt. Every platform capability and every consumer request now arrives twice: once in the platform, once in the wrapper that hides it.

open as a page

Beyond the doubled infrastructure bill, which recurring costs appear when a workload is kept live on two cloud platforms?

level: seniorimportance: should knowfreq 47%

basics

~20 s

The lasting costs are people and process: two control surfaces, two identity models, two quota regimes, two audits and two incident procedures. Add capacity bought twice so either side can carry the load alone, a split commitment discount, and continuous replication traffic charged on the way out.

open as a page

A term commitment has eighteen unused months left and the team has already chosen to leave — when do you actually cut over?

level: principalimportance: should knowfreq 36%

basics

~20 s

Start from whether the remaining months are owed regardless. If they are, they are not a reason to wait — the real comparison is the cost of carrying a rejected platform for eighteen months against the value of consolidating now, with the stranded amount as a fixed line in both options.

open as a page

You host a dozen services on one platform - where do you place the portable seam, and what do you leave unportable?

level: principalimportance: should knowfreq 40%

basics

~20 s

Place the seam only where a move would otherwise be expensive - usually one or two interfaces over the most proprietary, most widely used dependencies - and leave the rest deliberately unportable, recorded as an accepted cost with a named owner.

open as a page

After another company's outage your board demands a second cloud platform for a regulated payments ledger, so how do you decide whether that posture earns its tax?

level: principalimportance: should knowfreq 44%

basics

~20 s

Start from the failure being feared and check whether a second platform actually removes it, then price the posture's design, capacity, discount and people taxes against the value of surviving that specific failure. Most fears are answered by a cheaper posture than running live on two platforms.

open as a page

An in-house abstraction exposes only the features common to the platforms beneath it - what leaks through anyway?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Behaviour leaks even when the surface does not: latency and throughput, when a write becomes visible, quota and throttling shape, which errors are retryable, what a partial failure looks like, and how a caller's identity is established.

open as a page

A team holds a documented plan to leave its cloud platform but has never exercised it, so how would you test whether it is real?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Rehearse a slice of it end to end: stand one real service up on the second platform from its written description, load a realistic copy of its data, put real traffic on it, and measure the clock. The output is an elapsed time and a defect list, both of which expire.

open as a page

Two designs meet the requirement, one on components you could move anywhere and one on a proprietary managed capability — why might the harder-to-leave option still be right?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Portability is one cost among several, not the goal. Staying portable is paid every month forever; leaving is paid once, if ever. The proprietary option wins when the work it removes, compounded over years, outweighs an exit cost discounted by the chance of exiting.

open as a page