After another company's outage your board demands a second cloud platform for a regulated payments ledger, so how do you decide whether that posture earns its tax?
answer
- name the feared failure first
- insurance decision, not technology preference
- only total supplier loss needs this posture
- price design, capacity, discount, people
- cheapest posture covering that failure wins
basics
~20 sStart from the failure being feared and check whether a second platform actually removes it, then price the posture's design, capacity, discount and people taxes against the value of surviving that specific failure. Most fears are answered by a cheaper posture than running live on two platforms.
solid answer
~50 sTreat it as an insurance decision, not a technology preference. First, name the failure precisely: a whole platform lost across regions is rare; a single service degraded, a management API degraded, or a self-inflicted bad change are far more common — and a live second copy does not help with the last one, because it faithfully applies whatever you send it. Second, price the posture honestly: the intersection tax on design, capacity bought twice, a weakened commitment discount, and doubled operating surface for the life of the system. Third, compare the alternatives that address most of the same fear at a fraction of the cost — spreading within one platform, or a rehearsed exit plan, which in a regulated business is often what the supervisor is actually asking for. Then recommend a scoped answer and state what evidence would change it.
go deeper
Recall the framing rather than the answer: a second platform is insurance, and insurance is only worth buying against a failure that would actually happen and would actually be covered.
Be able to separate the failures. A live second platform covers losing a whole supplier; it does not cover a bad change you rolled out yourself, because the second copy applies that change too.
Price the posture completely — design constraint, doubled capacity, weaker commitment discount, doubled operating surface — and compare it against spreading within one platform or holding a rehearsed exit plan.
Own the recommendation: name the insured failure, show the full annual tax, recommend the cheapest posture that covers it, state what it leaves uncovered, and say what evidence would reverse the decision.
## Turn the demand into a question you can answer "Get us multi-cloud" is not a requirement; it is an emotional response to someone else's incident, expressed in architecture. The work is to convert it into a decision with inputs. Three questions do that, and they must be answered in order. ## Step one — which failure is actually feared? The posture is only worth its price if it removes the failure the board has in mind. Sort the candidates: | Feared failure | Does a live second platform remove it? | |---|---| | A whole platform unavailable across its regions | Yes — this is the one case only this posture covers | | One managed service degraded on one platform | Partly, and usually more cheaply answered inside one platform | | The management API degraded while workloads keep serving | Barely — running workloads are typically still serving; what stops is launching, scaling and failing over | | A bad configuration or release rolled out by you | No — the second live copy applies the same change and fails with it | | The supplier raising prices or changing terms at renewal | Partly — but a costed, rehearsed exit usually buys more leverage per unit of spend | | A supervisor asking about concentration risk | Usually no — what is asked for is evidence of a tested exit, not a live second estate | Only the first row is uniquely answered by running live on two platforms. That single observation resolves most of these conversations, and it is the part a lead is expected to bring. ## Step two — price the posture, all of it The honest cost has four components, and only one appears on an invoice: - **Design tax.** The workload is confined to capability both platforms express, so the richest managed tiers drop out and more components become yours to operate. This one is permanent and compounds, because it also slows adoption of anything new. - **Capacity tax.** Either side must carry the whole load alone, so you pay for roughly two estates. Sizing the second for half the load is legitimate but converts the promise into degraded-mode survival, which must be stated. - **Commercial tax.** Spend split across two suppliers earns a weaker commitment discount on each, so the unit cost rises as well as the unit count. - **People tax.** Two management APIs, two identity models, two quota regimes, two audit trails, two incident procedures — and an on-call population that has genuinely operated both. This recurs annually and resets with every new hire. ## Step three — compare the cheaper postures against the same fear Three alternatives address large parts of the same worry: 1. **Spread within one platform.** Independent failure domains inside one supplier cover a large share of real incidents at a fraction of the tax, with no design constraint at all. It does not cover the supplier failing entirely, which is exactly the residual risk to name. 2. **A rehearsed exit plan.** Cheap to hold, and when it carries a measured elapsed time and a dependency inventory, it is genuine evidence rather than an intention. In a regulated setting this is frequently the artefact the supervisor wants. 3. **A scoped live-on-both posture.** Apply the expensive posture only to the narrow tier whose unavailability is intolerable — often the authorisation path rather than the whole ledger — and leave the rest on one platform with the richest tiers it offers. ## Making the recommendation A lead-level answer has a shape: 1. State the failure being insured against, in one sentence, and its plausible frequency and business impact. 2. State the total annual tax of the posture, with the people and commercial components visible, not just infrastructure. 3. Recommend the cheapest posture that covers that failure, and name what it leaves uncovered. 4. State the evidence that would change the recommendation — a supervisory requirement that specifies a live second estate, a materially higher assessed probability of total supplier failure, or a contract term that makes the commercial risk unacceptable. 5. Give the board something concrete to approve now: fund a rehearsal of the exit plan and report its measured duration, then revisit with a real number instead of a headline. ## The trap worth naming out loud If the incident that triggered the demand was actually self-inflicted at the other company — a bad change, a mistaken deletion, an expired certificate — then the posture being demanded would not have prevented it there and would not prevent it here. Saying that plainly, with the alternatives costed, is the judgment the role is paid for. Agreeing to the posture because it is easier than disagreeing is how an organisation acquires a permanent tax that protects against a failure it was never facing.
- The board says the trigger incident took a competitor offline for a day. What do you need to know about it first?Whether it was the supplier's failure or the competitor's own. A bad change, a mistaken deletion or an expired certificate propagates to a second live platform just as readily, so the demanded posture would not have prevented it. If it genuinely was total supplier loss, the posture is on the table and the conversation moves to scope and price.
- If a supervisor requires the firm to address concentration risk, what usually satisfies it?Evidence, not a second live estate: a documented exit plan with a named owner, a dependency inventory, a measured rehearsal and a decision trigger. Requirements vary and some do go further, so read the actual obligation rather than assuming — but the common error is buying the most expensive posture to answer a question about documentation.
- How would you scope the posture if some form of it is genuinely required?Apply it to the narrowest tier whose unavailability is intolerable — often the authorisation path rather than the whole ledger — and keep everything else on one platform with its richest managed tiers. That contains the design and capacity taxes to a small surface while still delivering continuity where the business actually cannot absorb downtime.
- What should the recommendation include so it survives a change of leadership?The failure being insured against, the full annual tax with people and commercial components broken out, what the chosen posture leaves uncovered, and the specific evidence that would reverse the decision. Written that way, a successor can re-open it deliberately instead of rediscovering the argument from scratch.
saying these in an interview costs you the question
- Accepting the demand without naming the failure being insured against
- Assuming a live second platform protects against a bad release
- Pricing only the infrastructure line of the posture
- Answering a documentation requirement with a live second estate
- Applying the posture estate-wide instead of to the intolerable tier
- Claiming a second platform removes a degraded management API problem