A payments ledger must run live on two cloud platforms at once, so what does that constraint take away from its design?
answer
- design to the intersection, not the union
- only what both platforms express
- managed tiers become self-managed work
- the best capability is off the table
- one writer, or a conflict rule
basics
~20 sDesigning for two live platforms shrinks the usable capability set to their intersection: only what both offer, in a shape both express. The richest managed tiers drop out, more components become yours to run, and the data layer has to settle on one writer or on conflict resolution.
solid answer
~50 sThe constraint is the **lowest common denominator**. The workload may only use capability both platforms express, so wherever one platform's managed tier is richer, the design falls back to whatever both can do — usually a self-managed component, which moves the work onto your team. It also slows adoption: a capability only one side has cannot be used until the other has an equivalent, or until you decide the workload is no longer genuinely live on both. The hardest part is the data. A ledger needs one authoritative ordering of writes, and synchronous replication between two platforms pays a wide-area round trip on every commit. In practice you either accept a single writing side — at which point the second side is warm, not active — or you accept asynchronous replication and a conflict-resolution problem the business has to sign off on.
go deeper
Remember the core idea: if a workload must run on two platforms, it can only use what both of them offer. That intersection is smaller than either platform on its own.
Explain where the intersection bites — a managed tier replaced by a component your team now runs, and a capability on one side you cannot adopt until the other has it. Name the operating work that moves onto your team.
Demonstrate that you scope the constraint. Confine the live-on-both posture to the tier whose downtime is intolerable, and be precise about whether writes really go both ways or the design is actually a warm standby with a recovery point.
Own the long-run consequence: an estate frozen at the capability level of the slower platform, and a growing set of self-managed components. Decide which slice of the business genuinely justifies that and get the rest exempted.
## The tax the live-on-both posture levies on design When the same workload has to serve from two platforms simultaneously, every design decision is filtered through one question: *can both sides do this?* Anything only one side can do either gets rebuilt by hand on the other, or gets dropped. The resulting capability set is the **intersection** of two platforms rather than the union, and it is meaningfully smaller than either platform alone. This is the lowest-common-denominator tax, and it is paid every sprint, not once at the start. | Layer | What one platform alone would give you | What survives designing to the intersection | |---|---|---| | **Data engine** | the richest managed engine, with automated failover, backups, and tuning handled for you | a self-managed engine you operate, or two managed engines you must keep behaviourally identical | | **Identity** | platform-native workload identity, handed to the machine with no stored secret | a shared external identity provider plus a mapping on each side, and credential handling you own | | **Eventing** | platform-native eventing with delivery semantics you can build on | whatever ordering and delivery semantics both sides express, or a broker you run yourself | | **Operations** | one management API, one audit trail, one quota regime, one maintenance calendar | two of each, plus whatever you build to see them together | ## Where the intersection actually bites The damage is rarely a missing feature. It is three quieter effects: - **Managed work becomes your work.** Every capability you drop from a managed tier to a self-managed component is a component you now patch, size, back up, monitor and get paged for. That is the real trade the intersection forces, and it lands on the same team that is already operating two platforms. - **Adoption stalls.** A useful capability that appears on one side cannot be adopted until the other side has an equivalent. Over a few years this is the difference between an estate that keeps modernising and one frozen at the shape it had when the posture was chosen. - **The differences do not vanish, they move.** Two platforms are never identical in quota behaviour, throttling of the management API, maintenance windows, or how an outage is signalled. Designing to the intersection hides those differences from the application and relocates them into whatever your team builds and operates. ## The data layer is the hard part For a ledger the intersection tax is largest at the storage layer, because a ledger needs a single authoritative ordering of writes and money is not a value you are allowed to merge. The options are genuinely limited: 1. **Synchronous replication across platforms.** Every commit waits for the other side to acknowledge. The two platforms are typically much further apart than two failure domains inside one platform, so the wide-area round trip is added to every write. For a high-rate ledger that is usually rejected on latency alone. 2. **One writing side, one following side.** Writes go to one platform and replicate asynchronously to the other. This works and is common — but the posture is now warm standby, not live on both, and it inherits a recovery point measured by how far behind the follower is. Call it what it is rather than claiming active-active. 3. **Both sides writing, with conflict resolution.** Requires a partitioning rule that makes conflicts impossible (each side owns disjoint accounts), or an explicit business rule for resolving them. The first is achievable for some ledgers, the second is rarely acceptable for money. Option two is what most "active-active" payments estates actually run, and being precise about that is a senior-level signal in itself. ## What to do about it The useful move is to stop applying the constraint to the whole estate: - Apply the live-on-both posture **only to the tier whose unavailability is genuinely intolerable**, and let everything else use the richest tier its single platform offers. - Write down which capabilities the intersection costs you and what their absence is worth, so the decision is revisited deliberately rather than inherited. - Be honest in naming the posture. If writes only go one way, the design is a standby, and the recovery point and recovery time it implies should be stated and tested like any other recovery target. The intersection is not an argument against ever running on two platforms. It is the reason the posture is worth paying for in a narrow slice of the estate and rarely worth paying for across all of it.
- Why is the intersection tax a recurring cost rather than a one-off design decision?Because platforms keep shipping. Every capability that appears on one side is unusable until the other side has an equivalent, so the estate is permanently constrained to whichever platform moves slower. The cost shows up as slower adoption and as components you keep operating yourself long after a managed alternative existed on one side.
- A team says their ledger is active-active, but all writes go to one platform. What are they actually running?A warm standby. Both sides may serve reads, but a single writing side means the second platform has a lag, and that lag is a recovery point objective whether or not anyone wrote it down. The right response is to measure the lag, state the recovery point and recovery time it implies, and test a promotion.
- Does the intersection tax apply to the split-by-strength posture as well?No, and that is the posture's main design advantage. When each workload has one home, it can use the richest tier that home offers. The split posture still pays the operational tax of two platforms, but it pays no design tax, because no single workload has to be expressible on both.
saying these in an interview costs you the question
- Assuming both platforms offer equivalent managed tiers
- Calling a design active-active when only one side accepts writes
- Ignoring the wide-area round trip added to every synchronous commit
- Treating the dropped managed capability as free once self-managed
- Believing conflict resolution is acceptable for a money ledger
- Applying the intersection constraint to the entire estate at once