You are drafting an availability commitment for a reporting service built on three committed platform services in series — what can you honestly promise?
answer
- series multiplies, so the chain is worse
- worse than the weakest link
- tolerated downtime adds up
- redundancy multiplies the unavailabilities
- independence is the hidden assumption
basics
~20 sLess than the weakest link. Serial dependencies multiply, so a request path through three committed services promises the product of their figures — and the only ways to raise it are removing a link from the path, masking one with redundancy, or degrading without it.
solid answer
~50 sIf every request must traverse all three, the path's committed availability is the **product** of the three figures, which lands below the weakest of them because every factor is less than one. Stated as tolerated downtime, the allowances add: three links each permitted `T` minutes a month can, in the worst case, cost you `3T`. That is the ceiling your own commitment sits under, and the levers are structural rather than contractual — take a dependency off the critical path, serve a degraded answer when it is missing, cache or queue so a brief absence is invisible, or run a redundant copy so that a link fails only when both copies do. Two caveats belong in the same breath: the multiplication assumes the links fail independently, which correlated platform failures violate, and these are contractual figures rather than measured probabilities.
code
pseudocode · 18 linesserial(availabilities):
product = 1.0
for each a in availabilities:
product = product * a
return product // below every factor, since each a < 1
redundant(a, copies):
allFail = 1.0
for each copy in 1..copies:
allFail = allFail * (1 - a)
return 1 - allFail // above one copy, only if copies fail independently
entry = committedAvailabilityOf(managedEntryPoint)
compute = committedAvailabilityOf(computeTier)
store = committedAvailabilityOf(managedStore)
pathAsBuilt = serial([entry, compute, store])
pathWithPair = serial([entry, compute, redundant(store, 2)])go deeper
Recall the direction: chaining dependencies makes availability worse, not better, and a service is never as available as the best thing it depends on.
Do the arithmetic and say why: serial links multiply, so the product sits below every factor, and the tolerated downtime of each link accumulates across the window.
Show the levers in a real system — removing a dependency from the request path, defining what a degraded answer looks like, and making the weakest link redundant — and name what each one costs.
Own the published figure. Bring the computed ceiling with its assumptions, a proposed number with deliberate margin, and a clear statement of which exclusions you mirror upstream and which exposure the business is choosing to absorb.
## Why a chain promises less than any link in it Composite availability is the arithmetic every engineer meets the first time they have to put a number in front of a customer. If a request must pass through several components in series, the request succeeds only when **all** of them are working. With per-link availability `a1`, `a2`, `a3`, the path's availability is `a1 x a2 x a3`. Because each factor is below one, the product is below every factor. The chain is worse than its weakest link, not equal to it — and adding a fourth dependency makes it worse again, no matter how good that dependency is. This is the single most useful fact in this material, and it is counter-intuitive enough that the wrong answer ("the weakest link sets the number") is extremely common in interviews. The same statement in minutes is often more persuasive to non-engineers. If each link's commitment tolerates `T` minutes of downtime per measurement window, and the links fail at different times, the path can be down for up to `3T` in that window. The failures do not cancel; they accumulate. ## The levers, in the order they are usually worth pulling You cannot negotiate the arithmetic, so you change the topology: 1. **Take the dependency off the request path.** A service consulted synchronously on every request is a multiplicative term. The same service read from a periodically refreshed copy, or written to asynchronously through a queue, is not. 2. **Degrade instead of failing.** If the reporting service can answer from cached aggregates when a live store is unreachable, that store stops being a hard term in the product for the queries that can degrade. Decide per feature which answers may be stale and say so in the contract. 3. **Make a link redundant.** Two independent copies of a link both have to fail for the pair to fail: the pair's unavailability is `(1 - a)` multiplied by itself, so a weak link improves sharply. This is where most of the engineering money goes, and it only works to the degree the copies really are independent. 4. **Promise less, with margin.** Perfectly legitimate and often the right answer. Publish a figure comfortably under the computed ceiling and keep the difference as headroom for the things the model does not see. ## What the multiplication does not tell you Being honest about the limits of the model is what makes this a senior-plus conversation rather than a formula: - **Independence is an assumption, and often false.** Two links can share a failure domain, a shared platform dependency, or the same maintenance event, in which case they go down together and the product overstates what you have. Correlated platform failure is its own subject; here it is enough to say the arithmetic is an upper bound on what redundancy buys you. - **These are contractual figures, not measured probabilities.** A committed level says what the provider accepts being measured against, not what the service statistically does. Multiplying commitments yields a contractual ceiling for planning, not a prediction. - **The links have different carve-outs.** Each supplier excludes different things from its own measurement, so the felt availability of the chain is worse than the product of the measured figures. - **Remedies do not compose.** Your customer's credit is computed against what they pay you; the credits you can collect upstream are computed against what you pay each supplier. They are unrelated amounts, and the shortfall is yours. | Structure | Availability of the group | When it applies | |---|---|---| | Links in series | product of the links | every request needs all of them | | Independent copies in parallel | one minus the product of each copy's unavailability | a failure of one is masked | | Optional dependency with degradation | the path without it, for degradable requests | a stale or partial answer is acceptable | ## What to actually put in the contract Bring three things to the drafting conversation. First, the computed ceiling, with the chain drawn and the assumptions written down — which dependencies are on the request path, and which of them a request can survive without. Second, the figure you propose, below that ceiling, with the margin stated as a deliberate choice rather than a rounding. Third, the remedy and the exclusions, mirroring upstream where you must absorb something you cannot control. And resist the reflex to quote the platform's headline figure onward. Your service is not one of your suppliers' services; it is the composition of several of them plus your own code, and the composition is always the weaker object.
- Your chain's computed ceiling is below the figure sales has already quoted. What do you do?Say so with the chain drawn and the assumptions written down, then offer the options rather than a refusal: remove a dependency from the request path, add a redundant copy of the weakest link, define a degraded mode that keeps the promise for most requests, or lower the published figure. Each has a cost; the quote is a commercial decision made with them visible.
- When does adding a redundant copy of a link fail to improve the path?When the copies are not independent — sharing a failure domain, a common platform dependency, or the same maintenance event — and when the failover itself is a dependency that can fail. The arithmetic assumes both copies go down only by coincidence; anything that makes them fail together collapses the benefit while you keep paying for it.
- Can you offset your own shortfall with the credits you collect from your suppliers?Not meaningfully. Each supplier's credit is a share of what you pay that supplier, while your obligation is a share of what your customers pay you, and the two have no fixed relationship. Treat upstream credits as incidental recovery, never as a hedge against the commitment you publish.
- Does a chain of two links with different committed figures round to the weaker one?No — it is strictly below the weaker one, because the stronger link still contributes a factor less than one. The approximation is close when one link is far weaker than the others, which is why people get away with it, but the habit breaks as soon as the path grows several comparable dependencies.
saying these in an interview costs you the question
- Says the weakest link sets the path's availability
- Quotes a supplier's headline figure onward to customers
- Adds availabilities instead of multiplying them
- Assumes redundant copies always fail independently
- Treats contractual commitments as measured probabilities
- Plans to fund customer credits from supplier credits