A checkout API's monthly cloud spend rose sharply while orders also rose — which number tells you something got worse?
answer
- growth confounds the total
- divide by work actually done
- a denominator the business already counts
- retries inflate a request denominator
- fixed cost per unit falls by itself
basics
~20 sUnit cost: spend attributed to the service divided by a unit of work the business already counts, such as completed orders. A total confounds growth with regression, while a per-unit figure stays comparable as volume changes and survives the growth argument.
solid answer
~40 sA total bill answers "did we spend more", which nobody disputes, and never answers "did anything get worse", because growth and regression both push it up. Divide instead by a unit of work the business already counts and trusts — completed orders for a checkout API — and you get a figure that is comparable across periods of different size. Two refinements make it honest. **Split fixed from variable**: the fixed platform share divided by a growing denominator falls on its own, so an all-in figure flatters you for growing. And **choose the denominator carefully**: a technical denominator like total requests can be inflated by a client retry storm, which makes cost per request improve during an incident. Read it as a trend, and expect step changes where rates or commitments change.
code
pseudocode · 14 lineswindow = lastCompletePeriod
charges = chargesFor(scope = "checkout", period = window)
variableSpend = sum(c.amount for c in charges if c.scalesWithVolume)
fixedSpend = sum(c.amount for c in charges if not c.scalesWithVolume)
# a denominator the business already counts and trusts
orders = businessCounter("orders_completed", window)
unitCostAllIn = (variableSpend + fixedSpend) / orders
unitCostVariable = variableSpend / orders
# unitCostAllIn falls as orders grow even if nothing improved;
# unitCostVariable is the one that tests whether the design amortisesgo deeper
Recall that a bigger bill after more work is not automatically worse, and that the comparison you want is spend divided by the work done rather than spend on its own.
Explain how to construct the ratio: a numerator scoped to the service and a denominator the business already counts. Work through a case where spend rose and unit cost fell.
Show the refinements that make it honest — splitting fixed from variable so growth alone cannot flatter the figure, and picking a denominator a retry storm cannot inflate. Be able to say what a rising variable unit cost implies about the design.
The call you own is which unit the organisation is measured on and who is accountable for it, knowing that the number only works if spend can be attributed to the same thing the denominator counts — and that a good ratio on unnecessary work is still waste.
## Why the total cannot answer the question A monthly total moves for two completely different reasons: you did more work, or the same work became more expensive. The total is the sum of those and cannot separate them, which is why "spend is up a third" is a fact with no decision attached to it. If orders rose by half over the same period, then each order got **cheaper** — spend at 1.33x against work at 1.5x is about 0.89x per order — and the correct response to the rise is to do nothing. So the number that survives growth is a ratio: > **unit cost = spend attributable to the thing ÷ units of work the thing did** Everything interesting is in how you choose each half. ## Choosing the denominator | Denominator | What it is good for | How it misleads | |---|---|---| | Completed orders | A checkout path, where the business already counts and trusts this number | Undefined for services with no obvious business event | | Requests served | Technical services with no business event | A client retry storm inflates it, so unit cost *improves* during an incident | | Active tenants or accounts | Multi-tenant platforms, capacity planning per customer | Hides wide variation between a large tenant and a small one | | Gigabytes processed | Data pipelines, where volume is the work | Says nothing about whether processing that volume was worth doing | The rule of thumb is to pick something the business already counts, because a denominator invented by the platform team is one nobody else will accept in a review — and the review is where this number earns its keep. The retry-storm case is the sharp edge worth remembering: a denominator that moves for reasons unrelated to value delivered will happily report an improvement while the service is failing. ## Splitting fixed from variable Some of a service's spend scales with the work — per-request charges, runtime that scales out, bytes moved per order. Some does not: a baseline of always-on capacity, a managed component sized for peak, a monitoring footprint. Dividing the **all-in** figure by a growing denominator produces an improvement automatically, with no engineering behind it, purely because the fixed share is spread over more units. That means you want two figures: - **All-in unit cost** — what an order actually costs the business. This is the number for a pricing or margin conversation. - **Variable unit cost** — what the *next* order costs. This is the number that tests whether the design amortises, and the one that should be flat or falling. If the variable figure rises while volume rises, the design is not amortising: something in the path scales worse than linearly, or a fixed component is being duplicated per unit of work. That is a genuine finding, and it is invisible in both the total and the all-in figure. ## Reading the trend A few things to expect, so they are not mistaken for regressions: 1. **Step changes at rate boundaries.** Where volume tiers or a commitment change what you pay for the same usage, the unit figure steps rather than drifts. Annotate the step; do not investigate it twice. 2. **Seasonal swing on the all-in figure.** Fixed cost divided by a quiet month's volume is higher per unit. That is arithmetic, not regression. 3. **Improvement that is only dilution.** Volume grew, the fixed share spread further, nothing changed. The variable figure will say so. And a trap in the other direction: an excellent unit cost on a service the business does not need is still spend worth removing. The ratio tells you about efficiency per unit of work; it cannot tell you the work was worth doing. ## Where it fits among the other signals This is the signal that answers the question the others cannot: - A **budget threshold** tells you the cumulative total crossed a line, and says nothing about whether the work grew. - An **anomaly rule** tells you spend moved against the workload's own recent shape, and cannot distinguish a runaway from a launch that doubled traffic. - A **unit figure** tells you whether each unit of work got more expensive, which is the only one of the three that survives an argument about growth. It is also the one that takes the most setup, because the numerator has to be scoped to the same thing the denominator counts. Getting spend attributed to a service — rather than to an account that hosts nine of them — is a prerequisite and a separate subject with its own mechanisms. Without it the ratio is arithmetic performed on the wrong numerator, which is worse than no ratio at all, because it looks authoritative.
- Spend rose by a third and completed orders rose by a half. What does unit cost per order do?It falls, to roughly nine tenths of what it was — 1.33 divided by 1.5. Each order got cheaper even though the bill got bigger, so the rise is the expected consequence of doing more work rather than evidence of a regression.
- Why report a variable unit cost alongside the all-in one?Because the all-in figure improves automatically as volume grows: the fixed share is spread over more units with no engineering behind it. The variable figure is what the next unit of work actually costs, so it is the one that reveals a design that fails to amortise.
- What can a good unit cost hide?Absolute size and value. A service can have an excellent cost per unit of work and still be spend worth removing, because the ratio says nothing about whether the work needed doing. It also silently depends on the numerator being scoped to the same thing the denominator counts.
Fuel economy rather than the fuel bill. A bigger bill after a longer drive is not a fault; litres per hundred kilometres is the figure that tells you whether the car got worse.
saying these in an interview costs you the question
- Treats a rising total bill as proof something regressed
- Uses total requests as the denominator for a user-facing service
- Reports the all-in figure falling as an efficiency win during growth
- Picks a denominator only the platform team recognises
- Assumes a good unit cost means the spend is justified