skip to content

A mid-month spend forecast projects a large overrun from two weeks of actuals — what makes that projection wrong?

level: middleimportance: nice to knowfreq 33%

answer

  1. a straight line through a bumpy period
  2. days elapsed as the divisor
  3. one-off charges get amortised
  4. monthly charges counted as daily
  5. stored volume only ever grows

basics

~20 s

Most mid-period forecasts are run-rate extrapolation: spend so far divided by days elapsed, times days in the period. That breaks when a one-off charge gets multiplied across the period, a monthly-cadence charge is treated as daily, or stored volume is still growing.

solid answer

~50 s

A projection is almost always the simplest possible arithmetic — `spendSoFar / daysElapsed * daysInPeriod` — and the arithmetic assumes every remaining day looks like the average day so far. Five things routinely break that assumption: a **one-off charge** early in the period gets amortised across the whole of it; a charge that lands on a **monthly cadence** is treated as if it recurs daily; **weekday and weekend shape** makes the elapsed window unrepresentative; **stored volume grows monotonically**, so the average day so far was cheaper than the days to come; and a **deliberate change** mid-period makes the earlier days irrelevant. The fix is to forecast per charge dimension, split fixed from variable, and re-anchor after any deliberate change — then treat the number as a trigger to go and look, not as a figure to report.

code

pseudocode · 17 lines
pseudocode
daysElapsed  = 14
daysInPeriod = 30
charges      = chargesSoFar(scope = "checkout")
spendSoFar   = sum(c.amount for c in charges)

# naive: assumes every remaining day looks like the average day so far
naiveForecast = spendSoFar / daysElapsed * daysInPeriod

# decomposed: a one-off happens once, a monthly charge lands once,
# only the genuinely variable part is extrapolated
oneOff       = sum(c.amount for c in charges if c.recurrence == "one-off")
monthlyFixed = sum(c.amount for c in charges if c.recurrence == "monthly")
variable     = spendSoFar - oneOff - monthlyFixed

forecast = oneOff
         + monthlyFixed
         + (variable / daysElapsed * daysInPeriod)

go deeper

for a junior

Recall what the number is: spend so far, scaled up to a whole period. Know that it is an estimate built on an incomplete window rather than a figure anyone has committed to.

for a middle

Explain the arithmetic and at least two ways it breaks — a one-off charge amortised across the period, and a monthly-cadence charge treated as daily. Be able to say which direction each error pushes.

for a senior

Show the decomposition: split by recurrence, forecast per dimension, re-anchor after a deliberate change, and state the assumption next to the number. Explain why the projection is still the earliest signal in the billing system despite all of it.

for a principal

The judgment is what a projected figure is allowed to be used for. Treating it as a planning commitment rather than an investigative trigger is how a cost programme loses credibility the first time a one-off charge inflates it.

## What a forecast actually computes A month-end spend projection is rarely sophisticated. In its usual form it is: ```pseudocode forecast = spendSoFar / daysElapsed * daysInPeriod ``` Some implementations add a trend term, or fit the last few days more heavily than the first few. The assumption underneath all of them is the same: **the days that remain will resemble the days that have passed.** Everything that goes wrong with a forecast is a way that assumption is false. ## Five inputs that break it | Input | Direction of the error | Why | |---|---|---| | A one-off charge early in the period | Over-predicts | Divided by days elapsed and multiplied by days in the period, a single charge is counted many times over | | A charge on a monthly cadence | Over-predicts | A fee that lands once is treated as if it lands every day | | Weekday and weekend shape | Either way | A window that is mostly weekdays projects a weekday level across weekend days, and the reverse | | Monotonically growing stored volume | Under-predicts | Per-unit-time storage charges rise with the volume held, so the early days were genuinely cheaper than the late ones | | A deliberate change mid-period | Under-predicts a scale-up, over-predicts a scale-down | The elapsed window describes a system that no longer exists | The first two are the ones that produce the alarming projection on day two and the sheepish retraction on day four. A single large charge — a bulk data movement, a one-time retrieval, a migration — dominates a short elapsed window completely. The fourth is the one people get backwards. Storage charged per unit of volume per unit of time costs *more* per day as the stored volume grows, so a naive run rate built on the first fortnight of a growing dataset systematically **under**-predicts the period's total. It is the one direction in which a forecast is quietly reassuring while being wrong. ## Making the projection useful The repair is not a better curve fit; it is decomposition. 1. **Split the charges by recurrence** before extrapolating: one-off, monthly-cadence, and genuinely variable. Add the first two back once, and extrapolate only the third. 2. **Forecast per charge dimension and sum the parts.** A dimension that grows with stored volume and a dimension that tracks request count have different shapes, and a forecast on the total averages them into something that describes neither. 3. **Re-anchor after a deliberate change.** If capacity tripled on the 9th, the window that matters starts on the 9th, not on the 1st — and the forecast will be noisier for a few days because it has less data, which is honest rather than a defect. 4. **State the assumption next to the number.** "On the last six days' run rate, excluding the one-off on the 3rd" is a sentence someone can argue with. A bare projected figure is not. ## What the forecast is genuinely good for It is worth being clear about why teams keep a forecast in spite of all of the above: **a projection can cross a budget while there is still period left to act, and an actual cannot.** An actual figure crosses the line near the end, by construction. That is the whole value, and it is real — a threshold evaluated against a projected figure is the earliest signal in the entire billing system. Platforms differ on whether they let you set one; where yours does, it is the threshold worth having. It is also compounded by the reporting lag underneath it: the actuals feeding the projection are themselves hours old and provisional until the period closes, so an early-period projection is an extrapolation of an estimate. That is not an argument against forecasting. It is an argument against reporting the projected figure to anyone as though it were a commitment. ## How to read one in practice When a projection shows a large overrun, the useful sequence is short: - **Look at the daily series, not the projection.** A step tells you something changed on a date; a ramp tells you something is growing; a single tall bar tells you the projection is an artefact of one charge. - **Check for a one-off.** If removing one day's charge collapses the overrun, you have your answer and there is nothing to investigate. - **Check what changed on the day the curve bent.** Deployments, capacity changes, a new data feed, a retention setting. - **Only then ask what dimension it landed in**, which is where this stops being a forecasting question and becomes a question about what generates the charge — a different subject, owned elsewhere in this category.

  • A large one-off charge landed on day two. What does a naive run-rate projection do with it?
    It divides that charge by two days and multiplies it by the days in the period, so a single charge is counted roughly fifteen times over. That is why alarming projections cluster in the first days of a period and quietly deflate as the elapsed window grows.
  • Your dataset grows every day and nothing is ever deleted. Which way does a run-rate projection err?
    It under-predicts. Storage charged per unit of volume per unit of time costs more per day as the volume rises, so the average day so far was cheaper than any day still to come. This is the case where the forecast is comfortable and wrong.
  • Why is a threshold evaluated against a projection worth having at all, given how fragile projections are?
    Because a projection can cross a budget while there is period left to act and an actual cannot — an actual crosses near the end by construction. The fragility is a reason to treat the crossing as a prompt to look at the daily series, not as a fact to report.

saying these in an interview costs you the question

  • Reports a day-two projection as the expected period total
  • Assumes a one-off charge will repeat every day of the period
  • Thinks a growing dataset makes a run rate over-predict
  • Extrapolates the total instead of each charge dimension
  • Keeps forecasting from the start of the period after a deliberate change