What makes a burn-rate LLM spend alert better than a monthly-budget threshold?
answer
- an integral detects a step change late
- rate beats total
- two windows: fast plus stable
- alert at a dimension you can act on
- a page is not a cap
basics
~20 sA cumulative threshold fires long after the damage starts and names no cause. Burn-rate alerting compares current spend per hour against a trailing baseline, sliced by tenant and feature, so an anomaly pages within an hour and arrives with the culprit attached.
solid answer
~50 sConsider a budget alarm that trips at 70% of the monthly allowance on day nine because one customer scripted a nightly re-run. It is technically correct and operationally useless: the overspend began eight days earlier, and the alert says nothing about who or what caused it. A burn-rate alert instead watches cost per unit time against a trailing baseline — typically two windows, a short one for fast detection and a longer one to suppress flapping — and evaluates it per tenant, per feature and per model so the notification carries its own diagnosis. Pair it with **unit-cost** alerts: mean tokens per call and cost per successful invocation catch a prompt or effort change that doubles per-call spend at flat traffic, which no total-spend threshold sees quickly. Finally, keep alerting separate from enforcement — a page is not a control. Hard per-tenant and per-key caps belong in the request path, where a call can actually be refused, queued or downgraded.
go deeper
Know that spend should be watched as a rate over time, not only as a running monthly total, and that alerts need to say which tenant or feature moved.
Explain the two-window burn-rate pattern and the trailing baseline, and name the unit-cost metrics — tokens per call, cost per successful invocation — that catch regressions at flat traffic.
Demonstrate the full loop: burn-rate plus unit-cost alerts sliced by tenant and feature, synchronous caps in the gateway, per-feature kill switches, and cost-per-invocation as a canary gate before rollout.
Own the policy: which tenants get hard versus soft caps, what overage costs and who approves it, the escalation path when a whale's burn is legitimate, and the projection reported to finance each week.
## Why cumulative thresholds fail A month-to-date threshold is an integral of everything that already happened. Three problems follow. **It is late by construction.** Spend must accumulate before the line is crossed, so a step change in burn is detected only after it has run long enough to matter. The day-nine, 70%-of-budget page is an eight-day-old incident. **It carries no attribution.** "Spend is high" is not actionable. The on-call engineer starts an investigation from zero, which is exactly the time when the query they need is a per-tenant, per-feature breakdown. **It conflates growth with regression.** Healthy adoption and a runaway retry loop both push the total up. A threshold cannot distinguish them, so it either fires on good news or is set so loose it never fires at all. ## Burn-rate alerting Borrow the shape from SLO practice: alert on the *rate*, evaluated over two windows. - A **short window** (say 30–60 minutes) gives fast detection. - A **long window** (say 6–24 hours) must also be violated, which suppresses noise from spiky traffic. - The comparison baseline is a trailing period with the same weekly shape — the same hour last week beats "yesterday" for workloads with a weekday cycle, and matters a lot for business-hours SaaS. Evaluate it at the dimensions you can act on: total, per tenant, per feature, per model, per effort level. The per-tenant evaluation is what makes a scripted nightly re-run visible within an hour instead of a week, and it makes the alert self-diagnosing. ## Unit-cost alerting, which is the one people miss Totals move with demand; unit costs move with the system. The high-signal metrics are: - **Mean tokens per call**, split by prompt version — a prompt edit that adds context or a new tool definition shows here immediately. - **Reasoning tokens per call**, split by effort setting — a config change or harder traffic mix shows here. - **Cost per successful invocation** — charges retries and failures to the work that caused them, so a rising failure rate reads as a cost regression, which is what it is. - **Token-weighted cache-read share** — a collapsed share means prefixes stopped being reusable, and the input bill goes up with no other change. These are also the right gates for a deploy: compare cost per invocation on a canary against the current version and block a rollout that regresses it beyond a threshold. Catching it pre-rollout is far cheaper than any alert. ## Alerting is not enforcement An alert informs a human; a cap changes behaviour. Both are needed and they live in different places. - **Enforcement** sits in the request path: a gateway or middleware that keeps a per-tenant and per-API-key spend counter in a fast store, and once a cap is hit refuses the call, queues it, or serves a reduced-cost variant, with a clear error the product can render. It must be synchronous, so it cannot depend on the analytics pipeline. - **Tiering** the caps matters: free-tier and trial keys deserve tight, low ceilings because they are the natural target for abuse; paid tenants usually get soft caps that notify an account owner and allow overage rather than breaking their workflow mid-month. - **Kill switches** per feature and per model let you stop a specific runaway without taking the product down. ## Designing the alert so it is answerable A good spend page includes: which dimension violated (tenant, feature, model, effort), current burn versus baseline, the projected month-end figure if the burn continues, a link to representative calls, and the runbook step — throttle this key, roll back this prompt version, or accept it and raise the budget. An alert that says only "spend high" trains people to acknowledge and move on. ## Thresholds worth having anyway Cumulative budgets still belong on the board — as *financial* guardrails with a long fuse, and as a projection: "at the current 7-day burn, month-end lands at 1.4x budget" is a genuinely useful weekly signal. The mistake is making the cumulative line the primary detection mechanism rather than the reporting one.
- Why use two windows rather than one short window for burn-rate alerting?A single short window is noisy: normal traffic spikes, batch jobs and retries breach it constantly, and the alert gets muted. Requiring a longer window to be violated too means only sustained overspend pages, while the short window keeps detection fast. It is the same multiwindow trick SLO burn-rate alerts use to trade sensitivity against precision.
- Traffic is flat and per-call tokens are flat, but cost per successful invocation is climbing. What is happening?Work is failing and being retried, so the same number of successes is paid for more than once. Look at status breakdowns, tool errors, schema-validation failures and timeouts. This is why cost per successful outcome belongs on the dashboard: it is the only unit cost that makes reliability regressions visible as money.
- How would you set spend caps for free-tier keys versus enterprise tenants?Free and trial keys get tight hard caps enforced synchronously — they are the abuse surface and there is no revenue to protect. Enterprise tenants normally get soft caps that alert an account owner and permit overage, because refusing their traffic mid-month is a worse outcome than the extra spend. Both need a per-feature kill switch independent of the tenant caps.
- Can a burn-rate alert be gamed by a slow, steady increase?Yes — a gradual creep stays inside the trailing baseline because the baseline moves with it. That is what the cumulative projection is for: a weekly "at current burn, month-end is 1.4x budget" report catches drift that rate alerts normalise away. The two mechanisms cover opposite failure shapes and you want both.
saying these in an interview costs you the question
- Relying on a month-to-date percentage as the primary detection signal
- Alerting only on aggregate spend, with no per-tenant or per-feature slice
- Assuming an alert prevents overspend, with no cap in the request path
- Comparing burn to yesterday for a workload with a strong weekday cycle
- Watching totals only, so a prompt change that doubles per-call cost goes unnoticed