skip to content

How do you establish cost per tenant per invocation before pricing a new LLM feature?

level: principalimportance: should knowfreq 34%

answer

  1. price the invocation, not the call
  2. one click can fan out to a dozen calls
  3. the mean is not where margin lives
  4. warm-cache assumptions cool at GA
  5. a price needs an enforceable ceiling

basics

~20 s

Meter first, price second. Ship the feature behind a flag to real tenants with attribution on every call, roll costs up to the user-visible invocation, and price off the p95 tenant's distribution plus tail share — never the mean.

solid answer

~60 s

Say an accounting SaaS wants to sell an "explain this variance" feature and needs cost per tenant per invocation before the plan is announced. Run a shadow-metering period: enable it for a representative slice of tenants, and make every model call carry tenant, plan, feature and invocation id — including retries, sub-agent calls, embedding calls and any judge or grader running in production. Roll up to the **invocation**, because one user click may fan out to a dozen calls, and include the ones that failed or were abandoned. Price with a pinned rate card that separates input, output, reasoning, cache-write and cache-read, then look at the *distribution*: the mean is dominated by cheap invocations while margin is set by the p95–p99 tenant and by how much of total spend the tail holds. Then stress the assumptions — a cache-read share measured at pilot traffic, an effort-level mix, and a provider price you do not control — and price with headroom plus per-plan quotas rather than betting on the average.

go deeper

for a junior

Understand that one user action can trigger many model calls, so the cost that matters is the total for the whole invocation, not for a single call.

for a middle

Explain how attribution keys ride through the fan-out into every call — retries and sub-agents included — and how a rate card with separate token classes turns counters into money.

for a senior

Show that you read the distribution rather than the mean, include failed and retried work in the unit cost, and stress-test the cache-warmth and effort-mix assumptions before they carry a price.

for a principal

Own the commercial decision: which segment is profitable at which plan, where the fair-use line and overage sit, what headroom absorbs a provider price change, and the internal cost-per-invocation budget engineering must defend after launch.

## The unit you are pricing Almost every mistake here starts with metering the wrong unit. The billable event to a customer is a **user-visible invocation** — "explain this variance" once. Underneath, that may be a retrieval step, two or three tool calls, a reasoning pass, a retry after a schema failure, and a summarisation. Cost per *model call* is an engineering metric; cost per *invocation* is the commercial one. So the telemetry has to carry an invocation id that ties the whole fan-out together, propagated through sub-agents and background continuations, and the rollup has to sum every call under it — including the ones that produced nothing. ## Attribution that survives the fan-out Carry attribution keys — tenant, plan, feature, invocation id, prompt version — in the call context, and capture them in the same wrapper that captures token usage. Two paths leak in practice: retries performed inside an SDK or framework, and any component constructed with its own client (a sub-agent, an embedder, a production judge). If those are unattributed, cost per invocation looks great and the invoice disagrees, which is the worst possible discovery order when a price has already been published. ## Pricing the tokens Apply a rate card with separate rates for input, output, reasoning, cache-write and cache-read, and record which rate-card version priced each event. This matters twice: it makes historical margin stable when prices change, and it lets you re-run the whole model against a hypothetical price list — a provider increase, a switch to a different model tier — without touching stored counters. ## Read the distribution, not the average Per-tenant invocation cost is heavy-tailed, because tenants differ in data size, question difficulty and usage patterns. Useful views: - Cost per invocation at p50, p95 and p99, per tenant and overall. - Cost per tenant per month, and the **tail share** — what fraction of total spend the top 1% and top 5% of tenants account for. - Invocations per active user per month, since a plan sold per seat is exposed to usage intensity, not just cost per call. - Cost per *successful* invocation, so retried and abandoned work is charged where it belongs. A plan priced on the mean will be profitable on most tenants and structurally loss-making on the ones that matter. The decision is usually a combination: price for the p95, cap or meter beyond a fair-use line, and treat the extreme tail as an enterprise conversation rather than a plan tier. ## Stress the assumptions before they are load-bearing - **Cache warmth.** A cache-read share measured during a pilot with concentrated traffic may not survive general availability, where prefixes are colder and more varied. Model the margin at a materially lower read share. - **Effort mix.** If effort or thinking level is chosen per request class, the cost per invocation is a weighted average over a mix that will shift as usage broadens toward harder questions. - **Reasoning variance.** Thinking-token volume has no natural ceiling and scales with difficulty; the p99 invocation can be many times the p50, so a per-invocation cost ceiling in the harness is part of the pricing story, not an afterthought. - **Provider prices and model deprecation.** You do not control either. Ask what happens to the plan if the rate rises, or if the model you priced on is retired and the successor costs more. - **Growth in usage per seat.** Features that work get used more; model invocations per seat growing two to three times after launch. ## Guardrails as part of the price A price is only defensible if the product can hold the cost envelope. That means per-plan quotas, a per-invocation cost ceiling enforced in the harness, per-tenant caps in the request path with clear product behaviour when they bind, and overage pricing agreed in advance rather than improvised during an incident. Publish an internal cost-per-invocation target the engineering team defends the way it defends a latency budget, and review it when prompts, models or effort defaults change. ## Closing the loop Reconcile the modelled spend against the provider invoice every period, and report realised gross margin per plan rather than the pre-launch estimate. When the two diverge, the cause is almost always structural — an unmetered path, a colder cache than modelled, or a usage mix shift — and each of those is a signal worth acting on well before the next pricing round. Where the honest answer is that a segment cannot be served profitably at plan pricing, say so with the distribution in hand; that is the value of doing this before the announcement rather than after.

  • Why not price on mean cost per invocation if the tail is only 1% of tenants?
    Because that 1% can hold a large share of total spend, and per-seat pricing gives you no mechanism to recover it. Look at tail share explicitly: if the top percentile is a fifth of spend, the mean is describing tenants who were never the risk. Price for the p95, meter beyond a fair-use line, and handle the extreme tail contractually.
  • Which costs belong in cost per invocation besides the successful model call?
    Retries and failed calls, sub-agent and tool-driven calls, embedding or retrieval calls, any production judge or guardrail model, and continuations that finish after the user leaves. Cancelled generations count too, since tokens produced before the abort are typically billed. Excluding them makes the pilot look cheap and the invoice look wrong.
  • Your pilot showed a 70% cache-read share. How much should the plan depend on that?
    As little as possible. Pilot traffic is concentrated, so prefixes stay warm in a way that general availability rarely reproduces. Model margin at a substantially lower read share and treat anything above that as upside. If the plan is only viable at the pilot's share, you have priced a measurement artefact rather than the feature.
  • How do you keep the cost envelope from drifting after launch?
    Treat cost per invocation as a defended budget: canary every prompt, model and effort-default change against it, alert on unit-cost regressions, enforce a per-invocation ceiling in the harness, and report realised margin per plan next to the pre-launch model. Reconcile with the invoice each period so drift is caught structurally rather than at the next pricing review.

saying these in an interview costs you the question

  • Pricing from mean cost per model call rather than per invocation
  • Omitting retries, sub-agent and judge calls from the unit cost
  • Extrapolating a pilot's warm-cache discount to general availability
  • Announcing a plan with no per-tenant cap or fair-use line
  • Assuming provider prices and the chosen model will stay put

context