Which token counters belong on an LLM cost dashboard beyond input and output?
answer
- two counters no longer reproduce the invoice
- some billed generation is never displayed
- one input number hides two prices
- weight the hit ratio by tokens
- effort is a cost dimension, not a setting
basics
~20 sReasoning (thinking) tokens, which are billed at generation rates but never shown, and the cache-write versus cache-read split of input tokens, which carry very different unit prices. Without those three extra counters, a dashboard cannot explain the invoice.
solid answer
~50 sA usable cost line needs at least five counters per call: input tokens, output tokens, reasoning tokens, cache-write input tokens and cache-read input tokens. Reasoning tokens matter because they are billed like generated output while being invisible in the UI — on a hard question they can run several times the visible answer, so a counter built from displayed text can miss most of the bill. Cache-write and cache-read must stay separate because they are priced differently: a read of a warm prefix is typically far cheaper than the original write, so folding both into one "input tokens" number will either flatter or wreck your margin model depending on which price you apply. Alongside the counters, carry the dimensions that *explain* movement: model, effort or thinking setting, and token-weighted cache-hit share. Then a spend spike is attributable to a cause rather than just visible.
code
python · 27 linesRATES_PER_MTOK = {
"input": 3.00,
"output": 15.00,
"reasoning": 15.00,
"cache_write": 3.75,
"cache_read": 0.30,
}
def call_cost(counters):
return sum(
counters.get(name, 0) / 1_000_000 * rate
for name, rate in RATES_PER_MTOK.items()
)
naive = {"input": 32_000, "output": 2_000}
full = {
"input": 2_000,
"cache_read": 30_000,
"reasoning": 8_000,
"output": 2_000,
}
print(round(call_cost(naive), 4))
print(round(call_cost(full), 4))
print(round(call_cost(full) - call_cost(naive), 4))go deeper
Be able to name the counters beyond input and output — reasoning tokens and the cache-write versus cache-read split — and say that they are billed at different prices.
Explain why each counter exists and what it hides when merged, and show the arithmetic on a call where reasoning tokens outweigh a large cached prefix.
Demonstrate that you dashboard unit costs and explaining dimensions, not just totals: effort mix, token-weighted cache-hit share, retry-cost share, and cost per successful outcome.
Set the standard for what the org reports: a counter and dimension schema every service emits identically, so cost is comparable across features and a rate-card or model change can be modelled from stored counters rather than re-instrumented.
## Why two counters are not enough "Input tokens and output tokens" was an adequate model when a call was one prompt in, one completion out, at two flat prices. Current billing has more axes, and each one can dominate a bill on its own. A dashboard that cannot reproduce the invoice from its own counters cannot be used to price a feature or to explain a spike. ## The counters **Input tokens.** Everything sent for the call — user message, system instructions, conversation history, retrieved context, and the tool or function definitions. Tool schemas surprise people: a large toolset can contribute thousands of tokens to *every* call in a loop, so it is worth being able to see input tokens per call trending up after someone adds tools. **Output tokens.** The generated text or structured payload, normally the most expensive per token. **Reasoning (thinking) tokens.** Reasoning models generate an internal chain before answering, billed at generation rates and typically not returned verbatim. Their volume is highly variable and has no natural ceiling: on a hard task they can be several times the visible answer. Two consequences follow. First, they need their own dashboard line — averaged into "output" they hide the fact that difficulty, not traffic, moved the bill. Second, they must be paired with the **effort dimension**: providers expose an effort or thinking-budget dial per request, and the same traffic at a higher setting is a different cost profile. Meter effort as a label so you can see the mix. **Cache-write and cache-read input tokens.** Where a provider caches a repeated prompt prefix, writing the prefix and reading it later are billed at different rates — a read is typically a large discount, a write sometimes a premium over normal input. Recording one merged input number forces you to pick a single wrong price. It also destroys the most useful ratio on the dashboard: the **token-weighted cache-hit share**, meaning cache-read tokens divided by total input tokens. Weight by tokens, not by requests, because a hit on a 40k-token prefix and a hit on a 500-token prefix are not comparable events. (How caching is designed and how long entries live is a separate subject; here it is purely an accounting dimension.) ## Dimensions that explain the counters Counters tell you *how much*; dimensions tell you *why*. The high-value, low-cardinality ones are model id, effort or thinking setting, cache-hit bucket, feature, and call status. With those, "spend up 40% week over week" resolves into a concrete story: more traffic, a shift to a bigger model, a higher effort mix, a collapsed cache-hit share, or more retries. Without them you have a number and a guess. ## Cost per what? Raw totals track traffic, so express the same data as unit costs: cost per request, cost per user-visible invocation (one action may fan out to many calls), and cost per successful outcome — which is the honest one, because it charges failed and retried work to the tasks that caused it. A prompt edit that doubles input tokens at flat traffic barely moves the daily total for a day but jumps mean tokens per call immediately. ## Worked shape Take a call that reads a 30k-token cached prefix, adds 2k fresh input, thinks for 8k tokens and returns 2k visible tokens. A naive two-counter view records 32k in and 2k out. The real bill is dominated by the 8k reasoning tokens at generation price, while the 30k cached read is comparatively cheap — the opposite conclusion to the naive one, which makes prompt size look like the problem. Reserving separate counters is what lets you see that trimming the prompt saves little and capping effort saves a lot. ## Practical rules - Capture the counters in one client wrapper; never at individual call sites. - Store raw counters plus the rate-card version, and price separately. - Emit zero-valued counters explicitly so "no cache reads" is distinguishable from "not instrumented". - Keep a cost line for failed, cancelled and retried calls; they are real money. - Break out reasoning tokens on the *default* dashboard view, not behind a drill-down.
- Why weight cache-hit rate by tokens instead of by request count?Because the money is in tokens, not calls. A request-based ratio treats a hit on a 40k-token prefix and a hit on a 500-token prefix as equal, so it can read 90% while most billed input is still uncached. Cache-read tokens divided by total input tokens is the figure that moves with the bill, and it is the one to put next to spend on a dashboard.
- Traffic is flat but spend rose 30% overnight. Which dimension would you slice first?Slice by effort or thinking setting and by reasoning tokens per call, then by model. A deploy that raised the effort level, or a prompt change that makes the model think longer, moves cost with no traffic change at all. If those are flat, check cache-read share — a prefix edit can invalidate the cached prefix and push every call back to full input price.
- How would you present cost so it survives a traffic doubling?Publish unit costs alongside totals: cost per request, cost per user-visible invocation, and cost per successful outcome. Totals conflate demand with efficiency, so a doubling looks identical to a regression. Unit costs stay flat under pure growth and move only when the system itself got more expensive, which makes them the right thing to alert on.
- Should retried calls be counted against the feature or reported separately?Both. Charge them to the feature and invocation that caused them, so cost per successful outcome tells the truth, and also expose a retry-cost share as its own line. That share is an early warning: a rising fraction of spend on work that produced nothing usually means schema failures, tool errors or timeouts rather than growth.
saying these in an interview costs you the question
- Merging cache-write and cache-read into one input-token number
- Treating thinking tokens as free because they are not returned
- Reporting only totals, so a traffic rise looks like a regression
- Computing cache-hit rate per request rather than per token
- Excluding failed and retried calls from the cost line