Your gen_ai.usage token attributes don't match the provider bill — what do you check?
answer
- Missing is not the same as zero
- Streaming often reports no usage
- Two spans, one call
- Sampled spans cannot sum to a total
- Tokens are not the whole price
basics
~20 sCheck for missing usage on streamed calls, spans lost to sampling, the same call counted twice by overlapping instrumentation, retries and fallbacks recorded as extra spans, and cached or reasoning tokens the two attributes do not break out.
solid answer
~50 sWork the gap from both ends. **Under-count** usually comes from missing spans or missing attributes: streamed responses often carry no usage block unless the client explicitly requests it, so `gen_ai.usage.input_tokens` is absent rather than zero; head sampling means you are summing a fraction of traffic; and calls made outside instrumented code paths never appear. **Over-count** comes from duplication: a framework instrumentation and a provider instrumentation both wrapping the same call produce two spans with the same tokens, and retries or model fallbacks legitimately create extra spans your query is treating as extra requests. Then there is the pricing layer, which is not in the span at all: the two usage attributes do not distinguish cached input tokens or reasoning tokens, which are billed differently, and your backend's price table may be stale or missing the exact response model. Diagnose by picking one request, reading its raw spans, and reconciling that single call before touching aggregates.
go deeper
Know that gen_ai.usage.input_tokens and output_tokens can be missing rather than zero, especially for streamed responses, and that a missing attribute quietly disappears from a sum.
Explain the two directions of error — under-count from missing spans or attributes, over-count from overlapping instrumentation and retries — and check them by reading a single trace rather than tweaking the aggregate.
Demonstrate the full reconciliation: verify one known request end to end, split evaluation and background traffic by tag, account for sampling, and report the residual gap with a named cause instead of claiming an exact match.
Own where cost numbers come from at all — unsampled metrics rather than sampled traces for finance-facing figures, a governed price table including negotiated rates, and a stated accuracy target with the known unattributed remainder.
## Why this question is asked The cost dashboard is the first thing a business stakeholder looks at, and it is almost never exactly right on the first attempt. The span attributes are honest about what the instrumentation saw; the invoice is what the provider charged. The gap between them is a checklist, and knowing the checklist is what separates someone who has run this in production from someone who has read the docs. ## Under-count causes **Streaming without usage.** When a response is streamed, token counts are not implicit in the chunks — the provider reports them only if the client asked for usage to be included. If it did not, the instrumentation has nothing to record and the usage attributes are simply **absent**. Absent is not zero: a `sum()` skips it silently, so a service that streams everything can contribute almost nothing to your token chart while burning real money. This is the single most common cause, and it is worth checking first because it is easy to confirm — look for spans that have `gen_ai.request.model` but no `gen_ai.usage.output_tokens`. **Sampling.** If traces are head-sampled at, say, 10%, your token sum is a tenth of reality, and scaling it back up is only valid if sampling is uniform and independent of request size — which it often is not once someone adds a rule that always samples errors or slow requests. Token totals are one of the strongest arguments for deriving cost from metrics rather than from sampled spans. **Uninstrumented paths.** Calls from a background job, a different language service, a notebook, or a provider whose client the instrumentation does not cover never produce spans at all. The invoice sees them; your trace store does not. **Dropped spans.** Exporter queue overflow under load, or ingestion rejecting oversized spans, removes requests from the aggregate — and because oversized usually means large prompt, the loss is biased toward the expensive calls. ## Over-count causes **Overlapping instrumentation.** If both a framework-level instrumentation and the underlying provider-client instrumentation are active, one logical call can produce a parent and a child span that each carry usage. Summing across all spans then double counts. The fix is to sum at one level — filter by operation or by span depth — and to confirm by opening a single trace and counting how many spans in it carry usage attributes. **Retries and fallbacks.** A retried call is genuinely two provider requests and, if both consumed tokens, genuinely two charges — so this is often not a bug in your data but a bug in your interpretation. Distinguish "requests my application made" from "requests the provider served", and be explicit about which one the dashboard shows. **Evaluation and judge traffic.** LLM-as-judge scoring, synthetic data generation and test suites hit the same API key. If they are instrumented, they inflate your "application" tokens; if they are not, they inflate the invoice relative to your traces. Either way, tag them so they can be split out. ## The pricing layer Even with a perfect token sum, cost can be wrong because tokens are not the whole billing story: - `gen_ai.usage.input_tokens` is one number and does not separate cached from uncached input, which many providers bill at different rates. A prompt-caching rollout can cut the bill sharply with no visible change in the attribute. - Output tokens for reasoning-style models may include tokens you never see in the text, and those are billed. - The backend multiplies by its own price table, keyed by model. If your `gen_ai.response.model` is a snapshot the table does not know, the entry may fall back to a default price or be excluded entirely. - Batch or discounted tiers, and per-contract rates, are invisible to the tracing tool. ## How to actually diagnose it 1. **Reconcile one request, not the aggregate.** Make a single call with known content, find its trace, and compare its attributes against the provider's own usage reporting for that call. Nearly every cause above is visible at n=1. 2. **Count spans per logical call** in that trace to catch duplication. 3. **Query for spans missing usage attributes**, grouped by service and by streaming versus non-streaming. 4. **Check sampling configuration** before trusting any total, and prefer an unsampled metric for cost if the volume justifies it. 5. **Split traffic by tag** — application, evaluation, background jobs — so the comparison against the invoice is like for like. 6. **State the residual error.** "Traces account for 94% of the invoice, the remainder is uninstrumented batch jobs" is a far better answer than a number that claims to be exact. ## The framing to give an interviewer The span is a record of what the instrumentation observed; the invoice is a record of what the provider served. Reconciling them means naming every place those two populations differ — missing spans, missing attributes, duplicated spans, extra provider requests, and billing dimensions the attributes do not carry.
- Why does a fully streaming service often contribute almost nothing to a token dashboard?Because streamed responses typically do not include a usage report unless the client explicitly asks for it, so the instrumentation has nothing to put in gen_ai.usage.input_tokens or output_tokens. The attributes are absent, and aggregation functions skip absent values rather than treating them as zero — so the service looks free while spending normally. Confirm by querying for spans that have a model attribute but no usage attributes.
- How do you tell double-counting from legitimate retries when tokens exceed the invoice?Open one trace. Double-counting shows as a parent and child span for the same logical call, with identical token values and overlapping time ranges. A retry shows as sibling spans, sequential in time, usually with a failed status on the first. The first is a query bug you fix by summing at one instrumentation level; the second is real provider traffic your interpretation must account for.
- Why can token totals be right while the cost chart is still wrong?Because price is applied outside the span. The backend multiplies tokens by its own table keyed on model, so an unrecognised response-model snapshot, a stale published rate, a negotiated contract price, or batch-tier discounts all produce a wrong figure from correct tokens. Cached input tokens are the sharpest case: they are inside the same input_tokens number but frequently billed at a very different rate.
- Would you build cost reporting on sampled traces?Not as the source of truth. Head sampling means your sum is a fraction of reality, and scaling it up is only valid if sampling is uniform and independent of request size — which stops being true the moment someone adds a rule that always keeps errors or slow calls. Emit token counts as unsampled metrics for finance-facing numbers, and use traces for attribution and debugging.
saying these in an interview costs you the question
- Treats missing usage attributes as zero tokens
- Sums tokens across parent and child spans of one call
- Scales a sampled total up as if sampling were uniform
- Assumes input_tokens separates cached from uncached tokens
- Blames the provider before reconciling a single request