How do you compute the cost per million tokens of a self-hosted LLM endpoint?
answer
- Dollars per hour over tokens per hour
- the denominator is the hard part
- idle GPUs bill at full rate
- prefill and decode price differently
- goodput under SLO, not peak
basics
~20 sDivide the hourly cost of the GPU by the tokens it actually produces in that hour. GPU $/hr divided by (tokens per second times 3600), times one million. The number that matters is measured throughput at your real utilization, not the card's peak.
solid answer
~50 sThe formula is `cost per 1M tokens = (GPU $/hr) / (tokens/s x 3600) x 1,000,000`. Everything hard is in the denominator. Use **measured** throughput from a load test on your own prompt-length distribution at the concurrency your SLO allows — not a vendor peak, and not a saturation number you can never reach in production. Then divide by real **utilization**: a replica busy 20% of the day costs five times per token what its saturated rate suggests, because you rent the GPU by the hour whether it is generating or idle. Price input and output tokens separately, since prefill processes a whole prompt in one batched pass and is an order of magnitude cheaper per token than decode. Finally add the costs that are not the GPU: replicas held for redundancy, storage and egress for weights, the load balancer, and the on-call engineering time. Report a blended figure with the assumptions attached.
go deeper
Know that a GPU is billed per hour regardless of traffic, so cost per token is hourly price divided by the tokens actually produced in that hour.
Do the division with real numbers and explain why measured throughput at your concurrency, not a benchmark peak, belongs in the denominator.
Produce a defensible figure: separate input and output rates, a stated utilization assumption, replica count for availability, and the load test the throughput came from.
Own the commercial framing. Decide the unit economics target the product must hit, choose between on-demand, committed and spot capacity, and set the utilization floor that makes self-hosting rational at all.
## The formula, and why the denominator is the whole exercise ``` cost per 1M tokens = hourly_cost / (tokens_per_second * 3600) * 1e6 ``` Worked example: a card you pay $3/hour for, measured at 2,500 output tokens/s aggregate under your target concurrency, produces 9 million tokens per hour. That is about **$0.33 per million output tokens** — at full saturation. Every complication below makes the real number worse than that, and the gap between $0.33 and what you actually pay is where the interview lives. ## Complication 1: utilization You rent the GPU by the hour; you do not rent tokens. A replica that is genuinely busy 20% of the day costs the same $3/hour and produces one fifth the tokens, so the real figure is **$1.65 per million**. For interactive products with a daily traffic curve, average utilization of 20-40% is normal, and it dominates every other term in the calculation. This is also why the honest metric is cost per million tokens *at your traffic shape*, and why batch workloads — which can saturate a card overnight — look so much cheaper per token than chat. ## Complication 2: input and output tokens are not the same product Prefill runs the whole prompt through the model in one heavily parallel pass; decode emits one token per sequence per step and is memory-bandwidth bound. Prompt tokens are therefore vastly cheaper to serve than generated tokens, often by roughly an order of magnitude. Commercial APIs reflect this in their price sheets, and your internal figure should too. Compute two numbers, then blend them using your actual prompt-to-completion ratio — a RAG workload with 8,000-token prompts and 200-token answers has completely different economics from a chat workload with 200-token prompts and 800-token answers, even on identical hardware. ## Complication 3: measure throughput the way you will serve it The tokens/s you divide by must come from a load test that mirrors production: your prompt-length distribution, your output lengths, your concurrency, your context cap. Two traps: - **Peak-throughput numbers** come from saturating the server with unlimited concurrency, which usually violates your latency SLO. The number you can bank is *goodput* — tokens per second delivered by requests that met the SLO. - **Short synthetic prompts** overstate throughput badly, because prefill cost scales with prompt length and long prompts also consume the KV cache that would otherwise hold more concurrent sequences. ## Complication 4: the costs that are not the GPU - **Redundancy.** One replica is not a service. Two or three replicas for availability and rolling upgrades multiply the floor cost even at low traffic. - **Idle and warm capacity.** Weights are tens of gigabytes; you cannot scale to zero and back in seconds, so some capacity is paid for while idle by design. - **Storage and network.** Weight storage, image pulls, egress on every replica start. - **The rest of the stack.** Load balancer, observability, and the engineering and on-call time to keep an inference fleet healthy — real money that never appears in a $/hr line item. - **Commitment discounts.** Reserved or committed-use pricing can be a large fraction cheaper than on-demand, at the price of a fixed floor whether traffic arrives or not. Spot capacity is cheaper still and hostile to long-running stateful serving with slow cold starts. ## How to present the number Give a range, not a point, and attach the assumptions: hardware and hourly rate, measured input and output tokens/s at a stated concurrency and SLO, assumed utilization, replica count. "$0.40-$1.20 per million output tokens at 25-70% utilization on this SKU, measured with our production prompt mix" is a sentence a finance partner can use. A single unqualified number is a sentence someone will hold you to and you will miss. ## Where candidates go wrong Using peak throughput, ignoring idle time, pricing one replica, and quoting a single blended token price for a workload whose prompt-to-completion ratio varies wildly. Each of these is off by a factor, and they compound.
- Why should input and output tokens carry different internal prices?Because prefill processes an entire prompt in one parallel pass while decode emits one token per sequence per step against memory bandwidth. Prompt tokens are therefore far cheaper to produce — commonly by around an order of magnitude — so a single blended rate misprices any workload whose prompt-to-completion ratio differs from the one you measured. Compute both and blend with your real ratio.
- Your saturated benchmark says $0.30 per million tokens but the monthly bill implies $1.50. What explains the gap?Utilization, redundancy and idle capacity. The benchmark measured a fully loaded card; production traffic has a daily curve, you run more than one replica for availability, and you keep warm capacity because loading tens of gigabytes of weights takes minutes. Divide saturated cost by real utilization and multiply by replica count before quoting anything.
- How does batching change the number, and what does it cost you?Larger batches amortize each weight read over more sequences, so aggregate tokens/s — and therefore cost per token — improves substantially up to the point where the GPU or the KV cache saturates. The price is per-user latency: each sequence waits behind more work per step. The economically correct batch size is the largest one that still meets your time-per-token SLO.
saying these in an interview costs you the question
- Quoting vendor peak throughput as the serving rate
- Ignoring idle hours because the GPU 'was there anyway'
- Pricing a single replica as if it were the service
- Using one blended token price for every workload shape
- Forgetting engineering and on-call cost entirely