In DeepSeek's API response, what do prompt_cache_hit_tokens and prompt_cache_miss_tokens mean?
answer
- two buckets, one bill
- input is split, not added
- hit and miss are priced differently
- they sum to prompt_tokens
basics
~20 sDeepSeek splits a request's input tokens into two billed buckets: prompt_cache_hit_tokens, served from its context cache at a steeply discounted input rate, and prompt_cache_miss_tokens, billed at the full input rate. The two always sum to prompt_tokens.
solid answer
~40 sDeepSeek bills input at two different prices, so its `usage` object reports the split. `prompt_cache_hit_tokens` is the part of your prompt that DeepSeek recognised from its on-disk context cache and did not have to recompute; it is charged at a cache-hit input rate that recent price sheets put roughly an order of magnitude below the normal one. `prompt_cache_miss_tokens` is everything else in the prompt, charged at the standard input rate. The invariant to remember is that hit + miss equals `prompt_tokens` — they are a partition of the input, not an extra line item added on top. Output is untouched: `completion_tokens` is billed at the completion rate whether the prompt hit the cache or not. Nothing in the request enables this; the split simply appears, and a first-of-its-kind prompt legitimately reports zero hits.
code
json · 9 lines{
"usage": {
"prompt_tokens": 4200,
"completion_tokens": 180,
"total_tokens": 4380,
"prompt_cache_hit_tokens": 4096,
"prompt_cache_miss_tokens": 104
}
}go deeper
Know the two field names and that they are input-side only. Say plainly that a hit is cheaper input, not free input, and that output is billed the same either way.
Be ready to state the invariant that hit plus miss equals prompt_tokens, and to compute a per-call cost as a four-term sum over hit, miss and completion tokens at their three rates.
Show that you log all four counters per call and track the hit ratio as a deploy-safety metric, since a prompt change that kills reuse shows up here long before it shows up in a monthly invoice.
Own the fact that DeepSeek returns tokens, not money. Decide who converts usage into per-feature cost, how rate changes are version-controlled, and what hit-ratio regression is worth blocking a release for.
## The two fields Every DeepSeek chat completion response carries a `usage` object. Alongside the familiar `prompt_tokens`, `completion_tokens` and `total_tokens`, DeepSeek adds two vendor-specific counters: - `prompt_cache_hit_tokens` — the number of input tokens that DeepSeek served out of its context cache. - `prompt_cache_miss_tokens` — the number of input tokens it had to process normally. These two exist for exactly one reason: DeepSeek charges two different prices for input, and you cannot reconstruct the bill from `prompt_tokens` alone. They are a partition, so `prompt_cache_hit_tokens + prompt_cache_miss_tokens == prompt_tokens` on every response. If you add them *to* `prompt_tokens` when estimating spend, you will roughly double-count your input. ## Why a hit is cheaper, in billing terms When DeepSeek recognises that the leading portion of your prompt is one it has already processed and stored, it can reuse that stored state instead of recomputing it. The compute it skipped is the compute you are not charged for, so those tokens move onto a cache-hit input rate. Across DeepSeek's published price sheets the cache-hit rate has consistently been a small fraction of the miss rate — in recent generations of the price list, close to a tenth of it. The exact figures have changed several times as DeepSeek cut prices, so quote the ratio in an interview and look up the current numbers before you build a budget on them. Two things a cache hit is *not*: - It is not free. Discounted input is still input. A workload with a 100% hit rate still has a nonzero input line. - It is not a discount on output. `completion_tokens` is billed at the completion rate regardless of what happened on the input side. On `deepseek-reasoner`, the model's chain-of-thought tokens count toward `completion_tokens`, so they sit on the output meter too. ## Reading a real response Suppose a call reports `prompt_tokens: 4200`, `prompt_cache_hit_tokens: 4096`, `prompt_cache_miss_tokens: 104`, `completion_tokens: 180`. The reading is: about 4k tokens of the prompt were already known to DeepSeek and were charged cheaply; roughly a hundred tokens — the tail that differed from anything seen before — were charged at full input price; and 180 output tokens were charged at the output rate with no discount available. Per-call cost is a four-term sum: hit × hit-rate + miss × miss-rate + output × output-rate, divided by a million if rates are quoted per million tokens. That arithmetic is the whole point of the fields. A team that logs only `total_tokens` cannot tell a cheap call from an expensive one of the same length, because the same 4,380 total tokens can differ several-fold in price depending on where the hit/miss line fell. ## What zero hits means A response with `prompt_cache_hit_tokens: 0` is not an error and usually not a misconfiguration. The most common causes are benign: this is the first call with this prompt shape, so there was nothing stored yet; or the prompt is very short — DeepSeek stores cache content in units of 64 tokens and content below that threshold is not cached at all; or the entry aged out, since unused cache content is evicted automatically after a period of disuse. Cache content is also scoped to your own account rather than shared across users, so another customer's identical prompt does not warm yours. The practical consequence is that hit rate is a *statistic over traffic*, not a per-call guarantee. Alerting on "any call with zero hits" produces noise; tracking the hit ratio across a rolling window of calls tells you something real. ## Turning the fields into an operational signal The useful habit is to log all four numbers — hit, miss, completion, and the model ID — with every call, then aggregate. Two derived metrics pay for themselves. First, the cache-hit ratio, `hit / prompt_tokens`, which tells you whether a prompt-shape change quietly destroyed reuse: a deploy that drops the ratio from 0.9 to 0.05 is visible immediately in this number and invisible in latency dashboards. Second, the modelled cost per call from the four-term sum above, which lets you attribute spend per feature — DeepSeek's API returns token counts, not a currency amount, so the money figure is one you compute yourself. ## Common misreadings Saying "cache hits are free", adding the two buckets on top of `prompt_tokens`, expecting the output price to drop, or assuming you must send a cache-control marker to get hits are all wrong in ways an interviewer will hear straight away. The last one matters most: on DeepSeek this accounting appears with no request-side opt-in at all.
- Does a cache hit change how the output tokens are billed?No. Context caching discounts input only. `completion_tokens` is charged at the model's completion rate whether the prompt was fully cached or entirely new. On `deepseek-reasoner` the chain-of-thought tokens are counted into `completion_tokens` as well, so they sit on that same undiscounted meter. If your calls produce long answers, a high hit ratio can look impressive and still barely move the invoice.
- Why would a call report prompt_cache_hit_tokens of zero?Usually because nothing was there to hit. The prompt shape may be new to your account, the content may be below the 64-token unit DeepSeek uses for cache storage, or an entry may have been evicted after a period without use. Cache content is scoped per account, so someone else's identical prompt does not warm yours. Zero hits on a first call is expected behaviour, not a fault.
- How do you turn these counters into a cost figure the finance team can use?DeepSeek returns tokens, never a currency amount, so you compute it. Log hit, miss and completion counts plus the model ID for every call, keep the three per-million rates in version-controlled config, and multiply. Aggregate by feature tag to attribute spend. The account balance endpoint tells you remaining credit, not what any individual request cost, so per-request attribution has to be built on the usage object.
saying these in an interview costs you the question
- Says cache-hit tokens are free rather than discounted
- Adds hit and miss tokens on top of prompt_tokens, double counting
- Claims a cache hit also discounts output tokens
- Assumes you must send a cache-control marker to get hits
- Treats a first call with zero hits as a caching failure