skip to content

For a DeepSeek workload, how do you model per-request cost and prove caching paid off?

level: principalimportance: should knowfreq 34%

answer

  1. caching touches only one side
  2. three terms, one discount
  3. output rate is unchanged
  4. reasoner thinks on the output meter
  5. measure cost per unit of work

basics

~20 s

Model each call as a three-term sum: cache-hit input tokens, cache-miss input tokens and completion tokens, each at its own rate. Caching discounts only the first term, so an output-heavy or reasoner workload can show a 90% hit rate and barely move the bill.

solid answer

~50 s

DeepSeek returns tokens, never money, so the cost model is yours to build. Per call it is three terms: `prompt_cache_hit_tokens` at the discounted input rate, `prompt_cache_miss_tokens` at the full input rate, and `completion_tokens` at the output rate. Only the first term is what caching touches. That is why a high hit ratio can be a vanity metric: on `deepseek-reasoner`, chain-of-thought tokens land in `completion_tokens` and are billed as output, so a workload that thinks at length is dominated by the term caching cannot reach. To prove savings, log the three counters and the model ID with every call, tag them by feature, and compare modelled cost per unit of business work — per resolved ticket, per document — before and after the change, not cost per token. Then aim the real levers: output length, `max_tokens`, and whether the request needed the reasoner at all.

code

python · 17 lines
python
RATES = {  # illustrative USD per 1M tokens; pin real values with an effective date
    "deepseek-chat": {"hit": 0.028, "miss": 0.28, "out": 0.42},
}

def call_cost(usage: dict, model: str) -> float:
    r = RATES[model]
    return (
        usage["prompt_cache_hit_tokens"] * r["hit"]
        + usage["prompt_cache_miss_tokens"] * r["miss"]
        + usage["completion_tokens"] * r["out"]
    ) / 1_000_000

cached = {"prompt_cache_hit_tokens": 7936, "prompt_cache_miss_tokens": 64,
          "completion_tokens": 1500}
cold = {"prompt_cache_hit_tokens": 0, "prompt_cache_miss_tokens": 8000,
        "completion_tokens": 1500}
print(call_cost(cached, "deepseek-chat"), call_cost(cold, "deepseek-chat"))

go deeper

for a junior

Know that cost comes from tokens, not requests, and that the response tells you the counts. Be able to say caching discounts input only.

for a middle

Write the three-term formula from the usage object and explain why a cached prompt with a long answer barely changes the bill, especially on deepseek-reasoner.

for a senior

Show the instrumentation: log hit, miss and completion counts with a feature tag and model ID, keep rates in version-controlled config with effective dates, and decompose spend before choosing an optimisation.

for a principal

Own the measurement contract — cost per unit of business work rather than per token, alerting on that ratio rather than on total spend, controlled before/after comparisons, and prepaid balance treated as an availability risk.

## Why you have to build the model The DeepSeek API gives you token counts and nothing else. There is no cost field on the response, and the account balance endpoint reports remaining credit rather than what any individual request consumed. So any statement about per-feature spend, cost per customer, or the return on a caching change is something your own instrumentation produces. If nobody owns that model, the organisation is left comparing a monthly invoice against a vague memory of what shipped that month. ## The three-term model For a single call, cost equals hit tokens times the cache-hit rate, plus miss tokens times the full input rate, plus completion tokens times the output rate. All four inputs come from the response `usage` object plus the model ID; the three rates come from DeepSeek's price sheet and belong in version-controlled config, because they have changed several times as DeepSeek cut prices. Pin the rates with an effective date so historical cost figures stay reconstructible after a price change. The structure of that formula is the whole insight. Context caching moves tokens from the second term into the first. It cannot touch the third at all. ## The trap: hit rate as a vanity metric Consider a support-triage feature with a large stable instruction block. After a prompt restructure the hit ratio goes from near zero to about 0.95, and the team declares victory. The invoice moves four percent. What happened is that the input line was already the smaller share: the model was producing long answers, and output is priced well above even the full input rate. Ninety-five percent of a small number is still a small number. The effect is sharper on `deepseek-reasoner`. Its chain-of-thought tokens are counted into `completion_tokens` and billed as output, so a reasoning-heavy call can spend the large majority of its cost on tokens the user never sees and caching can never discount. Any cost model that ignores this will systematically under-predict reasoner spend, and any optimisation programme that starts with caching on a reasoner workload is optimising the wrong term. The corollary is the answer an interviewer is listening for: **before optimising, decompose actual spend into the three terms and optimise the biggest one.** Hit ratio is a diagnostic for whether reuse is working; it is not a proxy for money saved. ## Proving the change actually paid A credible before/after needs three things. **A stable denominator.** Cost per token always falls when you cache and tells you nothing about whether the business got cheaper. Measure cost per unit of work — per ticket resolved, per document processed, per active user per day. A change that halves per-call cost while doubling call volume is not a saving. **Attribution.** Tag every call with the feature, prompt version and model at emit time. Without a tag you can only see a fleet aggregate, and the moment two features share a key you can no longer say which one moved. **A controlled comparison.** Traffic mix drifts, so a week-over-week total is weak evidence. Compare the same cohort or the same replayed request set through both prompt versions, and hold the model ID fixed so a routing change does not get credited to caching. ## The levers that actually move the number Once the decomposition is in front of you, the choices become concrete. If output dominates: shorten what you ask for, set `max_tokens` as a real ceiling rather than a safety value, stop asking for restated context in the answer, and question whether the request needed `deepseek-reasoner` when `deepseek-chat` would do — the two models carry different price sheets and different output-token profiles. If input dominates and the hit ratio is low, the reuse work is the right work, and the fix is in how consistently the stable portion of your prompts is reproduced call to call. If the miss term is large but irreducible because every request is genuinely unique, caching was never going to help and saying so early is the valuable contribution. ## The governance angle At the level where this question is asked, the deliverable is not a spreadsheet — it is a durable capability. Someone must own the rate config and its effective dates, the cost-per-unit-of-work dashboard, and an alert on cost per unit rather than on total spend, since total spend rises with healthy growth and hides regressions. A prompt change that quietly halves the hit ratio should surface as an alert within a day, not as a surprise at the end of the month. One more governance point specific to DeepSeek: it is prepaid. Balance is therefore an availability signal as well as a cost one, since exhausting it fails requests outright. Forecasting from the same model that produces your per-call cost gives you the runway number to alert on. ## What a strong answer sounds like "Three terms, caching only touches one, decompose before optimising, measure cost per unit of work with tagged traffic, and expect reasoner workloads to be output-dominated." That is the shape. A weak answer is "turn on caching and watch the bill fall" — on DeepSeek there is nothing to turn on, and the fall may not come.

  • Your hit ratio is 0.95 and spend is flat. What is your first move?
    Decompose actual spend into the three terms rather than tuning caching further. A flat bill at that hit ratio almost always means output dominates — long generations, or `deepseek-reasoner` chain-of-thought billed into `completion_tokens`. The levers are then output length, a real `max_tokens` ceiling, dropping restated context from the answer format, and asking whether every request genuinely needed the reasoner rather than the chat model.
  • Why measure cost per unit of business work instead of cost per token?
    Because cost per token falls almost by definition when caching works, while the thing you are accountable for is what the feature costs to run. A change can halve per-call cost and still raise the bill by tripling calls, or leave per-token cost untouched while removing a whole retry loop. Pick a denominator the business recognises — per ticket, per document, per active user — and hold it fixed across the comparison.
  • How do you keep historical cost figures valid when DeepSeek changes prices?
    Store the rates as versioned config keyed by an effective date and stamp each logged call with the rate version used, rather than recomputing history at today's prices. Otherwise a price cut silently rewrites last quarter's numbers and makes any before/after comparison meaningless. Since DeepSeek is prepaid, also forecast runway from the same model and alert on remaining balance well above zero, because exhaustion is an outage.

saying these in an interview costs you the question

  • Assumes a high cache-hit rate always cuts the bill
  • Forgets reasoner chain-of-thought is billed as output tokens
  • Models spend from request counts instead of token counts
  • Treats cache hits as zero-cost input in the model
  • Reports cost per token instead of cost per unit of work

context