skip to content

What must you change in a DeepSeek API request to enable context caching?

level: juniorimportance: should knowfreq 52%

answer

  1. nothing to switch on
  2. no field, no header, no marker
  3. storage is charged at zero
  4. 64-token unit, automatic eviction

basics

~20 s

Nothing. DeepSeek's context caching is on by default for every account on both deepseek-chat and deepseek-reasoner. There is no cache parameter, marker or header to send, and DeepSeek charges nothing to store the cached content.

solid answer

~40 s

This is a trick question, and the answer is that you change nothing. DeepSeek runs context caching automatically for all users on both `deepseek-chat` and `deepseek-reasoner` — there is no opt-in field in the request body, no marker to place on a message, and no header to set. You discover whether it worked only after the fact, by reading `prompt_cache_hit_tokens` in the response's `usage` object. Storage is free: DeepSeek bills the discounted hit rate on matched input and charges nothing for keeping the content on disk. Two limits are worth knowing. Content is stored in units of 64 tokens, so very short prompts are never cached, and entries that go unused are evicted automatically, so a warm prefix is best-effort rather than guaranteed. Cache content is scoped to your own account.

code

python · 15 lines
python
from openai import OpenAI

client = OpenAI(api_key="sk-your-key", base_url="https://api.deepseek.com")

# No caching parameter appears anywhere in this request.
resp = client.chat.completions.create(
    model="deepseek-chat",
    messages=[
        {"role": "system", "content": "You are a contract analyst. " * 200},
        {"role": "user", "content": "Summarise the termination clause."},
    ],
)

# The only place caching shows up is the response.
print(resp.usage)

go deeper

for a junior

Answer confidently that there is no parameter to send: caching is automatic. Then name where you would look to confirm it worked, which is prompt_cache_hit_tokens in the response usage object.

for a middle

Add the economics and the limits: no storage fee, discount applies to input only, a 64-token storage unit below which nothing is cached, and automatic eviction of unused entries.

for a senior

Show how you would prove it in production — a two-call experiment on a real prompt, plus a dashboard on hit ratio so a prompt change that quietly kills reuse is caught at deploy time rather than at invoicing time.

for a principal

Speak to portability: a design that depends on automatic, unpriced caching behaves differently on a vendor where caching is explicit and metered, so treat the assumption as a documented constraint before making DeepSeek a single-vendor dependency.

## The answer is "nothing", and that is the point Interviewers ask this because candidates who have used other vendors reach for a parameter that does not exist here. On DeepSeek there is no request-side switch for context caching. It is enabled by default for every account, on both the `deepseek-chat` and `deepseek-reasoner` models, and the request body you send is exactly the one you would send if caching did not exist. The only way to know a hit occurred is to look at the response afterwards: `usage.prompt_cache_hit_tokens` reports how much of the input came out of the cache. This matters practically. If your job is to "turn on caching to cut the DeepSeek bill", there is nothing to turn on — the discount is either already reaching you or it is not, and the diagnosis lives in the usage numbers, not in the request. ## The cost model: free storage, discounted hits DeepSeek's context cache carries no storage fee. There is no per-hour charge for keeping a prefix resident and no separate write charge for putting it there. The entire commercial effect is on the input line: tokens that match cached content are billed at a cache-hit input rate that recent price sheets put far below the normal input rate, and everything else is billed normally. Output is unaffected. That free-storage property is unusual enough to be worth stating out loud in an interview, because it changes the calculus. When storage is free and there is no write premium, there is no downside case to worry about — no scenario where a low hit rate makes caching cost you more than not caching. The worst outcome is that you simply pay the ordinary input price. ## The limits that make hits fail Three facts explain nearly every "why did I get zero hits" question. **A 64-token storage unit.** DeepSeek stores cache content in units of 64 tokens; content below that size is not cached. Small prompts — a one-line classification instruction, a short user turn — will simply never register a hit, and that is expected rather than broken. **Automatic eviction.** Cached content that goes unused is cleared automatically after a period of disuse, on the order of hours to days rather than minutes. Nothing in the API lets you pin an entry, extend a lifetime, or pre-warm one deliberately. A workload with long idle gaps between calls of the same shape should expect the first call after each gap to miss. **Per-account scope.** Cache content belongs to your account. Another customer sending an identical system prompt does not warm yours, and yours does not leak to them. This is a common interview probe dressed up as a security question, and the correct answer is that entries are not shared across accounts. ## What you actually control Since there is no parameter, the only influence you have is the shape of the requests you send: how much stable content sits at the front of the prompt, and how consistently that content is byte-identical from call to call. The general design principles for structuring a prompt so a cache can match it are provider-independent and worth learning once, separately. What is DeepSeek-specific — and what this question is really testing — is that those principles are the *only* lever, because the API exposes no other. ## Model coverage and how to verify Both production model IDs benefit. On `deepseek-reasoner` the discount still applies to the input side only; the model's chain-of-thought output is billed as output regardless, so a reasoner workload with a fully cached prompt can still be dominated by generation cost. Verification is a two-call experiment anyone can run in a minute: send the same long prompt twice and compare `prompt_cache_hit_tokens` across the two responses. The first call typically reports zero, the second reports most of the prompt. If the second call also reports zero, the cause is one of the three limits above or a prompt that changed between calls — not a missing configuration flag. ## Interview traps The strongest wrong answers sound plausible: inventing a `cache_control` block on the first message, an `enable_cache` boolean, or a request header; claiming an hourly storage fee; or asserting that cache hits are free rather than cheap. Another trap is claiming caching must be requested per model. None of these exist on DeepSeek. Say "nothing to enable, read the usage object to confirm, storage is free" and you have covered the whole surface.

  • If there is no way to enable it, how do you prove caching is working for a given prompt?
    Run the same request twice and compare `usage.prompt_cache_hit_tokens`. The first call normally reports zero because nothing is stored yet; the second should report most of the prompt. In production, log the hit count with every call and chart `hit / prompt_tokens` over time, so a prompt edit that destroys reuse shows up as a step change in the ratio rather than as a surprise on the bill.
  • What happens to a cached prefix that goes untouched for a day?
    It is evicted. DeepSeek clears cache content that has not been used for a while, on the order of hours to days, and gives you no API to pin, extend or pre-warm an entry. The practical consequence is that a bursty workload pays full input price on the first call of each burst, so hit-rate targets should be set against your real traffic pattern rather than against an idealised steady stream.
  • Has DeepSeek ever offered pricing discounts beyond the cache-hit rate?
    Yes — DeepSeek has run an off-peak discount, a UTC-defined window during which token rates were reduced, with a deeper reduction for the reasoner than for the chat model. That programme has been revised as DeepSeek cut headline prices, so treat any specific window or percentage you remember as stale and read the current pricing page. The durable point for an interview is the mechanism: a time-window discount stacks on top of the hit/miss split rather than replacing it.

saying these in an interview costs you the question

  • Invents a cache_control field or enable_cache flag in the request
  • Claims DeepSeek charges an hourly cache storage fee
  • Expects caching to work on very short prompts
  • Assumes a cached prefix survives indefinitely once created
  • Believes one account's cached content can serve another account

context