What is the default TTL of Anthropic's prompt cache, and what does caching cost?
answer
- five minutes, sliding window
- every hit restarts the clock
- ttl field opts into one hour
- writes cost extra, reads cost a tenth
- one reuse repays a short-TTL write
basics
~20 sAnthropic cache entries live five minutes by default, and every hit refreshes that window. An optional one-hour TTL is requested with ttl inside cache_control. Writes cost a premium over base input price, reads a small fraction of it.
solid answer
~50 sBy default an entry in Anthropic's prompt cache lives about five minutes, and the clock is sliding — each hit refreshes it, so continuous traffic keeps a prefix warm indefinitely and a quiet gap lets it expire. You can opt into a one-hour lifetime by adding `"ttl": "1h"` alongside the type inside `cache_control`. Pricing has three input rates rather than one: a five-minute cache write costs roughly 1.25x the base input price, a one-hour write roughly 2x, and a cache read roughly 0.1x. The economics follow immediately. A five-minute write adds 25% but each reuse saves 90%, so a single reuse inside the window already pays for it; the one-hour tier adds a full 100%, so it needs about two reuses to break even and is aimed at bursty traffic with long gaps between calls.
go deeper
Know that cached prompt content is much cheaper to re-send but expires quickly — about five minutes by default — and that opting into a longer lifetime is an explicit field, not the default.
Explain the sliding refresh and the three input rates, and do the break-even arithmetic out loud: a short-TTL write costs 25% more but each reuse saves 90%.
An interviewer expects you to match TTL to observed traffic shape, spot where the write premium is being paid without reuse, and know that parallel cold starts each pay a write.
Own the policy across services: which prompts are cache-worthy at all, when the one-hour tier is justified, how warm-up is handled at deploy, and what the latency benefit is worth beyond the token savings.
## Lifetime An entry written by a `cache_control` breakpoint is short-lived. The default is roughly five minutes, and it is a *sliding* window: every read of that entry restarts the timer. This is the single most useful fact about the TTL, because it means the relevant question is not "how long does the cache live" but "how often does this prefix get used". A prompt hit every thirty seconds stays warm for hours; the same prompt hit every ten minutes never hits at all and pays the write premium every single time. An extended lifetime of one hour is available by adding a `ttl` field to the control object — `{"type": "ephemeral", "ttl": "1h"}` — and it slides the same way. It exists for workloads with real gaps: a support agent who returns to a session after a coffee break, an analyst iterating over a large document, a batch job that revisits the same corpus every few minutes. ## The three input prices Caching splits the input token price into three: - **Ordinary input** — the base rate, for everything not involved in the cache. - **Cache write** — about 1.25x base for the five-minute tier, about 2x base for the one-hour tier. You pay this once, on the miss that populates the entry. - **Cache read** — about 0.1x base. You pay this on every hit. The response tells you which is which through `cache_creation_input_tokens` and `cache_read_input_tokens`, so the bill is reconstructable per request. ## Break-even arithmetic Take base input as 1.0 per token. Without caching, N calls over the same prefix cost N. With five-minute caching, they cost 1.25 for the write plus 0.1 for each of the N-1 reads. At N = 2 that is 1.35 against 2.0 — already cheaper. The formal break-even is well under a single reuse: the 0.25 premium is repaid by 0.28 of one reuse's 0.9 saving. The one-hour tier pays 1.0 extra on the write, so it needs about 1.12 reuses — call it two calls within the hour — before it beats no caching, and it only beats the five-minute tier when the five-minute tier would have expired and re-written. That is the decision: choose one hour when your inter-request gap regularly exceeds five minutes but stays under an hour, and the prefix is large enough that re-writing it repeatedly hurts. Where caching loses is high-variance, low-reuse traffic: single-shot requests whose prefix will never be seen again pay the write premium for nothing. Marking such prompts is a small, permanent, invisible tax. ## Latency, not just cost The read price is the headline, but the latency effect is often what justifies the feature. A cached prefix skips most of the prefill work, so time to first token on a very long stable context drops substantially. For an agent that re-sends a large system prompt and tool set on every turn, that shows up directly in perceived responsiveness. ## Practical consequences Cache lifetime interacts with traffic shape, so treat it as a tuning parameter, not a constant. Two patterns are worth internalizing. First, low-QPS services benefit from the one-hour tier far more than busy ones, because a busy service refreshes the five-minute window for free. Second, a keep-alive strategy — issuing a cheap request purely to refresh a cache entry — is usually a false economy: the refresh call itself is billed, and if the prefix is large, the cheapest correct answer is the longer TTL rather than synthetic traffic. Finally, remember what the TTL does not do. It does not make a prefix match: an expired entry and a changed prefix look identical in the usage counters (a creation instead of a read), so when hit rate falls, check prefix stability before assuming the entry timed out.
- When is the one-hour TTL the wrong choice?When traffic is steady enough to refresh the five-minute entry anyway. Under continuous use the short TTL never expires, so paying the higher write premium buys nothing. It is also wrong for one-shot prefixes that are never reused, where any cache write is pure overhead, and for prompts that change faster than an hour, since a stale entry is never matched anyway.
- Does a cache hit make the request cheaper for output tokens too?No. Caching affects only the prompt side. Output tokens are generated fresh on every call and are billed at the normal output rate regardless of whether the prefix was read from cache. The savings come entirely from replacing base-rate input tokens with cache-read tokens, plus the reduced prefill latency.
- Two identical requests are sent in parallel against a cold prefix. What happens?Both are likely to miss and both pay the write premium, because the entry is not available for reads until the first request has processed the prefix. For a predictable warm-up, send one request first, wait for it to return, then fan out. This matters most at deploy time, when a fleet of workers all start with the same cold prompt.
saying these in an interview costs you the question
- Thinks the five-minute TTL is absolute rather than refreshed on hits
- Assumes cache reads are free
- Uses the one-hour TTL on steady high-volume traffic
- Believes caching also discounts output tokens
- Marks one-shot prompts and pays write premiums for nothing