skip to content

How does the Langfuse SDK cache prompts fetched with get_prompt, and what is the default TTL?

level: middleimportance: must knowfreq 66%

answer

  1. In-process, not shared between replicas
  2. Measured in seconds, not minutes
  3. Stale entry is served, then refreshed
  4. Only the first call blocks
  5. Zero disables it entirely

basics

~20 s

The Langfuse SDK keeps a client-side, in-process cache of fetched prompts with a default TTL of 60 seconds. After expiry the next get_prompt call serves the cached copy and refreshes in the background, so only the very first fetch in a process blocks on the network.

solid answer

~50 s

Without caching, every request would pay a round trip to Langfuse just to read a prompt, putting a third-party service on your critical path. So the SDK caches each fetched prompt in process, keyed by prompt name plus the version or label you asked for, with a default TTL of **60 seconds**. When the entry expires, the next `get_prompt` call still returns the cached copy immediately and triggers a background refresh, so the expiry itself is not a latency spike — only the first fetch after process start blocks. You override the TTL per call with `cache_ttl_seconds=`, and `cache_ttl_seconds=0` disables caching so every call hits the API — appropriate in a short-lived script or when demoing a change, not on a hot request path. Because the cache is per process, a label move propagates independently to each replica within roughly one TTL, so a promotion rolls across a fleet rather than flipping atomically.

code

python · 12 lines
python
from langfuse import Langfuse

langfuse = Langfuse()

# Default: cached for 60 seconds, refreshed in the background on expiry
prompt = langfuse.get_prompt("support-reply")

# Longer TTL on a very hot path: fewer reads, slower propagation
hot = langfuse.get_prompt("support-reply", cache_ttl_seconds=300)

# No caching at all: a network round trip on every call
fresh = langfuse.get_prompt("support-reply", cache_ttl_seconds=0)

go deeper

for a junior

Know that the SDK caches fetched prompts locally rather than calling the API every time, that the default TTL is 60 seconds, and that cache_ttl_seconds lets you change it.

for a middle

Explain the mechanics: the cache is per process and keyed by name plus label or version, expiry serves the stale copy and refreshes in the background, and only the first fetch in a process blocks.

for a senior

Show that you have operated this — a label move rolls out over roughly a TTL per replica so both versions serve traffic during the window, rollback has the same lag, and startup warm-up removes the cold-start blocking fetch.

for a principal

Own the tradeoff: the TTL is simultaneously a request-volume dial, a propagation-latency dial and a resilience dial. Decide it per surface alongside the fallback strategy, and set the expectation that prompt promotion is a gradual rollout, not an atomic switch.

## The problem the cache solves Server-side prompts turn a string literal into a network read. If every request called Langfuse for the prompt text, you would have added a third-party dependency to the hot path of your application, plus tens of milliseconds of latency per call and a large volume of read traffic. The SDK's client-side cache exists so that in the steady state, fetching a prompt is a dictionary lookup. ## How it behaves - **Scope**: in-process and in-memory. Each Python process holds its own cache; nothing is shared between replicas, and the cache dies with the process. - **Key**: the prompt name together with what you asked for — a specific `version=`, a specific `label=`, or the default `production` resolution. Fetching the same prompt under two labels caches two entries. - **Default TTL**: 60 seconds. - **Expiry behaviour**: expiry does not evict. When a cached entry is stale, `get_prompt` returns the stale value immediately and kicks off a background refresh, so the caller does not block. The refreshed value is served by subsequent calls. - **Cold start**: the first `get_prompt` for a given key in a process has nothing to serve, so it does block on the network. That design is deliberately biased toward availability and latency over freshness: a slightly old prompt is almost always better than a slow or failed request. ## Tuning the TTL `cache_ttl_seconds` is a per-call argument: - **Default (60s)** — fine for most services. A promotion is live everywhere inside about a minute. - **Longer (e.g. 300s)** — fewer reads, more resilience to a Langfuse blip, slower propagation. Reasonable for a stable prompt on a very high-QPS path. - **Shorter** — faster propagation at the cost of more requests; rarely worth it, since the background refresh already keeps the blocking cost at zero. - **`cache_ttl_seconds=0`** — caching off, network round trip on every call. Correct for a one-shot script, a notebook, or while iterating on a prompt and wanting each run to see the newest text. Wrong on a request path: you have re-introduced the dependency and the latency you were avoiding. ## Propagation is per-process, not fleet-wide This is the consequence engineers most often miss. Moving the `production` label is a single write on the Langfuse side, but each of your replicas only notices when its own entry expires and refreshes. With the default TTL and, say, 20 replicas, the fleet converges within roughly a minute — but during that minute **both versions are serving traffic**. Practical implications: - A promotion is a gradual rollout with an uncontrolled schedule, not an atomic switch. That is usually harmless and occasionally not — for instance if the prompt change is paired with a code change that expects a new variable. - A rollback has the same lag. In an incident, "I moved the label back" does not mean the bad prompt has stopped serving; it means it will stop within about a TTL. - To reason about which version served a given request, you need the version recorded on the generation itself. The label tells you the current pointer, not what was used a minute ago. ## Warm-up and long-lived processes Because the cost lands on the first call, a natural pattern is to fetch the prompts a service uses at startup, so the first real user request finds a warm cache. This also surfaces a missing prompt or a missing `production` label at boot rather than on the first customer request. In serverless environments where processes are short-lived and frequently cold, expect the blocking first fetch far more often — a case for a longer TTL, a warm-up in the init phase, and definitely for supplying a fallback. ## Interaction with failure The cache is also the first line of defence when Langfuse is unreachable: a process that already holds a copy keeps serving it rather than failing. The TTL you pick therefore doubles as a freshness-versus-resilience dial, not just a request-volume dial.

  • You move the production label and a colleague says the change is live. Is it?
    It is live on the Langfuse side, but not necessarily in your fleet. Each process serves its cached copy until that entry expires — roughly one TTL, 60 seconds by default — and refreshes in the background. So for about a minute both the old and new versions are serving real traffic, and the same lag applies to a rollback.
  • When is cache_ttl_seconds=0 the right setting?
    In short-lived contexts where freshness beats latency and volume is trivial: a one-off script, a notebook, a manual verification run right after promoting a version, or an offline batch job that must pin to the newest text. On a live request path it puts a network call and a third-party dependency back into every request, which is precisely what the cache exists to remove.
  • Why does an expired cache entry not cause a latency spike?
    Because expiry does not evict. On the first call after the TTI lapses, the SDK returns the stale cached prompt immediately and refreshes it in the background, so the caller never waits on the network. Only a key that has never been fetched in this process forces a blocking request.
  • Two services fetch the same prompt under different labels. How many cache entries is that?
    Two per process, because the cache key includes what you asked for — the name plus the specific label or version. Fetching "support-reply" under production and under staging caches them separately, and each expires and refreshes on its own schedule.

saying these in an interview costs you the question

  • Assumes prompts are fetched fresh on every call
  • Thinks the cache is shared across replicas or processes
  • Believes cache expiry blocks the next request
  • Expects a label move to reach every replica instantly
  • Sets cache_ttl_seconds=0 on a hot request path

context