What happens to langfuse.get_prompt() when the Langfuse API is unreachable?
answer
- Two cases, warm and cold
- Stale beats failing
- One argument saves the cold start
- The rescue value has no version number
- Bound the fetch, count the hits
basics
~20 sA process that already holds a cached copy keeps serving it, even past the TTL, so a warm process rides out the outage. A process with a cold cache fails unless you passed fallback=, which returns a prompt client built from the literal you supplied.
solid answer
~50 sThere are two very different cases. **Warm cache**: the SDK keeps serving the cached prompt rather than failing when the background refresh cannot reach the API, so a long-running process is largely immune to a Langfuse outage. **Cold cache**: a process that has never fetched this prompt has nothing to serve, and the call fails — which is why the `fallback` argument exists. `langfuse.get_prompt("support-reply", fallback="Answer the question: {{question}}")` returns a prompt client built from that literal, which you can `compile()` exactly like a real one; `prompt.is_fallback` tells you it happened. The catch: a fallback has no server-side version, so generations produced from it cannot be attributed to a prompt version in the metrics. You also bound the blocking call with `max_retries` and `fetch_timeout_seconds`. The rule of thumb is: every prompt on a request path gets a fallback, and services warm their prompts at startup.
code
python · 16 linesfrom langfuse import Langfuse
langfuse = Langfuse()
prompt = langfuse.get_prompt(
"support-reply",
fallback="Answer the customer question politely: {{question}}",
max_retries=2,
fetch_timeout_seconds=3,
)
if prompt.is_fallback:
# degraded mode: no server-side version to attribute generations to
metrics.increment("prompt.fallback_used", tags={"name": "support-reply"})
text = prompt.compile(question="Where is my order?")go deeper
Know that fetched prompts are cached so a brief outage usually goes unnoticed, and that get_prompt takes a fallback argument giving the SDK a literal prompt to use when it cannot reach the API.
Distinguish the warm case, where a stale cached prompt keeps being served, from the cold case, where nothing is cached and the call fails without a fallback. Mention is_fallback and the compile-compatible client it returns.
Show that you have run this: warm prompts at startup, always pass a fallback on the request path, bound the fetch with a timeout, emit a metric on fallback use, and know that fallback traffic loses version attribution.
Own the dependency decision — what degraded mode the business accepts when a third-party prompt store is down, who keeps fallbacks current, and whether the most critical surfaces should pin versions and ship their prompt with the build instead.
## Why this question is the real one Moving prompts to a server buys deploy-free iteration and costs you a runtime dependency. The interview question is always the same: what happens to your application when that dependency is down? A candidate who answers "the request fails" has not read the resilience design; a candidate who answers "nothing, it's cached" has not thought about cold starts. ## Case 1: the cache is warm If the process has fetched this prompt before, it holds a copy. When the TTL lapses and the background refresh cannot reach the API, the SDK does not throw away the entry — it keeps serving the stale prompt. In practice a long-running service that has been up for a while barely notices a Langfuse outage: prompts keep resolving from memory, and the only loss is that a promotion made during the outage will not land until connectivity returns. This is the reason the cache TTL is also a resilience dial, not just a freshness dial. A longer TTL means fewer refresh attempts, but the availability property comes from serve-stale-on-failure rather than from the TTL itself. ## Case 2: the cache is cold A freshly started process, an autoscaled-in replica, a serverless invocation on a new container — none of these have a cached copy. The fetch is the first thing that happens and it fails, so `get_prompt` raises. This is the dangerous case, because outages and deploys correlate: you restart to mitigate something, and every new pod comes up unable to resolve its prompts. ## The fallback argument `fallback=` is the answer to the cold case: ``` prompt = langfuse.get_prompt( "support-reply", fallback="Answer the customer question politely: {{question}}", ) ``` For a chat prompt the fallback is a list of role/content dicts instead of a string. The SDK returns a prompt client built from that literal, which supports `compile()` like any other, so the calling code is unchanged. `prompt.is_fallback` is `True` on such a client, which is what you branch on if you want to log loudly, emit a metric, or degrade the feature. What a fallback costs you: - **No version.** It is a local literal, not a stored version, so there is no version number to attach to the generations it produces. Traffic served during the outage is not attributable to a prompt version in the analytics. - **Drift.** The fallback is a string in your codebase and therefore does not update when the managed prompt does. If it is months out of date, your degraded mode is degraded in ways nobody has evaluated. Treat the fallback as code that needs periodic review — ideally the last-promoted text, refreshed when you deploy. ## Bounding the blocking call Even with a fallback in place, a cold fetch against a slow-but-not-dead API can hang. `get_prompt` accepts `max_retries` and `fetch_timeout_seconds` so the fetch fails fast into the fallback instead of stalling the request. Partial failure — high latency rather than a clean error — is the mode that actually hurts, so bounding it matters more than the retry count. ## Operational recipe 1. **Warm at startup.** Fetch every prompt the service uses during initialisation. This turns "first user request pays the cold fetch" into "boot pays it", and it surfaces a missing prompt or missing `production` label at deploy time. 2. **Always pass a fallback on the request path.** No exceptions for prompts that are "obviously always there". 3. **Bound the fetch** with a timeout, so degradation is fast rather than a slow stall. 4. **Observe fallback use.** Count `is_fallback` hits as a metric; a spike is both an outage signal and a warning that your traces for that window are missing version attribution. 5. **Keep fallbacks fresh.** Sync them to the last promoted version as part of a deploy, or they silently become an untested alternative behaviour. 6. **Consider pinning for the most critical surface.** A `version=`-pinned prompt still needs the network on a cold start, but pairing pinning with a fallback whose text matches that exact version gives you a degraded mode you have actually evaluated. ## What not to claim The cache does not survive the process, so "we're fine, it's cached" is only true for processes that were already running. And a fallback is a resilience mechanism, not a deployment mechanism — you never ship a prompt change by editing the fallback.
- Why can't you attribute generations to a prompt version while the fallback is in use?Because the fallback is a literal in your code, not a stored version — there is no version number for Langfuse to record on the generation. So traffic served during that window shows up in traces but not in per-version prompt metrics, which is worth knowing before you compare version performance across a period that included an outage.
- Your fallback string is a year old. Why is that a problem?It is an untested alternative behaviour. The managed prompt has been iterated on and evaluated; the fallback has not, so your degraded mode may be materially worse — or may expect variables the code no longer passes. Sync fallbacks to the last promoted version as part of the deploy, and review them like any other code.
- Why does warming prompts at startup help beyond latency?It moves the failure earlier and makes it visible. A missing prompt name, a prompt with no production label, or bad credentials fails at boot — where a health check or a rollout gate can catch it — rather than on the first customer request in production.
- Is a longer cache TTL a substitute for a fallback?No. A longer TTL only helps processes that already hold a copy; the resilience actually comes from the SDK serving stale entries when a refresh fails. A brand-new process has nothing cached regardless of TTL, so the cold-start case needs a fallback.
saying these in an interview costs you the question
- Says caching alone makes the app immune to a Langfuse outage
- Ignores the cold-start case in new or autoscaled replicas
- Ships prompt changes by editing the fallback string
- Assumes fallback-served traffic still records a prompt version
- Leaves the fetch unbounded so a slow API stalls requests