skip to content

How do you resume a long-running LlamaIndex Workflow after the process restarts?

level: seniorimportance: should knowfreq 44%

answer

  1. Two mechanisms, one durable
  2. Serialize the run's shared object
  3. from_dict rebuilds it against the workflow
  4. Step snapshots live in memory
  5. Store ids, never live clients

basics

~20 s

Serialize the run's Context with ctx.to_dict() and store it, then rebuild it with Context.from_dict(workflow, data) and pass it to the next run. WorkflowCheckpointer captures per-step snapshots, but they live in memory unless you persist them yourself.

solid answer

~50 s

There are two levels. For continuity across turns or restarts, serialize the `Context`: `ctx_dict = ctx.to_dict(serializer=JsonSerializer())`, write it to your own store keyed by session, then restore with `Context.from_dict(workflow, ctx_dict, serializer=JsonSerializer())` and pass that context into the next `run()`. Everything a step wrote through `ctx.store` comes back, including agent chat history. For step-level replay during development, wrap the workflow in a `WorkflowCheckpointer` and run through it. It records a checkpoint after each completed step — the input event, the output event and the context state — exposed on `.checkpoints`, and you can restart from any of them instead of paying for the earlier steps again. The caveat that matters in production: those checkpoints are held in memory by the checkpointer object, so a process restart loses them unless you persist the snapshots yourself. Also, only JSON-serializable state survives a plain JSON serializer — live clients and connections do not.

code

python · 26 lines
python
import asyncio
import json

from llama_index.core.workflow import Context, JsonSerializer, StartEvent, StopEvent, Workflow, step


class Counter(Workflow):
    @step
    async def bump(self, ctx: Context, ev: StartEvent) -> StopEvent:
        n = await ctx.store.get("n", default=0)
        await ctx.store.set("n", n + 1)
        return StopEvent(result=n + 1)


async def main() -> None:
    wf = Counter()
    ctx = Context(wf)
    await wf.run(ctx=ctx)

    blob = json.dumps(ctx.to_dict(serializer=JsonSerializer()))

    restored = Context.from_dict(wf, json.loads(blob), serializer=JsonSerializer())
    print(await wf.run(ctx=restored))


asyncio.run(main())

go deeper

for a junior

Know that a run's shared state lives in a Context object and that it can be serialized to a dict and rebuilt, which is what gives an agent memory across turns.

for a middle

Explain the round trip concretely — ctx.to_dict(serializer=...), your own store, Context.from_dict(workflow, data, serializer=...), then pass it into run() — and what a JSON serializer will and will not carry.

for a senior

Demonstrate operational judgment: checkpoints are in-memory replay aids, durability is your storage choice, resumed steps re-execute so side effects need idempotency keys, and saved state needs a version.

for a principal

Own the state architecture: decide what belongs in workflow context versus an external system of record, how long-lived approvals are represented, and how state schema changes roll out without stranding in-flight runs.

## Two different problems "Resume" means two things in LlamaIndex workflows, and interviews probe whether you separate them. 1. **Continuity of a conversation or a job across process boundaries.** The state you care about is what the steps accumulated: chat history, retrieved facts, counters, partial results. That lives in `Context`. 2. **Replaying a failed run from the step that failed** so you do not repeat expensive earlier steps. That is `WorkflowCheckpointer`. ## Context serialization A `Context` is the per-run shared object. Steps write to it with `await ctx.store.set(key, value)` and read with `await ctx.store.get(key)`. To carry it across runs, serialize: ``` ctx_dict = ctx.to_dict(serializer=JsonSerializer()) ``` Store that dict wherever your application keeps session state — Redis, Postgres, a document store; the framework does not choose for you. To resume: ``` ctx = Context.from_dict(workflow, ctx_dict, serializer=JsonSerializer()) result = await workflow.run(..., ctx=ctx) ``` Because the prebuilt agents are Workflows, the same pattern gives multi-turn agent memory: keep one `Context` per conversation, pass it into every `agent.run(...)`, and chat history plus any tool-written state persists. Serializer choice matters. `JsonSerializer` is safe and portable but only handles JSON-compatible values. A pickle-based serializer can round-trip richer objects at the usual cost — you are executing whatever is in the payload on load, so never restore untrusted bytes, and stored state becomes coupled to your class definitions. The disciplined answer is to keep only plain data in the store: strings, numbers, dicts, lists, and identifiers you can use to rebuild clients on the other side. ## WorkflowCheckpointer `WorkflowCheckpointer(workflow=my_workflow)` wraps a workflow. Run through the wrapper rather than the workflow directly, and after each completed step it records a checkpoint holding the last completed step's name, the input event, the output event and the serialized context state. The collected checkpoints are exposed on the checkpointer, grouped by run, and it can start a fresh run *from* a chosen checkpoint. This is invaluable when a workflow's third step is a slow, expensive LLM call and the fourth is the one you are iterating on — you replay from checkpoint three repeatedly instead of paying for the first three every edit. It is also how you triage a bad output: inspect the exact input event that the failing step received. The limitation to state plainly: checkpoints are accumulated by the checkpointer instance in memory. They are not a durable store, they grow with every step of every run, and they vanish with the process. Treating `WorkflowCheckpointer` as production fault tolerance is the mistake; treat it as a development and replay tool, and if you want durability, extract the checkpoint payloads and write them somewhere yourself. ## Human-in-the-loop is the same machinery A workflow that must wait for a person is a resume problem with a long gap. A step can emit an input-required event and wait for a human-response event; the run handler surfaces the request through the stream, and the answer is sent back into the context when it arrives. If the wait may outlive the process — an approval that takes hours — you serialize the context, drop the process, and rebuild the context when the human responds. That is why context serialization, not an in-memory pause, is the durable pattern. ## Design rules that keep this working - **Keep state small and plain.** Every megabyte in `ctx.store` is a megabyte serialized on each save. Store document ids, not documents. - **Never put live handles in the store.** Database sessions, HTTP clients and file objects do not serialize; rebuild them in the step from configuration. - **Version your state.** Restoring a context written by an older code path into a workflow whose steps now expect different keys fails at read time, in a step, not at load. Include a schema version and migrate or discard. - **Make steps idempotent where you can.** Resuming re-executes the step you resume into; if it charges a card or posts a message, guard it with a key you keep in the context. - **Watch the timeout.** The workflow timeout applies per run, so a resumed run gets a fresh budget — do not rely on it to bound total task time across resumes. ## What to say in an interview Name the two mechanisms, say which one is durable, and be explicit that persistence backend choice is yours: LlamaIndex gives you a serializable context and step snapshots, not a hosted state service. That distinction is what separates someone who has shipped a long-running workflow from someone who has read the quickstart.

  • Why is WorkflowCheckpointer not enough for production fault tolerance?
    Its checkpoints are accumulated by the checkpointer instance in memory, so they disappear when the process dies — exactly the event you wanted tolerance for. They also grow unboundedly across runs. Use it for development replay and post-hoc inspection of what a failing step received; for durability, serialize the context (or extract the checkpoint payloads) into your own store and treat that store as the source of truth.
  • What breaks when someone stores a database session or an HTTP client in ctx.store?
    Serialization fails, or worse, succeeds through a pickle-style serializer and restores a dead handle pointing at a closed socket. The fix is a discipline, not a serializer: keep only plain data in the context — ids, keys, small structures — and reconstruct clients inside the step from configuration. That also keeps the serialized payload small, which matters because it is written on every save.
  • How does human-in-the-loop approval fit this picture when a person may take hours to respond?
    It is the same resume problem stretched out. A step signals that input is required, the handler surfaces that through the run's event stream, and the response is delivered back into the context. If the wait can outlive the process, do not hold a paused run in memory: serialize the context, let the process go, and rebuild the context when the answer arrives. Guard the resumed step so a duplicate approval cannot double-execute a side effect.

saying these in an interview costs you the question

  • Treats in-memory checkpoints as durable persistence
  • Expects LlamaIndex to supply the storage backend
  • Stores live clients or sessions in workflow state
  • Assumes resuming skips the step it resumes into
  • Ignores schema drift between saved state and current steps

context