What does Koog's Persistency feature checkpoint, and what does a rollback not undo?
answer
- Snapshots run state, not the world
- Storage provider decides durability
- Rollback rewinds context, not consequences
- In-memory checkpoints die with the process
- Idempotent tools, compensating actions
basics
~20 sInstalling Koog's Persistency feature snapshots an agent run's execution state — its message history and where execution stood — into a storage provider, so a run can resume or roll back. It restores agent state only: side effects tools already performed stay done.
solid answer
~50 s`install(Persistency)` on a Koog agent, configured with a storage provider, records checkpoints of a run: the conversation the agent has accumulated plus the point in the strategy at which it stood, so the run can be resumed after a crash or rewound to an earlier decision. You can create checkpoints explicitly at chosen points, or enable automatic checkpointing so one is taken as execution proceeds. Two limits decide whether it helps in production. First, durability is the storage provider's job — in-memory checkpoints die with the process, so surviving a restart means a file- or database-backed provider, and the state you keep must be serialisable. Second, rollback is *not* transactional: restoring agent state cannot un-send an email, un-charge a card or un-write a row that a tool already executed, and tokens already spent stay spent. Design side-effecting tools to be idempotent, or pair them with compensating actions.
go deeper
Know that the feature exists to save an agent run's progress, that it is installed on the agent, and that where checkpoints are stored is a configuration choice.
Explain what a checkpoint captures — accumulated messages plus the point execution reached — and why an in-memory provider cannot survive a process restart.
Lead with the limit that matters: rollback restores agent state, not the effects tools already caused, and show how idempotent tools and compensating actions close that gap.
Own the policy questions: checkpoint retention and encryption for conversation data, resumability across strategy versions and deploys, and preventing concurrent resumption of the same run.
## What the feature is for Agent runs are long, non-deterministic and expensive. A run that dies at step nine of twelve has burned real money and real wall-clock time, and re-running it from scratch burns them again — and may not reproduce, because the model is sampling. Koog's Persistency feature exists to make a run's progress durable so it can be resumed, and rewindable so a bad branch can be abandoned without discarding everything before it. You turn it on by installing the feature on the agent and giving it a storage provider. From there, checkpoints can be taken at points you choose, or automatically as the run advances if you enable that. ## What a checkpoint holds Conceptually, a checkpoint captures the run's execution state: the messages the agent has accumulated so far, and where in the strategy execution stood, including the input that step was working on. That combination is what lets execution be re-entered rather than restarted — the agent resumes with the same context it had, at the same place. It does not capture the world. Nothing outside the agent process — your database, a third-party API, a queue, a mailbox — is part of the snapshot. ## Durability is the storage provider's decision The feature separates *taking* a checkpoint from *keeping* it. An in-memory provider is fine for tests and for rewind-within-a-run, and useless for crash recovery: when the JVM exits, the checkpoints exit with it. Surviving a restart or a redeploy means a provider that writes somewhere durable — a file, or your own implementation over a database. Two consequences follow. The state you checkpoint has to be serialisable, which constrains what you may park in it. And checkpoint writes are I/O on the hot path, so automatic per-step checkpointing on a chatty strategy is a real cost in latency and storage churn; checkpoint where the value is, typically before an expensive or irreversible step. ## The rollback illusion The most important thing to say in an interview is that rollback is not a database transaction. Restoring an earlier checkpoint rewinds the *agent's* state. It does not rewind anything a tool already did. If the agent called a tool that emailed a customer, refunded an order, posted to Slack or deleted a file, that happened. Rewinding to a point before the tool call means the agent may well decide to call it again — and now the customer has two emails. The same asymmetry applies to money and rate limits: tokens spent before the checkpoint you rewound to are spent, and provider rate-limit budget consumed is consumed. Persistency buys you recovery of context, not of consequences. ## Designing around it Three patterns handle this. Make side-effecting tools *idempotent*, keyed by something stable — an order id, a request id you generate before the call — so a replayed call is a no-op rather than a duplicate. Where idempotency is impossible, define a *compensating action* (refund for charge, retraction for publish) and make the recovery path invoke it rather than pretending the effect never happened. And *checkpoint immediately before* irreversible steps, so a resume re-enters at the effect rather than replaying a long tail of work that precedes it. Also think about ordering: a tool that both mutates the world and returns data the agent needs should, where possible, record its own outcome durably so that on resume the agent can read the outcome instead of repeating the mutation. ## Persistency versus memory A frequent confusion is between checkpointing and agent memory. Checkpointing is about one run's execution state — mechanical, short-lived, about resumption. Koog's memory feature is a different concern: durable facts extracted from runs, stored under concepts and scopes, and loaded back into later prompts so an agent recalls things across runs. Restoring a checkpoint replays a run; memory changes what a *new* run knows. In production you often install both, for different reasons. ## Operational questions to be ready for How long do checkpoints live, and who deletes them? They contain conversation content, which frequently contains personal data, so retention and encryption are governance questions, not implementation details. What happens if the strategy changes between checkpoint and resume — a checkpoint taken by yesterday's graph may reference a step that no longer exists, so treat resumability across deploys as a versioning problem. And can two workers resume the same checkpoint concurrently? If your storage does not prevent it, you have just double-executed the tail of a run.
- How does the Persistency feature differ from Koog's AgentMemory feature?Persistency checkpoints one run's execution state so it can resume or rewind; it is mechanical and run-scoped. AgentMemory is about knowledge: facts captured under concepts and scopes, stored durably, and loaded into later prompts so a future run starts knowing something. Restoring a checkpoint replays a run; memory changes what a new run knows. Production systems often install both for different reasons.
- Why can automatic checkpointing at every step be a bad default?Every checkpoint is serialisation plus a write on the execution path, so a chatty strategy pays latency and storage churn for snapshots you will never restore. It also multiplies the surface holding conversation content, which is usually personal data with retention obligations. Checkpoint where recovery has value — before expensive or irreversible steps — rather than uniformly.
- A checkpoint was written by yesterday's strategy and your deploy changed the graph. What now?Treat resumability across deploys as a versioning problem. A checkpoint references the point in a strategy that no longer exists, so resuming can fail or land somewhere semantically different. Options are stamping checkpoints with a strategy version and refusing to resume mismatches, draining in-flight runs before rollout, or keeping the old graph available until outstanding checkpoints expire.
- What stops two workers from resuming the same checkpoint at once?Nothing in the snapshot itself — you need it at the storage layer. Without a lease, a lock, or a claim marker on the checkpoint record, two consumers can both resume and both re-execute the remaining side-effecting steps. This is the same at-least-once problem as any job queue, and it is another argument for idempotent tools keyed on a stable identifier.
saying these in an interview costs you the question
- Calls rollback transactional, as if side effects unwind
- Uses in-memory checkpoints and expects restart recovery
- Confuses run checkpoints with long-term agent memory
- Checkpoints every step by default without weighing cost
- Ignores that checkpoints hold conversation content subject to retention rules