In CrewAI, what does setting memory=True on a Crew actually enable?
answer
- not one store, several
- two are vector-backed, one is not
- entities get their own store
- past task grades feed later runs
- injected into the prompt, not fetched by a tool
basics
~20 sThree stores switch on at once: short-term memory of recent interactions, entity memory about people and things mentioned, and long-term memory of past task results and their evaluations. CrewAI retrieves from all three and injects the result into each task's prompt.
solid answer
~40 s`Crew(memory=True)` turns on CrewAI's built-in memory stack rather than a single cache. Short-term memory and entity memory are vector-backed: recent interactions and extracted entities are embedded and stored so semantically similar context can be pulled back. Long-term memory is a SQLite table of past task executions along with a quality evaluation and improvement suggestions, keyed so a later run of a similar task can learn from it. Before each task, CrewAI queries all of them and prepends the hits to the agent's prompt as extra context. The costs are real: embedding calls on every write and read, extra prompt tokens, and an additional LLM call to evaluate task output for long-term memory. It also persists to disk, so state carries across process restarts until you reset it.
go deeper
Recall that memory=True is one flag that switches on several kinds of remembering, and that the crew then carries context between tasks instead of treating each one as isolated.
Name short-term, entity and long-term memory, say which are vector-backed and which is SQLite, and explain that retrieved context is prepended to each task's prompt automatically.
Quantify the cost — embedding calls per read and write, extra prompt tokens, an additional evaluation LLM call — and be ready to say when you would leave memory off entirely for a single-shot crew.
Own the policy: which crews are allowed to accumulate state, where that state lives, how it is reset between runs or tenants, and how you keep an automatic prompt-injection mechanism from making incidents unreproducible.
## One flag, three stores `memory=True` on a `Crew` is a convenience switch that enables several distinct mechanisms. Candidates who answer "it remembers the conversation" miss the point of the question — the interviewer wants the decomposition. **Short-term memory.** Interactions from the current execution are embedded and written to a vector store. When a later task runs, its description is embedded and semantically similar recent material comes back. It is "short-term" in intent — the recent working context of the crew — not in storage lifetime: it is written to the same on-disk store as the rest, so it survives a process restart until you reset it. That surprises people who expect a fresh crew object to start clean. **Entity memory.** As tasks run, CrewAI extracts entities — people, companies, products, concepts — with short descriptions and relationships, and stores them in vector storage too. The purpose is that when "Acme Corp" comes up again three tasks later, what the crew already established about Acme comes back with it, rather than the agent re-deriving it. **Long-term memory.** This one is not a vector store. Completed task executions are recorded in a SQLite database along with a quality score and improvement suggestions produced by evaluating the output. On a later run, when a task with a matching description executes, those past results and suggestions are surfaced so the crew can do better than last time. This is the piece that makes memory cross-run rather than merely cross-task. ## How the retrieved material reaches the model CrewAI assembles context from these stores before a task executes and prepends it to the prompt the agent sees. The agent does not call a "recall" tool; retrieval is automatic and invisible in the task definition. That is convenient and also the source of most confusion — when an agent behaves oddly, people look at the task description and never consider that injected memory is steering it. ## What it costs This is where mid-level answers separate from junior ones. - **Embedding calls.** Every write and every read of short-term and entity memory hits the embedding model. Default configuration uses OpenAI embeddings, so enabling memory silently introduces an API-key dependency and per-run spend even if your LLM is elsewhere. - **Prompt tokens.** Retrieved context is added to every task prompt. On long crews this measurably grows input cost and can crowd out the task's own instructions. - **An extra LLM call.** Producing the long-term memory evaluation (quality score and suggestions) means calling a model on task output, adding latency at the end of each task. - **Latency and variance.** Retrieval before each task adds round trips, and the injected content changes between runs, which makes reproducing a bad output harder. ## Configuring it The embedder is set with `Crew(embedder={"provider": "openai", "config": {"model": "text-embedding-3-small"}})`, which lets you point memory at a local or non-OpenAI embedding provider. State lives under the CrewAI storage directory, overridable with the `CREWAI_STORAGE_DIR` environment variable. You clear it with `crew.reset_memories(command_type="all")` — or narrower values for the individual stores — and from the CLI with `crewai reset-memories`. There is also an external route: `memory_config` lets you plug in a managed memory provider (Mem0, for example) with a `user_id`, which is how you get memory that is scoped to a person rather than to a storage directory. ## When to leave it off A single-shot crew that reads its inputs, produces a report and exits gains nothing from memory except cost and nondeterminism. Memory earns its keep when the same crew runs repeatedly over a related stream of work, when entity continuity across many tasks matters, or when you genuinely want later runs to be informed by how earlier ones were graded. "On by default because it sounds good" is the anti-pattern. ## Interview framing Name the three stores and say which are vector-backed and which is SQLite. Say that retrieval is automatic and prompt-injected. Then volunteer the cost profile — embedding calls, extra tokens, the evaluation call — and the fact that state persists across runs until reset. That combination reads as someone who has actually run a crew with memory on and watched the bill.
- Which of those stores survives a process restart?All of them. Long-term memory is a SQLite file and the vector-backed stores are written to the CrewAI storage directory on disk, so a brand-new crew object pointed at the same storage directory sees everything the previous process wrote. "Short-term" describes what the store is for, not how long it lives; only an explicit reset or a different storage location clears it.
- What extra model calls does memory=True introduce per task?At minimum an embedding call to write the interaction and one to embed the retrieval query for the vector-backed stores, plus an additional LLM call that evaluates the task output to produce the quality score and suggestions stored in long-term memory. On a ten-task crew that is a meaningful addition to both latency and spend.
- How would you tell whether memory is helping or hurting a crew's output?Run the same inputs with memory on and off and compare output quality, token counts and wall-clock time. Turn on verbose output to see what context is being injected before each task. If retrieved material is off-topic or contradicts the current task, memory is adding noise; that is usually a signal to disable it or narrow what the crew stores.
saying these in an interview costs you the question
- Says memory=True just keeps the chat history
- Assumes short-term memory disappears when the process exits
- Thinks the agent must call a tool to recall memory
- Ignores that memory adds embedding and evaluation calls
- Believes it works without an embedding provider configured