In CrewAI, how does Knowledge differ from crew memory at run time?
answer
- one you author, one the run writes
- read-only versus cumulative
- indexed at kickoff versus after each task
- only one of them makes reruns differ
- reset targets them separately
basics
~20 sKnowledge is grounding you attach up front and index at kickoff; it is read-only during the run and unchanged by it. Memory is written by the run itself — recent interactions, entities, past task evaluations — and grows every time the crew executes.
solid answer
~40 sThey are two separate retrieval surfaces with opposite write paths. **Knowledge** comes from sources you declare — `StringKnowledgeSource`, `PDFKnowledgeSource`, `CSVKnowledgeSource`, `CrewDoclingSource` and friends — passed to `Crew(knowledge_sources=[...])` or `Agent(knowledge_sources=[...])`. It is chunked, embedded and stored when the crew starts, and agents only ever read from it; nothing the crew does modifies it. **Memory**, enabled with `Crew(memory=True)`, is written by execution: short-term interactions and extracted entities go to vector storage, and completed task results plus their evaluations go to a SQLite long-term store. Both get retrieved and injected into the task prompt, which is why they feel similar in the prompt but behave completely differently operationally — knowledge changes when you change your documents, memory changes because the crew ran.
go deeper
Remember the direction of writes: you supply knowledge before the run, while memory is filled by the run itself. Both end up as extra context in the agent's prompt.
Explain that knowledge is chunked and embedded at kickoff and never modified by execution, while memory writes short-term, entity and long-term records as tasks complete, and say what that means for reproducibility.
Bring the operational angle: unbounded memory growth, reruns that differ because of injected state, reset granularity that spares an expensive knowledge index, and the per-run versus front-loaded cost profiles.
Own the classification policy for a platform — what is curated corpus, what is accumulated state, what must be a live tool — and the retention and reset rules that follow, since memory quietly turns a stateless crew into a stateful system.
## Two retrieval surfaces, one prompt By the time text reaches the model, knowledge chunks and memory hits look alike: both arrive as extra context prepended to the task prompt. The difference is everything around that moment — who writes them, when, and what invalidates them. ## Knowledge: authored, indexed once, read-only Knowledge is the corpus you choose. You declare sources and hand them to the crew or to a specific agent. At kickoff CrewAI chunks each source (`chunk_size` / `chunk_overlap` on the source), embeds the chunks with the configured embedder, and stores them. During execution, agent queries are matched against that store and the top chunks come back. Characteristics that follow: - **Deterministic content.** Two runs over the same documents retrieve from the same corpus. If output changes, it is not because the knowledge base drifted. - **Read-only during the run.** Nothing an agent does adds to knowledge. To change it you change the source documents and re-index. - **Scopable by role.** Agent-level `knowledge_sources` keep a corpus visible to one agent only. - **Cost is front-loaded.** You pay embedding for the whole document set once; per-run cost is just query embedding and prompt tokens. ## Memory: written by execution, cumulative Memory is the crew's own trail. With `memory=True`, short-term memory captures recent interactions in vector storage, entity memory captures the people and things that came up, and long-term memory records finished task executions in SQLite together with a quality evaluation and suggestions for next time. Characteristics that follow: - **Nondeterministic across runs.** The second run of an identical crew sees material the first run created. This is a feature when you want accumulation and a menace when you are debugging. - **Grows without bound.** Nothing prunes it for you; the store keeps filling until you reset it. - **Carries operational risk.** Whatever was in the run — including data belonging to whoever the run served — is now on disk and retrievable by any later run pointed at the same storage. - **Cost is per-run.** Embedding writes and reads plus an evaluation LLM call, every task, every time. ## Where each lives Both end up under the CrewAI storage directory (overridable with `CREWAI_STORAGE_DIR`): vector-store directories for knowledge, short-term and entity data, and a SQLite file for long-term memory. `crew.reset_memories(command_type="all")` clears the stores, and narrower command values target individual ones — which is exactly why the distinction matters operationally: you often want to wipe accumulated memory while leaving an expensively-indexed knowledge base alone. ## Choosing between them Ask what the material is: - A product manual, a style guide, a schema description, a set of past support tickets you curated — **knowledge**. It is authored, stable, and you want it identical on every run. - What this crew learned about the customer it is currently helping, which tasks scored badly last time, which entities have already been discussed — **memory**. - Something that must be current at the moment of use — a live price, a database row — **neither**; that is a tool. A frequent design error is stuffing changing operational data into knowledge sources and re-indexing constantly, or the reverse: relying on memory to hold reference material that should simply be a document. Memory is a lossy, semantically-retrieved, ever-growing store — it is a poor filing cabinet. ## Debugging implications When a crew produces a surprising answer, the diagnostic split is: did a knowledge chunk with wrong or outdated content get retrieved, or did memory inject something from a previous run? The first is fixed by editing documents and re-indexing; the second by resetting memory and asking whether it should be on at all. Being able to state that split cleanly is what a mid-level interviewer is listening for. ## Interview framing Lead with the write path — knowledge is authored and read-only, memory is produced by execution — then note that both are retrieved into the prompt the same way, then close with the operational consequences: determinism, growth, reset granularity, and cost profile.
- A crew gives different answers on two identical runs. Which surface do you suspect first?Memory. Knowledge is read-only during the run, so an unchanged document set retrieves the same corpus every time. Memory, by contrast, was written by the first run and is visible to the second. Confirm by resetting memory and re-running the same inputs; if the difference disappears, injected memory was the cause.
- Would you put a product manual in memory or in knowledge, and why?Knowledge. The manual is authored, stable reference material, and you want every run to see exactly the same content. Memory is a semantic, lossy, ever-growing store shaped by execution, so relying on it to hold reference text means the manual is only present if some earlier run happened to mention the relevant part.
- Can you clear accumulated memory without re-indexing an expensive knowledge base?Yes — reset is granular. crew.reset_memories takes a command type, so you can target short-term, long-term or entity memory individually rather than wiping everything. That matters when the knowledge base cost real embedding spend to build and you only want to discard what previous runs accumulated.
saying these in an interview costs you the question
- Calls knowledge "just another kind of memory"
- Expects agents to write new facts into knowledge sources
- Thinks memory reloads reference documents for you
- Assumes both are cleared by the same single command
- Uses knowledge for data that must be current at call time