Where does CrewAI write crew memory and knowledge embeddings on disk?
answer
- not your project folder
- a per-user app-data path by default
- one env var relocates all of it
- a SQLite file sits beside the vector dirs
- containers lose it on restart
basics
~20 sEverything lands in a CrewAI storage directory: vector-store folders for short-term, entity and knowledge data plus a SQLite file for long-term memory. The location defaults to a per-user application-data path and is overridden with the CREWAI_STORAGE_DIR environment variable.
solid answer
~40 sCrewAI keeps all persisted state in one storage directory. Inside it you get vector-store directories for short-term memory, entity memory and knowledge collections, and a SQLite database file for long-term memory holding past task executions and their evaluations. By default that directory is a per-user application-data path chosen by platform convention, not your project folder — which is why people are startled that a `git clean` does not reset a crew. Set `CREWAI_STORAGE_DIR` to relocate it, which is the standard move for containers (mount a volume), for CI (point at a scratch path so runs start clean), and for keeping one project's state out of another's. `crew.reset_memories(...)` and the `crewai reset-memories` CLI clear what is there; deleting the directory works too but throws away the knowledge index along with the memory.
go deeper
Know that CrewAI persists memory and knowledge to a directory on disk rather than only in RAM, and that CREWAI_STORAGE_DIR is the environment variable that says where.
Describe the layout — vector-store directories for short-term, entity and knowledge data plus a SQLite file for long-term memory — and explain why the default per-user path surprises people who expect project-local state.
Bring deployment reality: containers losing the store on restart and re-paying indexing cost, CI runs inheriting state, and the concurrency hazard of two processes sharing one directory.
Treat the directory as data at rest derived from production inputs: decide its lifecycle, isolation boundary and encryption, and set the platform rule for whether memory is filesystem-backed at all versus delegated to an external provider.
## One directory, several artifacts When a crew has `memory=True` or knowledge sources attached, CrewAI needs somewhere to persist embeddings and records. It uses a single storage root and lays out several artifacts inside it: - **Vector-store directories** for short-term memory, entity memory, and knowledge collections. These hold the embeddings and their metadata. - **A SQLite database file** for long-term memory — `long_term_memory_storage.db` — containing rows for completed task executions along with the quality score and suggestions recorded for them. The important structural fact is that memory and knowledge are separate collections in the same place, which is what makes granular reset possible. ## Where the root actually is By default CrewAI picks a per-user application-data location following platform convention (the usual OS-specific app-data path), not the working directory of your script. Two consequences follow, and interviewers like both: 1. **State outlives your project directory.** Wiping the repo, re-cloning, or `git clean -xdf` does not reset a crew. People chasing "my crew still remembers the old data" often look everywhere except the app-data path. 2. **State is shared by anything using the same default.** Multiple projects, or multiple runs of the same project under the same user, land in the same root unless you separate them. ## Overriding with CREWAI_STORAGE_DIR `CREWAI_STORAGE_DIR` sets the root explicitly. The situations where you should always set it: - **Containers.** The default path lives inside the container filesystem and vanishes on restart — memory silently resets, or worse, re-indexes knowledge on every cold start and burns embedding spend. Point it at a mounted volume if you want persistence, or accept ephemerality deliberately. - **CI.** Point it at a scratch directory per job so tests start from a known-empty state and do not inherit whatever the last run wrote. - **Project isolation.** Give each project (or each tenant-serving process) its own root so their memories cannot retrieve each other's content. - **Serverless.** Read-only or ephemeral filesystems will fail or lose data at the default location; set it to writable scratch space and design for the store not surviving. ## Concurrency Multiple processes writing the same storage root is asking for trouble: you have a SQLite file and file-backed vector collections being mutated concurrently. Treat the storage directory as owned by one process at a time. If you need to scale out, give each worker its own root, or move memory to an external provider configured through `memory_config` rather than sharing a directory across replicas. ## Clearing it Three levels of blunt instrument: - `crew.reset_memories(command_type="all")` in code, with narrower command values for individual stores. - `crewai reset-memories` from the CLI, with flags for what to clear — handy when you have no running crew object. - Deleting the directory. This works but is indiscriminate: it takes the knowledge index down with the memory, so the next kickoff pays the full embedding cost of re-indexing your documents. ## Operational hygiene Because the directory holds embeddings derived from whatever the crew processed, it is data at rest with the same sensitivity as the inputs. Anything confidential a crew handled with memory on is now sitting in a vector store on that machine. That belongs in your threat model: encrypt the volume if the inputs warrant it, keep the path off shared machines, and reset it between tenants rather than between convenience points. ## Interview framing A strong answer names the two artifact kinds (vector directories plus a SQLite long-term file), says the default root is a per-user app-data path rather than the project directory, names `CREWAI_STORAGE_DIR` as the override, and then volunteers a real scenario — containers losing state on restart, CI runs inheriting state, or two processes fighting over one directory. Reciting only the environment-variable name is a thin answer.
- What goes wrong if two crew processes share one storage directory?They contend over a SQLite file and file-backed vector collections at the same time, which risks lock errors and corrupted or interleaved state, and it lets each process retrieve what the other wrote. Give each worker its own CREWAI_STORAGE_DIR, or move memory to an external provider configured through memory_config instead of sharing a filesystem path.
- Why might deleting the storage directory be the wrong way to clear memory?Because knowledge collections live there too. Deleting the directory discards the indexed document corpus along with accumulated memory, so the next kickoff re-chunks and re-embeds every source and you pay that embedding cost again. Use reset_memories with a specific command type, or the CLI reset, when you only want to drop what previous runs accumulated.
- How should the storage directory be treated in a container deployment?Decide explicitly whether state should survive. The default path is inside the container filesystem, so it disappears on restart and knowledge is re-indexed on every cold start. Set CREWAI_STORAGE_DIR to a mounted volume for persistence, or leave it ephemeral on purpose and accept the re-indexing cost, but do not let the behaviour be accidental.
saying these in an interview costs you the question
- Assumes state is written into the project directory
- Thinks re-cloning the repo resets a crew's memory
- Shares one storage directory across parallel workers
- Believes deleting the folder only clears memory
- Never sets the path in containers and wonders why state vanishes