In CrewAI, how do you attach a PDF or text file as crew knowledge?
answer
- one class per file type
- paths are not arbitrary
- a folder at project root
- knowledge_sources= on Crew or Agent
- chunk_size and chunk_overlap on the source
basics
~10 sCreate a source object such as PDFKnowledgeSource(file_paths=["report.pdf"]) and pass it as Crew(knowledge_sources=[source]). CrewAI resolves those paths inside a knowledge/ folder at the project root, so the file must sit there.
solid answer
~40 sCrewAI ships one class per source type — `StringKnowledgeSource`, `TextFileKnowledgeSource`, `PDFKnowledgeSource`, `CSVKnowledgeSource`, `JSONKnowledgeSource`, `ExcelKnowledgeSource` and `CrewDoclingSource` for the Docling-backed converter. You instantiate the one you need and pass a list to `Crew(knowledge_sources=[...])`, or to a single `Agent(knowledge_sources=[...])` if only that agent should see it. The file-based sources take `file_paths`, and those paths are resolved **relative to a `knowledge/` directory at the project root** — the single most common beginner error is putting the PDF next to the script and getting a file-not-found. At kickoff CrewAI splits each source into chunks (`chunk_size` / `chunk_overlap` are constructor arguments), embeds them with the configured embedder, and stores them so agents can query them while running tasks.
code
python · 12 linesfrom crewai import Agent, Crew, Task
from crewai.knowledge.source.pdf_knowledge_source import PDFKnowledgeSource
from crewai.knowledge.source.string_knowledge_source import StringKnowledgeSource
# resolved as <project-root>/knowledge/handbook.pdf
handbook = PDFKnowledgeSource(file_paths=["handbook.pdf"], chunk_size=2000, chunk_overlap=200)
policy = StringKnowledgeSource(content="Refunds are approved up to 200 USD without escalation.")
support = Agent(role="Support Lead", goal="Answer policy questions", backstory="Veteran support agent")
task = Task(description="Can a 150 USD refund be approved directly?", expected_output="Yes or no with the rule cited", agent=support)
crew = Crew(agents=[support], tasks=[task], knowledge_sources=[handbook, policy])go deeper
Be able to name a couple of source classes, show the one-liner that passes them to Crew(knowledge_sources=[...]), and state that file paths live under a knowledge/ folder in the project root.
Explain the kickoff pipeline — split into chunks, embed, store, retrieve top matches into the prompt — and say which knobs (chunk_size, chunk_overlap, result limits) you would turn when grounding is poor.
Talk about scoping sources per agent to keep retrieval precise, about the one-off embedding cost and API-key dependency of indexing, and about when the data belongs behind a tool because it must be fresh.
Own the decision of what belongs in a shared indexed corpus versus a per-request or per-tenant one, and be ready to justify the re-indexing and cost story when the source documents change frequently.
## What "knowledge" means in CrewAI Knowledge is the grounding material you attach to a crew *before* it runs: documents, records or raw strings that the agents should be able to consult while doing their work. It is distinct from tools (which fetch data at call time) and from memory (which accumulates from the run itself). You declare it once, CrewAI indexes it, and agents retrieve from it automatically as part of task execution. ## The source classes Each kind of input has its own class, imported from `crewai.knowledge.source.*`: - `StringKnowledgeSource(content="...")` — inline text, useful for policies or facts you generate in code. - `TextFileKnowledgeSource(file_paths=["notes.txt"])` - `PDFKnowledgeSource(file_paths=["report.pdf"])` - `CSVKnowledgeSource(file_paths=["sales.csv"])` - `JSONKnowledgeSource(file_paths=["config.json"])` - `ExcelKnowledgeSource(file_paths=["book.xlsx"])` - `CrewDoclingSource(file_paths=[...])` — routes the file through Docling, which handles a broader set of document formats and preserves more structure than the plain readers. All of them accept `chunk_size` and `chunk_overlap`, so you can tune how the document is cut up without writing your own splitter. ## The `knowledge/` directory rule File-based sources do **not** take arbitrary filesystem paths. CrewAI resolves `file_paths` against a `knowledge/` directory at the root of your project. So `PDFKnowledgeSource(file_paths=["handbook.pdf"])` looks for `<project-root>/knowledge/handbook.pdf`, and `file_paths=["docs/handbook.pdf"]` looks for `<project-root>/knowledge/docs/handbook.pdf`. Passing an absolute path to a file outside that tree, or leaving the PDF beside your `main.py`, produces a file-not-found error at kickoff. This trips up almost everyone once; remembering it is worth an easy interview point. ## Crew level versus agent level `Crew(knowledge_sources=[...])` makes the material available across the crew. `Agent(knowledge_sources=[...])` scopes it to that one agent, which matters when different roles should not see each other's material (a legal reviewer's contracts versus a marketer's brand guide) and when you want to keep each agent's retrieval set small and on-topic. You can use both together. ## What happens at kickoff When the crew starts, CrewAI takes each source, produces its chunks, embeds them with the configured embedder (OpenAI embeddings by default, which is why an `OPENAI_API_KEY` is required unless you override `embedder`), and writes them into the knowledge store on disk. During task execution the agent's query is embedded and the closest chunks are pulled back and injected into the prompt as extra context. Because the indexing happens on the way in, the very first kickoff after adding a large document is slower and costs embedding tokens; subsequent runs reuse what is already stored. ## Tuning retrieval If agents get too much or too little grounding, the knobs are the source's `chunk_size`/`chunk_overlap` and a `knowledge_config` on the agent or crew, built from `KnowledgeConfig`, which controls how many results come back and the minimum similarity score they must clear. Smaller chunks give sharper matches but lose surrounding context; larger chunks carry context but dilute the embedding and eat prompt budget. ## Common mistakes Attaching a 300-page PDF and expecting the agent to "read" it — it never sees the whole thing, only the retrieved chunks. Assuming knowledge updates itself when the file changes — you generally need to reset the knowledge store so the new content is re-embedded. And confusing knowledge with tools: if the data must be current at call time (a live API, a database query), that is a tool, not a knowledge source. ## Interview framing A strong answer names two or three concrete source classes, states the `knowledge/`-relative path rule, mentions both the crew-level and agent-level attachment points, and finishes by saying what happens mechanically — chunk, embed, store, retrieve into the prompt — rather than describing knowledge as something the model "has read".
- What is the practical difference between attaching a source at crew level and at agent level?Crew-level sources are available to the whole crew; agent-level sources are scoped to that single agent. Agent level is how you keep roles from seeing each other's material and how you keep each retrieval set small and on-topic, which improves match quality. You can combine both — shared corpus on the crew, role-specific documents on the agent.
- When should a document be a tool instead of a knowledge source?When freshness matters. Knowledge is chunked and embedded up front, so it reflects the file as of indexing time. If the agent needs the current value — a live price, a row in a database, today's tickets — that belongs behind a tool the agent calls during the run, not in a knowledge source that would go stale between kickoffs.
- Why does the first kickoff after adding a big PDF cost more and take longer?Because indexing happens on the way in: CrewAI splits the document into chunks and sends every chunk to the embedding model before the agents start working. That is a one-off embedding cost proportional to document size. Later runs reuse the stored vectors, so only the query embeddings are paid per run.
saying these in an interview costs you the question
- Thinks the agent reads the whole PDF into its context
- Puts the file next to the script instead of knowledge/
- Confuses a knowledge source with a tool call
- Believes editing the file re-indexes it automatically
- Assumes knowledge works without any embedding model configured