What does SimpleDirectoryReader do in LlamaIndex, and how do you control which files it loads?
answer
- local folder in, Documents out
- one knob whitelists extensions
- another overrides the parser per type
- a flag makes ids survive re-runs
- PDFs do not arrive one per file
basics
~10 sSimpleDirectoryReader walks a folder, picks a reader based on each file's extension, and returns Document objects with file metadata attached. You narrow what it touches with input_files, required_exts, exclude, exclude_hidden, recursive and num_files_limit.
solid answer
~40 s`SimpleDirectoryReader` is the default entry point for local ingestion. You give it `input_dir` (or an explicit `input_files` list), it selects a per-extension reader for each file — text, PDF, DOCX, CSV and so on — and `load_data()` returns a `List[Document]`. Selection is controlled by `recursive=True` to descend into subfolders, `required_exts=[".pdf"]` to whitelist extensions, `exclude` for glob patterns, `exclude_hidden`, and `num_files_limit`. Loading behaviour is controlled by `file_extractor` to override the reader for an extension, `file_metadata` to attach your own metadata per path, and `filename_as_id=True` to give each Document a stable id instead of a random UUID. By default each Document already gets metadata such as `file_path`, `file_name`, `file_type`, `file_size` and modification timestamps. `load_data(num_workers=N)` parallelises across files.
code
python · 14 linesfrom llama_index.core import SimpleDirectoryReader
def file_metadata(path: str) -> dict:
return {"team": "support", "source_path": path}
reader = SimpleDirectoryReader(
input_dir="./data",
recursive=True,
required_exts=[".md", ".pdf"],
exclude=["**/drafts/*"],
file_metadata=file_metadata,
filename_as_id=True,
)
docs = reader.load_data(show_progress=True, num_workers=4)go deeper
Know the two-line usage — construct with a directory, call load_data(), get Documents — and be able to name required_exts and recursive as the way to narrow the file set.
Explain the parser selection: extension-to-reader mapping, file_extractor overrides, and the fact that granularity varies by format so a PDF arrives as one Document per page. Mention the default file metadata fields.
Focus on re-ingestion: filename_as_id for stable identity, file_metadata as the cheap place to attach filter keys, num_workers and iter_data for large corpora, and when a directory scan stops being viable against a changing source.
Own the boundary — decide whether file enumeration and change detection belong inside the reader at all, or whether an external inventory drives input_files so ingestion is incremental and auditable rather than a full rescan.
## What it is `SimpleDirectoryReader` (importable from `llama_index.core`) is the loader almost every LlamaIndex tutorial opens with, and the one most production pipelines still use for filesystem-backed corpora. It does three things: enumerate files, choose a parser per file type, and return `Document` objects with useful metadata already attached. ``` from llama_index.core import SimpleDirectoryReader docs = SimpleDirectoryReader("./data").load_data() ``` That one line hides several decisions, and interviews poke at exactly those. ## Choosing the file set - `input_dir` — the folder to scan. Non-recursive by default; `recursive=True` descends into subdirectories. - `input_files` — an explicit list of paths, used instead of `input_dir` when the caller already knows the file set (for example from a change-detection job). - `required_exts=[".pdf", ".md"]` — whitelist of extensions; everything else is skipped. - `exclude=["**/drafts/*"]` — glob patterns to skip. - `exclude_hidden` — on by default, so dotfiles and dot-directories are ignored. This bites people whose corpus genuinely lives under a hidden path. - `num_files_limit` — a cap, useful for smoke tests against a huge corpus. - `fs` — an fsspec filesystem, which is how you point the same reader at object storage instead of local disk. Unsupported extensions are skipped rather than crashing the run; `raise_on_error` changes error handling for files that do fail to parse. ## Choosing the parser The reader maps each extension to a reader class. Support for rich formats lives in the `llama-index-readers-file` package, which the `llama-index` meta-package pulls in — that is why a PDF "just works" but the same code in a minimal `llama-index-core` install does not. Granularity differs by format, and this matters more than people expect: a plain text or markdown file becomes one Document, while the default PDF reader produces **one Document per page**. Your "1,000 files" can easily become 30,000 Documents. Downstream chunking and any per-Document metadata cost scale with that number. You override the mapping with `file_extractor`, a dict from extension to a reader instance: ``` SimpleDirectoryReader("./data", file_extractor={".pdf": MyPdfReader()}) ``` Note that `file_extractor` chooses *how* a matched file is read; it does not filter which files are matched — that is `required_exts`. ## Metadata Every Document gets default file metadata: `file_path`, `file_name`, `file_type`, `file_size`, plus creation and modification timestamps. Pass `file_metadata=lambda path: {...}` to add your own — tenant, department, source URL, anything you will later want to filter retrieval on. Because Document metadata is copied onto every node parsed from it, this callback is the cheapest place in the whole pipeline to make filtered retrieval possible. ## Identity By default each Document gets a generated id, fresh on every run. `filename_as_id=True` makes the id derive from the path instead. That single flag is what makes re-ingestion idempotent: deduplication, upserts, ref-doc deletion and refresh all key on document id, and all of them silently degrade into "everything is new" when ids churn. ## Loading modes - `load_data()` — eager, returns the full list. - `load_data(show_progress=True, num_workers=4)` — parallel across files; worth it for large corpora, and the point at which per-file parser cost becomes visible. - `iter_data()` — yields batches instead of materialising everything, for corpora that will not fit in memory. - `aload_data()` — the async variant. ## When it stops being the right tool It is a *directory* reader: no incremental change detection, no auth, no pagination, no rate limiting. Once the source is an API, a database, a ticketing system or a bucket with millions of objects and a change feed, you either drive `input_files` from an external change list or write a reader of your own. Reaching for it as a batch job over an ever-growing folder, with no stable ids and no dedup, is the classic way to end up with a vector store full of duplicates.
- How many Documents does a 40-page PDF produce with the default reader?Forty — the default PDF reader emits one Document per page, not one per file. That inflates document counts, changes what "one source" means for deduplication and citation, and means page-level metadata is available for free. If you want the whole file as a single Document, supply your own reader through `file_extractor` that concatenates pages.
- Why would you set filename_as_id=True?To make document identity stable across runs. Without it every load produces fresh generated ids, so a docstore or refresh check sees entirely new documents each time and duplicates the corpus. With it the id derives from the path, which is what deduplication, upserts, `delete_ref_doc` and `refresh_ref_docs` all key on.
- The reader skips your files entirely and returns an empty list — where do you look?Check `required_exts` for a missing or mistyped extension, `recursive` if the files are in subfolders, and `exclude_hidden` if the corpus lives under a dot-directory. Also confirm the format's reader package is installed — `llama-index-core` alone does not bring the file readers, so unsupported extensions are simply skipped.
saying these in an interview costs you the question
- Thinks it returns ready-to-embed chunks rather than Documents
- Assumes it recurses into subdirectories by default
- Believes file_extractor filters which files are loaded
- Expects one Document per PDF file
- Never sets a stable id and re-runs the loader nightly