How do you write a custom LlamaIndex loader by subclassing BaseReader?
answer
- one method, one return type
- stable identity comes from the source system
- attach filter keys before parsing, not after
- stream variant for paged sources
- resist splitting at load time
basics
~20 sSubclass BaseReader and implement load_data(), returning a list of Document objects with text, a metadata dict, and a stable id_ derived from the source system. Use lazy_load_data() to stream large sources instead of materialising everything.
solid answer
~50 sA custom loader subclasses `BaseReader` from `llama_index.core.readers.base` and implements `load_data(**kwargs) -> List[Document]`. Its whole job is fetching and normalising: pull the records from the API, database or file format, turn each into a `Document` with `text`, a `metadata` dict, and an `id_` taken from something durable in the source (a primary key, a URI, a content hash). For large or paged sources implement `lazy_load_data()` and yield Documents instead, so the pipeline can stream. Subclass `BasePydanticReader` instead when the reader itself needs to be serialisable as part of a stored pipeline. Two design rules matter more than the API: do not chunk inside the reader — that belongs to the node-parsing stage, and doing it early hides the source unit from citation and dedup; and do not let ids churn between runs, because deduplication, upserts and ref-doc deletion all key on them.
code
python · 22 linesfrom typing import Iterable
from llama_index.core import Document
from llama_index.core.readers.base import BaseReader
class TicketReader(BaseReader):
def __init__(self, client):
self.client = client
def lazy_load_data(self, project: str) -> Iterable[Document]:
for ticket in self.client.iter_tickets(project):
yield Document(
text=f"{ticket['title']}\n\n{ticket['body']}",
metadata={
"project": project,
"status": ticket["status"],
"url": ticket["url"],
},
id_=f"ticket:{ticket['id']}",
)
def load_data(self, project: str) -> list[Document]:
return list(self.lazy_load_data(project))go deeper
Know that a custom loader subclasses BaseReader and implements load_data() returning a list of Documents, and that its job is fetching and converting source data, not indexing it.
Explain what goes into each Document — text, metadata, a stable id_ — plus lazy_load_data for streaming and the file_extractor route for wiring a format reader into a directory scan.
Argue the boundaries: why chunking stays out of the reader, why ids must come from the source system, how per-record failures are handled without silently dropping documents, and how you test a reader in isolation.
Set the ingestion contract across teams — the required metadata schema, who owns document identity, and whether a bespoke reader is worth maintaining versus normalising upstream into a common store that one reader can serve.
## The interface ``` from llama_index.core.readers.base import BaseReader from llama_index.core import Document class TicketReader(BaseReader): def load_data(self, project: str) -> list[Document]: ... ``` That is nearly the whole contract. `load_data` returns `List[Document]`. `lazy_load_data` is the streaming variant that yields Documents one at a time; implement it when the source is large or paged, and callers that only need to iterate never have to hold the corpus in memory. `BasePydanticReader` is the variant to subclass when the reader object itself must be serialised — for example when it is part of a pipeline definition that gets persisted or shipped elsewhere — because it declares its configuration as typed, serialisable fields. ## What belongs inside the reader **Fetching and auth.** Pagination, retries, rate limiting, credentials. The reader is the only place in the pipeline that knows the source protocol, so this logic has nowhere better to live. **Normalisation to text.** Whatever the source's native shape — HTML, a JSON record, a row — the reader decides what the retrievable *text* is. For a structured record this usually means rendering fields into readable prose rather than dumping JSON, because the embedding model sees exactly this string. **Metadata.** Every filter key you will ever want at query time has to be attached here: tenant, author, status, timestamps, a URL you can cite. Document metadata propagates to every node parsed from that Document, so this is the cheapest possible place to set it. Keep values scalar — vector stores generally filter on flat, primitive-valued metadata, and nested dicts are a common source of "the filter silently matches nothing". **Identity.** Set `id_` from the source system's own key. This is the step people skip and later regret: with generated ids, a nightly re-run looks like an entirely new corpus, so a docstore-backed pipeline inserts duplicates instead of upserting, and `delete_ref_doc` has no name to delete by. ## What does not belong inside the reader **Chunking.** It is tempting to return pre-split 500-token Documents. Don't. Node parsing is a separate, tunable stage, and returning chunks as Documents means the true source unit no longer exists as an object — citation points at a fragment, deduplication happens per fragment, and re-tuning chunk size means re-fetching from the source instead of re-running a local transformation. **Embedding.** That is a transformation in the pipeline, not a loader concern. **Business filtering that belongs to retrieval.** Load broadly with good metadata and filter at query time; a reader that hard-codes "only published documents" forces a re-ingest when the rule changes. ## Two ways to plug it in A reader for a *format* is usually wired through `SimpleDirectoryReader`'s `file_extractor` dict, so directory scanning, exclusion rules and default file metadata still apply and only the parsing step is yours: ``` SimpleDirectoryReader("./data", file_extractor={".eml": EmailReader()}) ``` A reader for a *source system* is standalone: you call `load_data()` yourself and hand the Documents to an index constructor or an `IngestionPipeline`. ## Practical concerns - **Failure granularity.** One unparseable record should not kill a 100k-document run. Catch per-record, log the identifier, and keep going — but count the failures, because a silent drop rate is worse than a crash. - **Determinism.** The same source state should produce the same text and the same ids. Non-deterministic rendering makes content hashes churn and defeats caching. - **Testing.** A reader is unusually easy to unit-test: fixture in, `List[Document]` out. Assert ids, metadata keys and text shape. This is where most ingestion bugs are cheapest to catch.
- Should a custom reader split long records into smaller Documents?No. Return the natural source unit and let a node parser split it downstream. Chunking in the reader destroys the source-unit identity that citation, deduplication and ref-doc deletion depend on, and it freezes a chunking decision into the fetch step — re-tuning chunk size then means re-hitting the source instead of re-running a local transformation.
- When would you subclass BasePydanticReader instead of BaseReader?When the reader itself has to be serialisable — persisted as part of a pipeline definition or moved between processes. BasePydanticReader declares its configuration as typed fields so it round-trips, whereas a plain BaseReader holding an arbitrary client object generally does not.
- One record in a 100k-record source fails to parse. What should the reader do?Handle it per record: catch, log the source identifier, increment a failure counter, and continue. Aborting the whole run over one malformed record is brittle, but swallowing failures silently is worse — the pipeline finishes green while the corpus quietly loses documents, and nobody notices until retrieval misses them.
saying these in an interview costs you the question
- Pre-chunks inside the reader and calls each chunk a Document
- Generates a fresh document id on every run
- Dumps raw JSON as the document text
- Puts nested dicts in metadata and expects filters to work
- Thinks a custom reader must also produce embeddings