skip to content

How do you write a custom LlamaIndex loader by subclassing BaseReader?

level: middleimportance: should knowfreq 38%

answer

  1. one method, one return type
  2. stable identity comes from the source system
  3. attach filter keys before parsing, not after
  4. stream variant for paged sources
  5. resist splitting at load time

basics

~20 s

Subclass BaseReader and implement load_data(), returning a list of Document objects with text, a metadata dict, and a stable id_ derived from the source system. Use lazy_load_data() to stream large sources instead of materialising everything.

solid answer

~50 s

A custom loader subclasses `BaseReader` from `llama_index.core.readers.base` and implements `load_data(**kwargs) -> List[Document]`. Its whole job is fetching and normalising: pull the records from the API, database or file format, turn each into a `Document` with `text`, a `metadata` dict, and an `id_` taken from something durable in the source (a primary key, a URI, a content hash). For large or paged sources implement `lazy_load_data()` and yield Documents instead, so the pipeline can stream. Subclass `BasePydanticReader` instead when the reader itself needs to be serialisable as part of a stored pipeline. Two design rules matter more than the API: do not chunk inside the reader — that belongs to the node-parsing stage, and doing it early hides the source unit from citation and dedup; and do not let ids churn between runs, because deduplication, upserts and ref-doc deletion all key on them.

code

python · 22 lines
python
from typing import Iterable
from llama_index.core import Document
from llama_index.core.readers.base import BaseReader

class TicketReader(BaseReader):
    def __init__(self, client):
        self.client = client

    def lazy_load_data(self, project: str) -> Iterable[Document]:
        for ticket in self.client.iter_tickets(project):
            yield Document(
                text=f"{ticket['title']}\n\n{ticket['body']}",
                metadata={
                    "project": project,
                    "status": ticket["status"],
                    "url": ticket["url"],
                },
                id_=f"ticket:{ticket['id']}",
            )

    def load_data(self, project: str) -> list[Document]:
        return list(self.lazy_load_data(project))

go deeper

for a junior

Know that a custom loader subclasses BaseReader and implements load_data() returning a list of Documents, and that its job is fetching and converting source data, not indexing it.

for a middle

Explain what goes into each Document — text, metadata, a stable id_ — plus lazy_load_data for streaming and the file_extractor route for wiring a format reader into a directory scan.

for a senior

Argue the boundaries: why chunking stays out of the reader, why ids must come from the source system, how per-record failures are handled without silently dropping documents, and how you test a reader in isolation.

for a principal

Set the ingestion contract across teams — the required metadata schema, who owns document identity, and whether a bespoke reader is worth maintaining versus normalising upstream into a common store that one reader can serve.

## The interface ``` from llama_index.core.readers.base import BaseReader from llama_index.core import Document class TicketReader(BaseReader): def load_data(self, project: str) -> list[Document]: ... ``` That is nearly the whole contract. `load_data` returns `List[Document]`. `lazy_load_data` is the streaming variant that yields Documents one at a time; implement it when the source is large or paged, and callers that only need to iterate never have to hold the corpus in memory. `BasePydanticReader` is the variant to subclass when the reader object itself must be serialised — for example when it is part of a pipeline definition that gets persisted or shipped elsewhere — because it declares its configuration as typed, serialisable fields. ## What belongs inside the reader **Fetching and auth.** Pagination, retries, rate limiting, credentials. The reader is the only place in the pipeline that knows the source protocol, so this logic has nowhere better to live. **Normalisation to text.** Whatever the source's native shape — HTML, a JSON record, a row — the reader decides what the retrievable *text* is. For a structured record this usually means rendering fields into readable prose rather than dumping JSON, because the embedding model sees exactly this string. **Metadata.** Every filter key you will ever want at query time has to be attached here: tenant, author, status, timestamps, a URL you can cite. Document metadata propagates to every node parsed from that Document, so this is the cheapest possible place to set it. Keep values scalar — vector stores generally filter on flat, primitive-valued metadata, and nested dicts are a common source of "the filter silently matches nothing". **Identity.** Set `id_` from the source system's own key. This is the step people skip and later regret: with generated ids, a nightly re-run looks like an entirely new corpus, so a docstore-backed pipeline inserts duplicates instead of upserting, and `delete_ref_doc` has no name to delete by. ## What does not belong inside the reader **Chunking.** It is tempting to return pre-split 500-token Documents. Don't. Node parsing is a separate, tunable stage, and returning chunks as Documents means the true source unit no longer exists as an object — citation points at a fragment, deduplication happens per fragment, and re-tuning chunk size means re-fetching from the source instead of re-running a local transformation. **Embedding.** That is a transformation in the pipeline, not a loader concern. **Business filtering that belongs to retrieval.** Load broadly with good metadata and filter at query time; a reader that hard-codes "only published documents" forces a re-ingest when the rule changes. ## Two ways to plug it in A reader for a *format* is usually wired through `SimpleDirectoryReader`'s `file_extractor` dict, so directory scanning, exclusion rules and default file metadata still apply and only the parsing step is yours: ``` SimpleDirectoryReader("./data", file_extractor={".eml": EmailReader()}) ``` A reader for a *source system* is standalone: you call `load_data()` yourself and hand the Documents to an index constructor or an `IngestionPipeline`. ## Practical concerns - **Failure granularity.** One unparseable record should not kill a 100k-document run. Catch per-record, log the identifier, and keep going — but count the failures, because a silent drop rate is worse than a crash. - **Determinism.** The same source state should produce the same text and the same ids. Non-deterministic rendering makes content hashes churn and defeats caching. - **Testing.** A reader is unusually easy to unit-test: fixture in, `List[Document]` out. Assert ids, metadata keys and text shape. This is where most ingestion bugs are cheapest to catch.

  • Should a custom reader split long records into smaller Documents?
    No. Return the natural source unit and let a node parser split it downstream. Chunking in the reader destroys the source-unit identity that citation, deduplication and ref-doc deletion depend on, and it freezes a chunking decision into the fetch step — re-tuning chunk size then means re-hitting the source instead of re-running a local transformation.
  • When would you subclass BasePydanticReader instead of BaseReader?
    When the reader itself has to be serialisable — persisted as part of a pipeline definition or moved between processes. BasePydanticReader declares its configuration as typed fields so it round-trips, whereas a plain BaseReader holding an arbitrary client object generally does not.
  • One record in a 100k-record source fails to parse. What should the reader do?
    Handle it per record: catch, log the source identifier, increment a failure counter, and continue. Aborting the whole run over one malformed record is brittle, but swallowing failures silently is worse — the pipeline finishes green while the corpus quietly loses documents, and nobody notices until retrieval misses them.

saying these in an interview costs you the question

  • Pre-chunks inside the reader and calls each chunk a Document
  • Generates a fresh document id on every run
  • Dumps raw JSON as the document text
  • Puts nested dicts in metadata and expects filters to work
  • Thinks a custom reader must also produce embeddings

context