skip to content

In LangChain, what does a document loader return, and when do you use lazy_load()?

level: juniorimportance: should knowfreq 48%

answer

  1. Loaders emit a carrier type, not chunks
  2. page_content plus a metadata dict
  3. Eager list versus streaming iterator
  4. Metadata is what filters and citations use
  5. lazy_load yields one Document at a time

basics

~20 s

LangChain document loaders return Document objects, each with page_content text and a metadata dict (source, page). load() builds the full list in memory; lazy_load() yields Documents one at a time, so large corpora stream into splitting and embedding.

solid answer

~40 s

Every loader implements the `BaseLoader` interface, so it produces `Document` objects — a `page_content` string plus a `metadata` dict — not raw strings and not chunks. `load()` returns the whole `list[Document]` at once, which is fine for a handful of files; `lazy_load()` returns an iterator that yields Documents as they are parsed, so a 50k-file crawl never holds everything in memory. Async variants (`aload`, `alazy_load`) exist for I/O-bound sources. The granularity is loader-specific: a PDF loader typically emits one Document per page, a web loader one per URL. Loaders do not chunk — splitting is a separate step — and the metadata they attach (`source`, `page`, and anything you add) is what later powers metadata filters and citations, so it is worth normalising at load time rather than after embedding.

code

python · 5 lines
python
from langchain_community.document_loaders import PyPDFLoader

loader = PyPDFLoader("handbook.pdf")
for doc in loader.lazy_load():
    print(doc.metadata["page"], len(doc.page_content))

go deeper

for a junior

Be able to say that loaders return Document objects with page_content and metadata, and that splitting is a separate step you do next.

for a middle

Explain the load() versus lazy_load() tradeoff in memory terms, and show how metadata set before splitting propagates to every chunk.

for a senior

Talk about ingestion as a pipeline with durable, incremental writes, and about validating extraction quality before you pay for embeddings.

for a principal

Own the ingestion contract across sources: normalised metadata keys, stable document ids, and a re-index and delete story that survives loader changes.

## What a loader actually produces A LangChain document loader is a thin adapter from some source — a PDF, a URL, a CSV, an S3 prefix — into the framework's one universal carrier type, `Document`. A `Document` has two things that matter: `page_content`, the text that will eventually be embedded, and `metadata`, a plain dict of whatever the loader knew about the text's origin. Loaders in the community distribution live under `langchain_community.document_loaders` and all implement `BaseLoader`. The important mental correction for newcomers: a loader does **not** chunk. It does not size anything to your embedding model's window. It converts bytes into Documents at whatever natural granularity the format has — one per PDF page, one per web page, one per CSV row — and stops. Splitting is a separate, explicit stage. ## load() vs lazy_load() `load()` returns `list[Document]`: everything, materialised. `lazy_load()` returns an `Iterator[Document]`: one at a time, parsed on demand. In current versions `load()` is effectively `list(self.lazy_load())`, so the two produce identical Documents — the difference is purely memory and time-to-first-document. That difference stops being academic at scale. A directory of ten thousand PDFs loaded eagerly holds every page's text in the process at once, before a single embedding call has been made. Streaming lets you pipeline: load a Document, split it, embed the chunks, write them to the store, discard, repeat. Async equivalents (`aload`, `alazy_load`) matter when the source is network-bound, because the bottleneck is latency rather than CPU. ## Metadata is the part people underestimate Only `page_content` is embedded. The metadata dict rides along untouched, and it is the reason a retrieval system is operable at all: - **Citations.** After retrieval you show the user `metadata["source"]` and `metadata["page"]`. If the loader did not record them, you cannot add them later without re-ingesting. - **Filters.** Vectorstore retrievers accept store-specific metadata filters in `search_kwargs`. A filter on `tenant_id` or `doc_type` only works if that key exists on every chunk. - **Re-indexing.** To update or delete a document's chunks you need a stable identifier on them. Stamp your own key at load time. Splitters copy the parent Document's metadata onto every chunk they emit, so anything you set before splitting propagates. Anything you forget does not. ## Where loaders sit in the pipeline The canonical ingestion path is load → split → embed → write. Each stage has a different failure signature, and conflating them makes debugging retrieval quality miserable. If answers cite the wrong page, that is usually loader granularity or metadata. If answers cite the right document but the wrong passage, that is splitting. If nothing relevant comes back at all, that is retrieval configuration or an embedding mismatch. ## Failure modes worth naming - **Garbage extraction.** A PDF loader over a scanned document returns Documents with empty or whitespace `page_content`; they embed to nothing useful and quietly pollute the index. Check for near-empty Documents before embedding. - **Eager loads in a job with a memory limit.** The process dies during ingestion with nothing written; `lazy_load()` plus incremental writes makes progress durable. - **Loader granularity mismatch.** A CSV loader emitting one Document per row produces thousands of tiny Documents that a splitter will happily pass through unchanged, giving you retrieval units too small to answer anything. - **Unnormalised sources.** The same file loaded from two paths yields two different `source` values and duplicate chunks that both surface in results. The practical habit: after wiring a loader, print `len(docs)`, one sample `page_content[:300]`, and one sample `metadata`. Almost every ingestion bug is visible in those three values before you have spent a cent on embeddings.

  • If only page_content is embedded, why bother normalising metadata at load time?
    Because metadata is copied onto every chunk by the splitter and is the only handle you have afterwards. Citations read `source` and `page`; metadata filters in a retriever's `search_kwargs` need a consistent key on every chunk; and deleting or refreshing one document's chunks requires a stable id you stamped before embedding. Adding it later means re-ingesting the corpus.
  • How would you spot a PDF that loaded but extracted nothing useful?
    Inspect the Documents before embedding: flag any whose `page_content` is empty or below a small character threshold, and sample a few for readable text. Scanned or image-only PDFs come back blank or as ligature noise from a text-layer parser, which embeds into meaningless vectors that still get returned as neighbours. Those files need an OCR-capable loader instead.
  • Does a loader guarantee one Document per file?
    No — granularity is loader-specific. A PDF loader commonly emits one Document per page, a CSV loader one per row, a web loader one per URL, and a directory loader delegates to a per-file loader and concatenates the results. You should check `len(docs)` against your expectation rather than assume a one-to-one mapping.

saying these in an interview costs you the question

  • Thinks the loader chunks text to the embedding window
  • Assumes every loader returns exactly one Document per file
  • Believes metadata is embedded along with page_content
  • Calls load() on a huge corpus and is surprised by memory use
  • Adds source metadata only after embedding, then cannot cite

context