How do you decide between LlamaParse and LlamaIndex's built-in file readers for a PDF corpus?
answer
- a PDF describes pages, not structure
- tables and columns are where it breaks
- no text layer means nothing extracted
- hosted parser, per-page price, outbound data
- parse once, store the result
basics
~20 sDecide by how much of the corpus is layout-dependent. Built-in readers extract a raw text stream and lose tables, column order and scanned pages; LlamaParse is a hosted parser returning structured markdown, at a per-page price, added latency, and data leaving your environment.
solid answer
~50 sBuilt-in file readers pull the embedded text stream out of a PDF. That is fine for single-column, text-native documents and fails predictably elsewhere: tables collapse into ragged lines, multi-column pages interleave, headers and footers pollute chunks, and scanned pages yield nothing at all because there is no text layer. `LlamaParse` is a hosted parsing service that returns structured output — typically markdown via `result_type="markdown"` — preserving tables and heading hierarchy, and it is wired into ingestion through `SimpleDirectoryReader(file_extractor={".pdf": parser})` or called directly. The decision is a cost-and-control one: per-page pricing across the corpus, network latency and failure modes in the ingest path, and documents leaving your environment for a third-party service. The way to make it evidence-based is to build a small evaluation set of questions whose answers live in tables and figures, run both parsers, and compare retrieval quality — then persist parsed output as a durable artifact so you never pay to parse the same page twice.
code
python · 11 linesfrom llama_parse import LlamaParse
from llama_index.core import SimpleDirectoryReader
parser = LlamaParse(result_type="markdown") # reads LLAMA_CLOUD_API_KEY
docs = SimpleDirectoryReader(
input_dir="./reports",
required_exts=[".pdf"],
file_extractor={".pdf": parser},
filename_as_id=True,
).load_data()go deeper
Know that PDFs do not parse cleanly by default — tables and multi-column pages come out garbled — and that LlamaParse is the hosted option LlamaIndex offers for harder documents.
Explain the failure modes concretely and show the wiring: LlamaParse slots in through SimpleDirectoryReader's file_extractor with result_type set to markdown, leaving the rest of the pipeline unchanged.
Diagnose from symptoms — empty documents from scans, interleaved columns, tables that never retrieve — and put reliability around a hosted parser: retries, rate limits, a dead-letter path, and caching so pages are not re-parsed.
Own the tradeoff explicitly: corpus size times per-page price against retrieval value, data-residency constraints, a mixed routing policy by document family, and parsed output as a durable versioned artifact so the cost is paid once.
## What actually differs A PDF is a page-description format, not a document format. The built-in readers do the honest thing available to them: extract the embedded text stream in roughly the order the file stores it. Consequences are systematic, not random: - **Tables** become sequences of loose cell strings. The row/column relationship — which is usually the *answer* — is gone before chunking even starts. - **Multi-column layouts** interleave. Two columns of a research paper can alternate line by line, producing chunks that are locally fluent and globally nonsense. - **Headers, footers, page numbers** land inside every chunk. - **Scanned pages** have no text layer, so extraction yields empty or near-empty documents. Nothing downstream reports this as a failure; you simply get an index that silently omits part of the corpus. - **Structure is lost.** No heading hierarchy means no way to attach section context to a chunk. LlamaParse is a hosted service built for this problem. It returns structured output — `result_type="markdown"` is the common choice — so tables arrive as markdown tables, headings as headings, and reading order is reconstructed. It accepts a natural-language `parsing_instruction` to bias how a document family is handled, and it exposes async loading for throughput. Because it is a service, it needs an API key and it sends your documents outbound. ## The decision axes **How much of the corpus is layout-dependent?** Not "is it PDF" — "do the answers live in tables, figures or multi-column text?" A corpus of markdown-exported reports gains nothing. A corpus of financial statements or lab reports gains most of its usable content. **What fraction of pages have no text layer?** Scanned material is a hard yes for a layout-aware parser; the built-in path returns nothing at all, so the alternative is not "worse quality", it is "absent documents". **Corpus size times page price.** Per-page pricing turns a 2-million-page archive into a real budget line. Note the asymmetry: parsing is a *one-time* cost per document version, while retrieval quality is paid back on every query. That usually favours parsing well — provided you never re-parse. **Data residency and confidentiality.** Sending contracts, medical records or regulated documents to a third-party service is a governance decision, not an engineering one. Where that is disallowed, the alternatives are a self-hosted layout model or restructuring upstream so the source system emits structured content directly. **Ingest-path reliability.** A hosted parser puts a network dependency and a rate limit inside ingestion. It needs retries, backpressure and a dead-letter path for documents that fail — none of which a local reader needed. **Latency profile.** Bulk backfill tolerates slow parsing; user-uploads-a-file-and-asks-a-question does not. Those two paths may warrant different parsers. ## Treat parsed output as an artifact The single most valuable architectural move here is to stop treating parsing as a step inside the ingestion pipeline and start treating its output as a stored artifact keyed by document id and content hash. Parse once, write the markdown to object storage, and let chunking, embedding and re-embedding read from there. Then re-tuning chunk size, swapping embedding models, or rebuilding the index after a bad deploy costs nothing extra — whereas a pipeline that re-parses on every rebuild pays the per-page price again each time, which is how parsing budgets get burned. ## Making it evidence-based Do not decide from the sample page in the docs. Take 30–50 real questions whose answers you know live in tables, figures or multi-column sections. Ingest a corpus slice both ways. Measure whether the right passage is retrieved at all — recall at k on known-correct sources — before measuring answer quality, because a parser failure shows up as a retrieval miss, not as a bad generation. A mixed policy is often the right answer: cheap local parsing for the text-native majority, the paid parser routed only at the layout-heavy minority, with routing driven by a detectable property such as an absent text layer or a detected table density. ## Wiring Because LlamaParse plugs in through `file_extractor`, adopting it does not restructure the pipeline — directory scanning, exclusion rules, default file metadata and everything downstream are unchanged. That reversibility is worth stating in an interview: the choice is swappable per file type, which makes a staged rollout by document family straightforward.
- How would you prove the parser change actually improved the system?Build an evaluation set of questions whose answers you know sit in tables, figures or multi-column text, record the source passage for each, and measure retrieval recall at k under both parsers before looking at answer quality. A parser failure surfaces as the right passage never being retrieved, so measuring generation first hides the cause.
- Data residency rules forbid sending documents to a hosted parser. What are your options?Run a layout-aware parser inside your own boundary, or push the problem upstream so the source system emits structured content — many report generators can export the underlying data directly. Failing both, scope RAG to the text-native subset and be explicit that table-dependent questions are out of scope rather than silently unanswerable.
- Why store parsed output as a separate artifact rather than parsing inside the pipeline?Because parsing is the expensive, externally-priced step and everything after it is cheap and frequently re-run. Persisting parsed markdown keyed by document id and content hash means re-chunking, swapping embedding models or rebuilding a corrupted index costs nothing extra, while an inline parse pays the per-page price on every rebuild.
saying these in an interview costs you the question
- Assumes any PDF reader recovers table structure
- Forgets scanned pages have no text layer to extract
- Treats parsing cost as recurring rather than one-time per version
- Ignores that documents leave the environment for a hosted parser
- Chooses a parser from a demo page rather than an eval set