skip to content

LlamaIndex

You will learn the data framework for RAG: ingestion and loaders, node parsing and chunking, indexes and vector stores, retrievers and query engines, agentic workflows, and evaluation. Interviewers ask because LlamaIndex is organised around the data pipeline, which forces you to hold opinions on chunking, index choice, and how retrieval quality gets measured.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

explore

questions

page 2 of 2

When would you replace a LlamaIndex FunctionAgent with a hand-written Workflow?

level: principalimportance: should knowfreq 36%

basics

~20 s

When the control flow is known in advance. A prebuilt agent pays an LLM turn to decide every step; a hand-written Workflow encodes the sequence in typed events and steps, making it cheaper, deterministic and testable — at the price of owning the loop yourself.

open as a page

How would you prove a chunking change actually improved a LlamaIndex RAG pipeline?

level: principalimportance: should knowfreq 38%

basics

~20 s

Freeze an eval set, then measure retrieval and generation separately: RetrieverEvaluator with hit_rate and MRR over labelled query/node pairs, and BatchEvalRunner with the faithfulness and correctness evaluators over answers. Hold the judge model fixed, and report cost and latency next to quality.

open as a page

When should a LlamaIndex app move off SimpleVectorStore to Chroma, Pinecone or Weaviate?

level: principalimportance: should knowfreq 38%

basics

~20 s

Move when the index stops fitting one process: the corpus exceeds memory, several replicas must share it, writes must be incremental rather than a whole-file rewrite, or durability and concurrent access become requirements. Until then SimpleVectorStore is faster to iterate on.

open as a page

How do you decide between LlamaParse and LlamaIndex's built-in file readers for a PDF corpus?

level: principalimportance: should knowfreq 34%

basics

~20 s

Decide by how much of the corpus is layout-dependent. Built-in readers extract a raw text stream and lose tables, column order and scanned pages; LlamaParse is a hosted parser returning structured markdown, at a per-page price, added latency, and data leaving your environment.

open as a page

Your LlamaIndex query engine misses its p95 latency budget — which knobs do you trade?

level: principalimportance: should knowfreq 40%

basics

~20 s

Attribute the time first: measure retriever.retrieve() alone against the full query() to see whether retrieval, postprocessing or synthesis dominates. Then trade deliberately — candidate pool size, reranker placement, response mode and streaming each buy latency back at a different quality cost.

open as a page

showing 31–35 of 35