skip to content

Which components make up a Haystack indexing pipeline, and what does each do?

level: juniorimportance: must knowfreq 72%

answer

  1. Five stages, files in, store out
  2. Everything after the converter speaks List[Document]
  3. Hygiene, then cut, then vectorise
  4. Splits inherit the parent's meta
  5. Writer emits a count, not documents

basics

~10 s

A Haystack indexing pipeline runs a converter (file bytes to Document), then DocumentCleaner (strip whitespace and boilerplate), DocumentSplitter (cut into smaller Documents), a document embedder (attach vectors), and DocumentWriter (persist to a DocumentStore).

solid answer

~40 s

An indexing pipeline in Haystack is an ordinary `Pipeline` whose components are ordered by what each one hands to the next. A converter such as `TextFileToDocument`, `PyPDFToDocument` or `MarkdownToDocument` turns file paths or `ByteStream` sources into `Document` objects and records `file_path` in their `meta`. `DocumentCleaner` normalises the text — removing empty lines, extra whitespace, optionally repeated headers and footers. `DocumentSplitter` cuts each Document into many smaller Documents, copying the parent's meta into every split and adding `source_id`, `split_id` and `page_number`. A document embedder such as `SentenceTransformersDocumentEmbedder` or `OpenAIDocumentEmbedder` sets `.embedding` on each Document. Finally `DocumentWriter` writes them into the `DocumentStore` and returns `documents_written`. Every stage passes `List[Document]` in and out, which is why the order is interchangeable in principle but load-bearing in practice: splitting after embedding would leave the splits without vectors.

code

python · 23 lines
python
from haystack import Pipeline
from haystack.components.converters import TextFileToDocument
from haystack.components.preprocessors import DocumentCleaner, DocumentSplitter
from haystack.components.embedders import SentenceTransformersDocumentEmbedder
from haystack.components.writers import DocumentWriter
from haystack.document_stores.in_memory import InMemoryDocumentStore

store = InMemoryDocumentStore()

indexing = Pipeline()
indexing.add_component("converter", TextFileToDocument())
indexing.add_component("cleaner", DocumentCleaner())
indexing.add_component("splitter", DocumentSplitter(split_by="word", split_length=200, split_overlap=20))
indexing.add_component("embedder", SentenceTransformersDocumentEmbedder(model="sentence-transformers/all-MiniLM-L6-v2"))
indexing.add_component("writer", DocumentWriter(document_store=store))

indexing.connect("converter", "cleaner")
indexing.connect("cleaner", "splitter")
indexing.connect("splitter", "embedder")
indexing.connect("embedder", "writer")

result = indexing.run({"converter": {"sources": ["handbook.txt"], "meta": {"team": "platform"}}})
print(result["writer"]["documents_written"])

go deeper

for a junior

Be able to name the five stages in order and say in one sentence what each one produces. Interviewers mostly want to hear that converters make Documents and the writer puts them in a store.

for a middle

Explain why the order is load-bearing: embeddings attach to whatever Document exists at that moment, so splitting after embedding leaves chunks unvectorised. Know that splits inherit the parent's meta.

for a senior

Show that you treat indexing as a pipeline you can test and re-run: routing mixed file types, deciding which metadata to attach at conversion time, and knowing that adding vector search later means a full re-index.

for a principal

Own the cost model. Embedding dominates indexing cost and latency, so batching, incremental re-indexing and deciding what actually needs a vector are the tradeoffs you should be framing, not the component list.

## The shape Haystack has no dedicated "indexing" object. An indexing pipeline is just a `Pipeline` whose components happen to move raw files into a `DocumentStore`. The canonical chain is: **Converter → DocumentCleaner → DocumentSplitter → DocumentEmbedder → DocumentWriter** What makes the chain composable is that four of the five stages speak the same currency: they take `documents: List[Document]` and return `documents: List[Document]`. Only the converter is different — it takes `sources` (file paths, strings, or `ByteStream` objects) and an optional `meta`. ## Stage by stage **Converters** live in `haystack.components.converters`: `TextFileToDocument`, `PyPDFToDocument`, `MarkdownToDocument`, `HTMLToDocument`, `DOCXToDocument`, `CSVToDocument`, `PPTXToDocument`, plus service-backed ones such as `TikaDocumentConverter` and `AzureOCRDocumentConverter`. Each produces one Document per source and puts `file_path` into `meta`. The `meta` argument accepts either one dict applied to every source, or a list of dicts aligned positionally with `sources` — that list form is how per-file metadata (tenant, product, publication date) enters the corpus. Because each converter handles one format, mixed corpora usually put a `FileTypeRouter` in front to fan sources out by MIME type. **DocumentCleaner** does text hygiene, not chunking. Its flags are `remove_empty_lines`, `remove_extra_whitespaces`, `remove_repeated_substrings` (useful for page headers and footers that PDF extraction repeats), `remove_substrings`, `remove_regex`, plus `unicode_normalization` and `ascii_only`. One flag matters more than it looks: `keep_id`, which defaults to `False`. Because a Document's auto-generated id is a hash of its content and meta, cleaning changes the id unless you ask it not to. **DocumentSplitter** turns one Document into many. `split_by` chooses the unit ("word", "sentence", "line", "page", "passage", "period", or "function" for a custom callable), `split_length` how many units per chunk, `split_overlap` how many are repeated between neighbours, and `respect_sentence_boundary` avoids cutting mid-sentence when splitting by word. Critically, the splitter copies the parent's `meta` into every split and adds `source_id` (the parent Document's id), `split_id`, `split_idx_start` and, where known, `page_number`. That copy is why metadata must be attached *before* splitting for filtered retrieval to work on chunks. **Document embedders** are the indexing-side half of Haystack's embedder pair. `SentenceTransformersDocumentEmbedder`, `OpenAIDocumentEmbedder`, `AzureOpenAIDocumentEmbedder`, `HuggingFaceAPIDocumentEmbedder` and friends take `documents` and return the same Documents with `.embedding` populated. Their text-embedder twins (`SentenceTransformersTextEmbedder`, `OpenAITextEmbedder`) take a single `text` string and are used at query time. Mixing them up is the classic first-day error: a document embedder cannot embed a query string, and a text embedder cannot fill in `Document.embedding`. **DocumentWriter** wraps `document_store.write_documents(...)`. It takes `document_store` and a `policy` (a `DuplicatePolicy`) at construction, takes `documents` at run time, and outputs `documents_written` — an integer, not the Documents. That means the writer is normally a terminal component; nothing downstream can consume its output as Documents. ## Why the order matters The stages are ordinary components, so nothing stops you wiring them differently — but the semantics change. Embedding before splitting produces vectors for whole documents that the splits then discard, so the writer stores chunks with no embedding and vector retrieval silently returns nothing. Cleaning after splitting wastes work and can leave chunk boundaries in odd places. Splitting before adding per-file metadata means the metadata never reaches the chunks. ## Pure-keyword pipelines If you are only doing BM25 retrieval, the embedder drops out entirely: converter → cleaner → splitter → writer is a complete pipeline, and `InMemoryDocumentStore` will still serve `InMemoryBM25Retriever`. Adding the embedder later means re-running the whole pipeline, because embeddings are a property of the stored Document, not something the store computes for you. Haystack never embeds implicitly — if no embedder ran, the vectors are simply absent.

  • What is the difference between a document embedder and a text embedder in Haystack?
    A document embedder (for example `SentenceTransformersDocumentEmbedder`) takes `documents` and writes the vector into each `Document.embedding`; it is the indexing-side component. Its text twin (`SentenceTransformersTextEmbedder`) takes a single `text` string and returns an `embedding` for a query. They must use the same model, or query and document vectors live in incomparable spaces.
  • Where does file-level metadata enter the pipeline, and how does it reach the chunks?
    Converters accept a `meta` argument — one dict for all sources, or a list aligned with `sources` for per-file values. Whatever lands on the parent Document is copied by `DocumentSplitter` into every split, alongside the `source_id`, `split_id` and `page_number` it adds itself. Metadata attached after splitting never reaches the chunks.
  • Can you skip the embedder entirely?
    Yes, if retrieval is keyword-only. Converter, cleaner, splitter and writer are enough to serve BM25 retrieval from a store. Haystack never embeds implicitly, so the Documents simply carry no `.embedding` — and adding vector retrieval later means re-running the pipeline with an embedder in it.

saying these in an interview costs you the question

  • Thinking the document store embeds documents automatically on write
  • Using a text embedder to embed documents at index time
  • Splitting after embedding and expecting chunks to have vectors
  • Attaching metadata after the splitter and expecting chunks to carry it
  • Believing DocumentWriter outputs the written Document objects

context