What does meta_fields_to_embed do on a Haystack document embedder?
answer
- It changes the input, not the output
- Titles ride along with the chunk
- Content in the store stays as written
- Filtering never looks at the vector
- High-cardinality values are pure noise
basics
~20 sIt prepends the named metadata values to the document text before the embedding model sees it, joined by embedding_separator. The vector then encodes that metadata, while the stored Document's content and meta are left unchanged.
solid answer
~50 sDocument embedders such as `SentenceTransformersDocumentEmbedder` and `OpenAIDocumentEmbedder` accept `meta_fields_to_embed=[...]` and `embedding_separator` (default a newline). At run time the embedder builds the text it sends to the model by joining the listed meta values and the document's `content` with that separator — so a chunk tagged `{"title": "Transaction isolation"}` is embedded as the title followed by the chunk text. Only the embedding changes: `Document.content` and `Document.meta` are stored exactly as they were, so nothing about filtering or the text shown to a generator is affected. It is the cheapest fix for context-free chunks — a chunk saying "this is disabled by default" is meaningless alone but findable once its section title rides along in the vector. The cost is tokens on every document plus the risk that noisy or high-cardinality meta (ids, timestamps, file paths) dilutes the semantic signal and pulls unrelated chunks together.
code
python · 18 linesfrom haystack import Document
from haystack.components.embedders import SentenceTransformersDocumentEmbedder
embedder = SentenceTransformersDocumentEmbedder(
model="sentence-transformers/all-MiniLM-L6-v2",
meta_fields_to_embed=["title"],
embedding_separator="\n",
)
embedder.warm_up()
doc = Document(
content="This is disabled by default and must be enabled per environment.",
meta={"title": "Read committed isolation"},
)
result = embedder.run(documents=[doc])
print(len(result["documents"][0].embedding)) # 384
print(result["documents"][0].content) # unchanged - only the vector saw the titlego deeper
Know that it puts selected metadata values in front of the chunk text before embedding, and that the stored content is unchanged. Naming the title-alongside-chunk use case is enough.
Explain the mechanics — joined by embedding_separator, absent keys skipped, only the vector affected — and why it is not the same mechanism as metadata filtering.
Show judgment about which fields earn their tokens: short and query-like helps, high-cardinality identifiers dilute the vector, and the query side has no metadata, so the match is asymmetric. Treat it as an A/B against an eval set.
Own the lifecycle cost. Any change to what goes into the vector means a full re-embed of the corpus, so decide early whether context enrichment lives in the embedding, in hybrid keyword matching, or in the chunking strategy itself.
## What it does mechanically A Haystack document embedder does not necessarily embed `Document.content` alone. Before calling the model it assembles a string: `<meta value 1><separator><meta value 2><separator><content>` using exactly the meta keys listed in `meta_fields_to_embed`, in that order, joined with `embedding_separator` (default `"\n"`). Keys that are absent from a given document are skipped rather than erroring. The resulting vector is written to `Document.embedding`; `Document.content` and `Document.meta` are untouched, and the store receives the document with its original text. Most document embedders expose the pair — `SentenceTransformersDocumentEmbedder`, `OpenAIDocumentEmbedder`, `AzureOpenAIDocumentEmbedder`, `HuggingFaceAPIDocumentEmbedder`, `CohereDocumentEmbedder` among them. The related `prefix` and `suffix` arguments do a similar thing with a constant string rather than per-document metadata. ## Why it exists Chunking destroys context. A 200-word chunk lifted out of the middle of a manual may read "This is disabled by default and must be enabled per environment." — a perfectly good sentence that no query can find, because it never names the feature. The information a searcher would use lives in the section heading, which the splitter left behind on the parent document or in a metadata key. Embedding the title alongside the chunk restores exactly that. It is the smallest possible version of the context-enrichment family: no extra LLM calls, no rewritten text, one constructor argument. ## Why it is not free **Token cost.** Every document carries the extra text, on every re-index. With an API embedder that is a real line item, and with a local model it is real latency. **Truncation.** The model's input window applies to the *assembled* string. Prepend a long meta value to a chunk already near the limit and the tail of the actual content is silently cut. Budget the meta text as part of the chunk size. **Signal dilution.** The vector is a single point representing everything you fed it. Adding a long, generic title to a short, specific chunk shifts the vector toward the title, so chunks from the same section start looking alike and retrieval loses the ability to distinguish them. Embedding a UUID, a file path, a timestamp or any high-cardinality identifier is pure noise — it cannot help matching and it costs tokens. **Asymmetry with the query.** Queries are embedded by a *text* embedder with no metadata to attach. So you are matching a metadata-enriched document vector against a plain query vector. That works when the metadata is the kind of thing users actually type — a product name, a document title, a section heading — and works badly when it is not. ## What it is not It is not a way to make metadata filterable. Filtering reads `Document.meta` through the filter syntax and is entirely independent of what went into the vector; a field can be filterable without being embedded, and vice versa. It is also not a way to store metadata — the meta dict is persisted regardless. And it does not append dimensions to the vector: the vector's dimensionality is fixed by the model. The distinction matters in an interview because the two mechanisms answer different questions. "Only search within the 2024 policy documents" is a filter. "Find the chunk about retries even though the chunk never says which service it is about" is `meta_fields_to_embed`. ## Choosing fields Good candidates are short, human-meaningful, and would plausibly appear in a query: document title, section heading, product or component name, sometimes the document type. Bad candidates are ids, hashes, file paths, timestamps, ingestion bookkeeping and anything with thousands of distinct values. A useful discipline is to treat it as a retrieval-quality experiment rather than a default: index the same corpus with and without the field, run your evaluation set against both, and keep the version that wins. Because the setting changes the vectors, switching it later requires re-embedding the entire corpus — and, since content-hash ids are unaffected by embeddings, a re-run with `OVERWRITE` will correctly replace them in place. ## A word on symmetry If the enriched documents perform badly for short keyword-ish queries, that is usually the asymmetry biting. The remedies are to shorten the embedded metadata, to restrict it to one high-value field, or to move the requirement into hybrid retrieval instead — a keyword match on the title achieves much of the same effect without touching the vector space.
- Does embedding a metadata field make it filterable?No — the two are independent. Filters read `Document.meta` through the filter syntax regardless of what text went into the embedding, and a field can be filterable without being embedded. `meta_fields_to_embed` only changes the string sent to the embedding model, which changes which chunks are semantically near a query.
- Why does the query side have no equivalent setting?Queries are embedded by a text embedder that receives a bare string with no metadata attached, so an enriched document vector is always matched against a plain query vector. That asymmetry is fine when the embedded metadata is the sort of thing users type — a title or product name — and harmful when it is bookkeeping the user would never mention.
- What happens if the metadata plus content exceeds the model's input limit?The assembled string is truncated by the model or its client, and since the metadata is prepended, it is the tail of the actual chunk content that is lost. Budget the metadata text as part of the chunk size, or keep the embedded fields short — one heading rather than a full breadcrumb path.
- You change meta_fields_to_embed on a live index. What must you do?Re-embed the whole corpus: existing vectors were built from a different string and are no longer comparable with newly written ones, so retrieval quality degrades unevenly. Embeddings are not part of the content-hash id, so a full re-run with DuplicatePolicy.OVERWRITE replaces documents in place rather than duplicating them.
saying these in an interview costs you the question
- Thinking it makes metadata available to filters
- Believing it adds extra dimensions to the vector
- Embedding ids, timestamps or file paths as if they help matching
- Assuming the stored document content is modified too
- Forgetting that changing it invalidates every existing embedding