skip to content

How do you keep LlamaIndex metadata out of embeddings while the LLM still sees it?

level: seniorimportance: should knowfreq 42%

answer

  1. metadata is rendered into the embedded string
  2. two independent exclusion lists
  3. hidden does not mean unstored
  4. one enum shows what each model sees
  5. short chunks suffer most

basics

~10 s

Set excluded_embed_metadata_keys on the Document or node to hide keys from the embedded text, and excluded_llm_metadata_keys to hide keys from the prompt. Verify with get_content(metadata_mode=MetadataMode.EMBED) or MetadataMode.LLM.

solid answer

~40 s

In LlamaIndex, metadata is not just a filter sidecar — by default it is rendered into the text that gets embedded *and* into the text handed to the LLM. Each `Document`/`TextNode` carries `excluded_embed_metadata_keys` and `excluded_llm_metadata_keys`; a key listed there is still stored and still filterable, but is dropped from that particular rendering. You inspect the result with `node.get_content(metadata_mode=MetadataMode.EMBED)` versus `MetadataMode.LLM`, `MetadataMode.ALL` and `MetadataMode.NONE`. Formatting is controlled by `metadata_template`, `metadata_seperator` and `text_template`. This matters because noisy keys — file paths, ids, byte sizes, timestamps — dilute a short chunk's embedding and consume its token budget, while a genuinely semantic key like a document title can *improve* retrieval. Set the exclusions on the Document at ingestion time so they propagate to every node parsed from it.

code

python · 17 lines
python
from llama_index.core import Document
from llama_index.core.schema import MetadataMode

doc = Document(
    text="Refunds are processed within 5 business days.",
    metadata={
        "title": "Refund policy",
        "file_path": "/srv/corpus/policies/refunds.md",
        "file_size": "2048",
        "tenant_id": "acme-42",
    },
    excluded_embed_metadata_keys=["file_path", "file_size", "tenant_id"],
    excluded_llm_metadata_keys=["file_size", "tenant_id"],
)

print(doc.get_content(metadata_mode=MetadataMode.EMBED))
print(doc.get_content(metadata_mode=MetadataMode.LLM))

go deeper

for a junior

Know that metadata attached at load time is not purely a sidecar — it can end up in the text that gets embedded and shown to the model — and that LlamaIndex has fields to exclude keys from each.

for a middle

Name the two exclusion lists and the MetadataMode values, and explain the templates that control rendering. Be able to show what a node will actually embed rather than assuming it is the raw text.

for a senior

Reason about the consequence: metadata noise dominating short chunks, prompt-token cost across ten retrieved chunks, reranker inputs, and the fact that changing the rendering forces a re-embed of the whole corpus.

for a principal

Set the policy — a corpus-wide convention for which keys are semantic, which are citation-only, and which are internal — and treat the embedded-text recipe as a versioned property of the index, since changing it is a full re-embedding event.

## The behaviour people miss A LlamaIndex node is not embedded as `node.text`. It is embedded as a rendering of metadata plus text. Load a folder with `SimpleDirectoryReader` and every node quietly carries `file_path`, `file_name`, `file_type`, `file_size`, `creation_date` and `last_modified_date` — and by default those strings sit in front of the chunk text when the embedding is computed and when the chunk is placed in the prompt. On a 1,000-token chunk that is noise you can absorb. On a 128-token chunk it can be a third of the tokens, and the embedding drifts toward "this is a file from a directory" and away from the sentence's meaning. Symptoms are recognisable: chunks from the same directory cluster together regardless of topic, similarity scores compress into a narrow band, and short chunks retrieve worse than long ones. ## The controls Four fields on every `Document`/`TextNode`: - `excluded_embed_metadata_keys: list[str]` — keys omitted when building the text to embed. - `excluded_llm_metadata_keys: list[str]` — keys omitted when building the text shown to the LLM. - `metadata_template` — how one key/value pair is rendered, e.g. `"{key}: {value}"`. - `metadata_seperator` — what joins the rendered pairs (note the spelling in the library). Plus `text_template`, which controls how the rendered metadata block and the content are combined. The two exclusion lists are independent, and that independence is the whole point. A key can be: - **In both renderings** — a document title or section heading. Genuinely semantic, helps the embedding, and helps the model know what it is reading. - **Embed-excluded, LLM-visible** — a source URL or filename. Useless as embedding signal, essential when you want the model to cite where an answer came from. - **Excluded from both** — internal ids, byte sizes, ingestion timestamps, tenant keys. Still stored, still available for metadata filtering at retrieval time, invisible to both models. Exclusion never removes metadata from storage. Filtering at query time reads the stored `metadata` dict, so a fully excluded key still works perfectly as a filter. ## Verifying it ``` from llama_index.core.schema import MetadataMode print(node.get_content(metadata_mode=MetadataMode.EMBED)) print(node.get_content(metadata_mode=MetadataMode.LLM)) ``` This is the highest-value five-line diagnostic in a LlamaIndex ingestion review, and hardly anyone runs it. `MetadataMode.ALL` shows everything, `MetadataMode.NONE` shows bare text. Printing the EMBED rendering of a handful of real nodes tells you immediately how much of each chunk's embedding budget is going to bookkeeping. ## Where to set it Set the exclusion lists on the **Document**, at ingestion time. Node parsing copies metadata and the exclusion lists onto every node produced from that Document, so one assignment covers the whole file. Setting them per node after parsing works but is error-prone and easy to apply inconsistently across sources. If a metadata extractor adds keys later — a generated title, a summary, extracted questions — decide the visibility of those keys explicitly too. An extracted summary is usually worth embedding; extraction bookkeeping is not. ## Second-order effects - **Token budget.** Metadata counts toward the chunk's tokens in the prompt. Ten retrieved chunks each carrying five metadata lines is a real slice of the context window spent on file paths. - **Reranking.** A cross-encoder reranker also scores the rendered text, so metadata noise degrades the second stage as well as the first. - **Re-embedding.** Changing what is embedded changes the vectors. Adjusting exclusion lists on an existing corpus means re-embedding it, and mixing old and new renderings in one store makes similarity comparisons quietly inconsistent — decide the policy before the first large ingest. - **Leakage.** Anything LLM-visible can be reproduced in an answer. Internal paths, employee ids and tenant identifiers belong in `excluded_llm_metadata_keys` for reasons that are about disclosure, not retrieval quality. ## The rule of thumb Ask of each key: would a human searching for this passage type these words? If yes, embed it. Would the model need it to answer or to cite? If yes, show it. Otherwise store it, filter on it, and hide it from both.

  • If a key is excluded from embedding, can you still filter retrieval on it?
    Yes. Exclusion only affects the rendered text passed to the embedding model or the LLM; the key stays in the node's stored metadata dict, and metadata filters at query time read that dict. This is exactly how tenant or permission keys are handled — fully invisible to both models, fully available as a hard filter.
  • Which metadata is actually worth embedding?
    Keys a searcher would plausibly phrase in their query: a document or section title, a product name, a jurisdiction, a version. Those add discriminative signal. File paths, byte sizes, ingestion timestamps and internal ids add none and dilute short chunks, so they belong in excluded_embed_metadata_keys even when you keep them for filtering or citation.
  • You change the exclusion lists on a corpus that is already indexed. What must happen?
    Re-embed it. The vectors were computed from the old rendering, so leaving them in place means the store mixes two different text conventions and similarity comparisons become inconsistent. Treat the embedded-text recipe as part of the index's identity, decide it before the first large ingest, and version it so you know when a re-embed is due.

saying these in an interview costs you the question

  • Believes only node.text is embedded
  • Thinks excluding a key deletes it from storage
  • Leaves file paths and sizes in the embedded text
  • Assumes metadata does not consume prompt tokens
  • Changes exclusion lists without re-embedding the corpus

context