How do you design document metadata in Haystack for tenant-filtered retrieval?
answer
- Filters see chunks, not files
- Attach before the splitter or never
- Stable and low-cardinality wins
- Denormalised permissions are a rewrite bill
- Scope belongs in the outermost AND
basics
~20 sAttach the scoping keys at conversion time so DocumentSplitter copies them into every chunk, keep them low-cardinality scalars the store can index, and make the tenant condition a mandatory outermost AND assembled by shared code rather than by each caller.
solid answer
~50 sFiltering in Haystack happens on the stored `Document`, and retrieval returns chunks — so every field you will ever filter on has to live on the **chunk**. Converters accept a `meta` argument (one dict, or a list aligned with `sources`), and `DocumentSplitter` copies the parent's meta into every split while adding `source_id`, `split_id` and `page_number`. That ordering is the design rule: scope metadata is attached before splitting or it never arrives. Prefer stable, low-cardinality scalars — `tenant_id`, `source`, `doc_type`, a date — because the backing store indexes them and because anything volatile becomes a re-indexing obligation. Permissions are the hard case: denormalising an ACL into chunk metadata makes filtering trivial but turns every permission change into a rewrite of every affected chunk, so fast-changing entitlements usually belong in a resolved-at-query-time filter instead. Build the filter dict in one helper that always wraps caller conditions in an `AND` with the tenant condition, so no query path can widen the scope by omission.
go deeper
Know that filters run against the chunks in the store, and that metadata reaches chunks because DocumentSplitter copies the parent document's meta into every split.
Explain the ordering rule — attach at conversion time, before splitting — and why patching metadata onto stored chunks is not possible when ids are hashes of content and meta.
Show that you choose filterable fields deliberately: stable, low-cardinality scalars, a durable source key for deletion, and a mandatory scope condition built in shared code rather than per call site.
Own the coupled decisions: denormalised versus query-time permission resolution, shared versus per-tenant indexes, and the fact that both are effectively fixed once the corpus is embedded. Be able to defend the write-amplification and recall tradeoffs with numbers.
## The rule that follows from the pipeline Haystack's filters evaluate against a `Document`, and after an indexing pipeline the documents in the store are chunks. So the design question is not "what metadata does this file have" but "what metadata does each chunk of this file need to carry". The mechanism decides the timing. Converters take `meta` — a single dict applied to all sources, or a list aligned positionally with `sources` for per-file values. `DocumentCleaner` preserves meta. `DocumentSplitter` copies the parent's meta into every split and adds `source_id`, `split_id`, `split_idx_start` and `page_number`. Therefore: **anything attached before the splitter reaches every chunk; anything attached after it does not.** Teams that discover a missing filter field after go-live are always re-running the whole pipeline, because there is no supported way to patch metadata onto stored chunks — a write replaces a document wholesale, and the id is a hash of content plus meta, so an "update" is a delete and an insert. ## What belongs in metadata The fields that earn their place are the ones queries must be scoped by, plus the ones you need for lifecycle management: - **Scope**: `tenant_id`, workspace, project, language, region. - **Provenance**: a stable source key (`file_path` is provided free by converters, but a durable URI or CMS id survives file moves), plus `source_id`/`split_id` from the splitter for reassembling neighbours. - **Selection**: `doc_type`, product, version, effective date — the dimensions users actually narrow by. - **Lifecycle**: enough to delete a source's chunks in one filtered call. Prefer scalars and short lists over nested objects: filter field paths are flat (`meta.tenant_id`), and support for deeply nested meta paths varies by backend. Prefer low cardinality: a store can index and filter `doc_type` cheaply; a per-document UUID is only ever an equality lookup. And prefer stable: every mutable field is a future re-index. ## The permission problem Two strategies, and the choice is the actual interview content. **Denormalise entitlements into the chunk.** Store `allowed_groups` on every chunk and filter with `in`. Retrieval is a single filtered query, correct by construction, and fast. The cost is write amplification: reassigning one document's permissions rewrites all its chunks, and a group membership change that affects thousands of documents is a bulk re-index. Acceptable when permissions are coarse and change rarely — a fixed set of departments, a public/internal/confidential classification. **Resolve at query time.** Keep only stable identifiers on the chunk (owning space, classification) and translate the caller's identity into a filter over those identifiers at request time, using your authorization service as the source of truth. Permission changes take effect immediately, with no re-index. The cost is that the filter's cardinality can explode — a user with access to ten thousand spaces produces an unusable `in` list — which pushes you toward coarser groupings than the permission model natively has. Most production systems land on a hybrid: coarse, slow-moving scopes denormalised into metadata for the store to filter on, and a fine-grained check applied to the retrieved set before anything reaches the generator. That belt-and-braces arrangement also protects against the failure mode nobody plans for — a store or filter bug returning too much. ## One tenant per index, or one shared index? A shared index with a `tenant_id` filter is operationally simple, keeps one embedding model and one set of dashboards, and scales to many small tenants. Per-tenant indexes give hard isolation, per-tenant deletion that is a single drop, and independent re-index schedules, at the cost of many small indexes and per-tenant operational surface. The usual shape is a shared index by default with the largest or most regulated tenants promoted to their own — but the decision must be made before the corpus exists, because migrating between the two is a full re-embed either way. ## Making the filter unforgettable A mandatory scope that each caller assembles by hand will eventually be omitted. Wrap it: one function that takes the caller's optional conditions and returns `{"operator": "AND", "conditions": [tenant_condition, *caller_conditions]}` so the scope is structurally outermost and cannot be widened by anything a caller passes. Because retrievers also accept a constructor-level `filters` default that run-time input overrides, do not rely on that default for security — a run-time filter replaces it, which is exactly the silent widening you were trying to prevent. Then test it as a security control: a case asserting that a query issued as tenant A never returns a tenant B chunk, running against the real store rather than the in-memory one, since filter edge cases are translated per backend. ## Verify, do not assume, how filters interact with vector search Stores differ in how strictly a metadata filter is applied relative to the approximate vector search — whether candidates are constrained up front or filtered after the fact — and that difference shows up as reduced recall for narrow filters, not as an error. It is a property of the backing store, so measure it on the store you actually deploy: run a query with a very selective filter and check you get the number of results you expected rather than a short list.
- Why must scope metadata be attached before the splitter?Because DocumentSplitter produces new Documents and copies the parent's meta into each split. Anything attached afterwards lives only on objects the store never sees. There is no patch path either: a write replaces a whole document and the auto id hashes content plus meta, so adding a field to stored chunks means deleting and re-indexing them.
- When is denormalising an ACL into chunk metadata the wrong call?When entitlements change often or are fine-grained per user. Every permission change becomes a rewrite of every affected chunk, and a group membership change can turn into a bulk re-index. Coarse, slow-moving classifications denormalise well; anything that moves at user-management speed should be resolved into a filter at query time instead.
- What are the tradeoffs of one index per tenant versus a shared filtered index?Per-tenant indexes give hard isolation, single-operation tenant deletion and independent re-index schedules, but multiply operational surface and fragment small corpora. A shared index with a tenant filter is simpler and cheaper but makes correctness depend on every query carrying the filter. Common practice is shared by default, with large or regulated tenants promoted — decided up front, since migrating means re-embedding.
- Why is a retriever's constructor-level filters default a poor security boundary?Because a run-time filters value overrides it for that run. Any caller passing their own filters silently drops the default scope, which is exactly the widening you were guarding against. Build the filter in one shared helper that always nests caller conditions inside an AND with the mandatory scope, and test that a query for one tenant never returns another's chunks.
saying these in an interview costs you the question
- Attaching filter metadata after the splitter and expecting chunks to have it
- Assuming stored chunk metadata can be updated in place
- Denormalising fast-changing per-user permissions into every chunk
- Relying on each caller to remember the tenant filter
- Assuming a narrow filter never costs vector-search recall