In Haystack retrievers, how do init-time top_k and filters differ from run-time ones?
answer
- one place is configuration, one is per request
- serialised pipelines only carry one of them
- depth simply overrides
- filters have a policy, depth does not
- replace versus merge, and who is trusted
basics
~20 sConstructor values are the component's defaults and are captured in its serialised form. Values passed in Pipeline.run's per-component input dict apply to that run only. top_k always overrides, while filters are combined or replaced according to the retriever's filter_policy setting.
solid answer
~50 sHaystack retrievers accept `top_k` and `filters` in two places, and they behave differently. What you pass to the constructor becomes the component's default and travels with `to_dict`, so it is what a serialised pipeline or YAML file carries. What you pass at run time — `pipeline.run({"retriever": {"top_k": 3, "filters": {...}}})`, or from an upstream component connected to those sockets — applies to that invocation only. For `top_k` the run-time value simply wins. For `filters` the outcome depends on `filter_policy`: the default `FilterPolicy.REPLACE` discards the constructor filters entirely, while `FilterPolicy.MERGE` combines them with the run-time ones. That distinction matters for anything security-shaped: a tenant or ACL restriction set at construction under REPLACE vanishes the moment a caller supplies its own filter, so make such restrictions MERGE, or apply them in a component the caller cannot address.
go deeper
Know that top_k and filters can be set when you build the retriever or passed in the run input dict, and that the run-time value takes effect for that call.
Explain that only constructor arguments are serialised, that top_k simply overrides, and that filters follow filter_policy with REPLACE as the default and MERGE as the combining option.
Demonstrate the security angle: a scoping filter under REPLACE is not a boundary, both legs of a hybrid pipeline need the same treatment, and the negative case belongs in a test rather than in an assumption.
Own where trust sits in the request path. Decide as policy that authorization scoping is constructed server-side per request, never accepted from the entry-point input dict, and that serialised pipeline artefacts reflect real deployed configuration.
## Two places to configure the same thing Every Haystack retriever takes `top_k` and `filters` in its constructor and also exposes them as run-time inputs. The design is deliberate: the constructor holds the pipeline's stable configuration, and the run-time dict holds per-request variation. ```python retriever = InMemoryBM25Retriever(document_store=store, top_k=10, filters={"field": "meta.lang", "operator": "==", "value": "en"}) pipeline.run({"retriever": {"query": q, "top_k": 3}}) ``` ## Why the distinction is not cosmetic **Serialisation.** Haystack components implement `to_dict`/`from_dict`, and a pipeline can be dumped to YAML and reloaded. Only constructor arguments are captured there. Anything you supply per run is, by definition, absent from the serialised pipeline — so a colleague who loads the YAML sees `top_k=10` and has no way of knowing that the service always overrides it with 3. If a value is really fixed for your deployment, put it in the constructor so the artefact tells the truth. **Testability.** Constructor defaults let you run the pipeline with a minimal input dict in tests. Run-time overrides let one deployed pipeline serve several callers with different budgets without rebuilding components. **Wiring.** Run-time inputs need not come from the caller at all: because they are ordinary input sockets, an upstream component can be connected to `retriever.filters`, which is how you build a pipeline that derives filters from the query (for example, a router or a custom component that extracts a date range and emits a filter dict). ## `top_k` is a plain override Nothing subtle happens here: if a run-time `top_k` is present it is used for that run, otherwise the constructor value applies. Note that `top_k` is a *retrieval depth*, not a quality control — a large value in a hybrid pipeline feeds a bigger candidate pool to the joiner and ranker, and the cost lands on the ranker, which is usually the expensive component. ## `filters` and `filter_policy` Filters are the interesting case, because "the caller supplied filters" and "the pipeline was configured with filters" are different intentions that must be reconciled. Retrievers take a `filter_policy` argument for exactly this: - **`FilterPolicy.REPLACE`** (the default) — run-time filters, when present, replace the constructor filters wholesale. Absent run-time filters, the constructor's apply. - **`FilterPolicy.MERGE`** — the two sets are combined, so the constructor filters continue to constrain the result no matter what the caller sends. The default is REPLACE because the common case is a pipeline with a default scope that callers legitimately re-aim. The trap is using a constructor filter as a security boundary. Under REPLACE, a `tenant_id` restriction baked into the retriever disappears the instant a request supplies any filter of its own — and the request that supplies it may be a user-authored query in an agentic system. Either set MERGE, or, more robustly, keep the authorization filter out of caller-addressable inputs entirely: construct a per-request retriever inside a trusted layer, or compose the filter in a component the caller cannot pass inputs to. Treat the entry-point dict as untrusted input, because in an API-fronted pipeline it is. ## Practical guidance 1. Put the value in the constructor when it is a property of the deployment; pass it at run time when it is a property of the request. 2. Never rely on REPLACE semantics to preserve a scoping filter — assert the behaviour in a test that supplies a caller filter and checks the restriction still holds. 3. If several retrievers must share a scope (both legs of a hybrid pipeline, say), remember that each has its own filters and its own policy. A filter applied to only one leg lets restricted documents in through the other, and the joiner will happily fuse them. 4. When filters come from an upstream component, keep the component that builds them small and unit-test it directly; a malformed filter dict is a run-time error deep inside the store, not a wiring error.
- Why is a constructor-level tenant filter risky under the default filter_policy?Because the default is REPLACE: as soon as a run supplies any filters, the constructor's are discarded, taking the tenant restriction with them. In an API-fronted or agent-driven pipeline the run-time dict is attacker-influenced input, so the scope silently disappears. Use FilterPolicy.MERGE, or build the retriever per request inside a trusted layer so the restriction is never caller-addressable.
- In a hybrid pipeline, what do you have to remember about filters on the two legs?Each retriever carries its own filters and its own filter_policy — there is no pipeline-wide scope. If you filter the dense leg but not the BM25 leg, out-of-scope documents enter through the lexical branch and the joiner fuses them in as normal. Apply the same filter and policy to every retriever, and test the negative case rather than assuming symmetry.
- Why might you connect an upstream component to a retriever's filters socket instead of passing filters from the caller?Because filters are an ordinary input socket, a component can compute them from the query — extracting a date range, a product line, or a resolved tenant — and the logic stays inside the pipeline where it is versioned, serialised and testable. It also keeps the derivation out of the caller's hands, which matters when the filter encodes a scope you do not want the caller choosing.
saying these in an interview costs you the question
- Thinks constructor filters always survive a run-time filter
- Assumes filter_policy affects top_k as well
- Believes run-time overrides are captured in the serialised pipeline
- Applies a scoping filter to only one leg of a hybrid pipeline
- Treats top_k as a relevance threshold rather than a retrieval depth