When must a custom Haystack component implement to_dict and from_dict?
answer
- Default works until JSON cannot express it
- One-to-one init parameter to attribute
- Convert out, convert back, then delegate
- Type key is an import path
- Configuration only, never warmed state
basics
~20 sOnly when default serialization cannot round-trip your constructor arguments. Haystack's default maps each init parameter to a same-named attribute and needs JSON-friendly values, so sets, callables, custom objects and secrets require explicit conversion in to_dict and reconstruction in from_dict.
solid answer
~40 sA component serializes to `{"type": "<module path>.<ClassName>", "init_parameters": {...}}`. If you write no `to_dict`, Haystack falls back to `default_to_dict`, which has two requirements: a 1:1 mapping between `__init__` parameter names and attributes of the same name, and values JSON can express. Satisfy both and you need nothing. The moment an init argument is a `set`, a callable, a device handle, a `Secret` or any custom object, you implement `to_dict()` to convert it to a serializable form and `from_dict()` to convert it back before delegating to `default_from_dict(cls, data)`. Two traps follow. The `type` string is an import path, so a class defined in a notebook or under `__main__` serializes fine and then fails to load elsewhere. And `to_dict` must report configuration, never run-time state, so a warmed instance and a fresh one serialize identically.
code
python · 24 linesfrom haystack import component, default_from_dict, default_to_dict
@component
class TagFilter:
def __init__(self, allowed: set[str]):
self.allowed = allowed
def to_dict(self) -> dict:
return default_to_dict(self, allowed=sorted(self.allowed))
@classmethod
def from_dict(cls, data: dict) -> "TagFilter":
data["init_parameters"]["allowed"] = set(data["init_parameters"]["allowed"])
return default_from_dict(cls, data)
@component.output_types(kept=list[str])
def run(self, tags: list[str]):
return {"kept": [tag for tag in tags if tag in self.allowed]}
original = TagFilter(allowed={"rag", "search"})
print(original.to_dict())
print(TagFilter.from_dict(original.to_dict()).run(tags=["rag", "other"]))go deeper
Know that a component can be saved as a dict of its type and its constructor arguments, and that simple components with plain string and number configuration need no extra code for that to work.
Explain the two requirements of the default path — matching init parameter and attribute names, and JSON-expressible values — and show the convert-then-delegate shape of to_dict and from_dict.
Name the traps you have hit: import paths that do not exist outside a notebook, callables and secrets needing their own helpers, non-deterministic set ordering, and warmed state leaking into the serialized form. Test the round trip.
Own the consequence for the platform: once pipeline definitions are versioned artifacts, component class paths and init-parameter names become a compatibility contract, and renaming a component is a migration rather than a refactor.
## What serialization is for Haystack lets you turn a pipeline into a dict or a YAML file and load it back somewhere else. That is what makes a pipeline a reviewable artifact rather than a pile of glue code: you can diff it, promote it between environments, or let a non-Python tool render it. Component serialization is the unit that makes it work — a pipeline's serialized form is essentially its components' serialized forms plus the connections between them. ## The serialized shape Every component becomes: ``` { "type": "my_package.filters.TagFilter", "init_parameters": {"allowed": ["a", "b"]} } ``` `type` is the fully qualified import path of the class. `init_parameters` is what will be passed to `__init__` on the way back. Note what is *not* there: no run-time state, no loaded model, no outputs. Serialization captures how the component was configured, and reconstruction rebuilds it from scratch. ## The default path and its two requirements If you implement nothing, Haystack uses `default_to_dict`/`default_from_dict`. For that to work: 1. **1:1 naming.** Each `__init__` parameter must be stored on an attribute of the same name. If your constructor takes `allowed` and you store `self._allowed`, discovery breaks. 2. **JSON-expressible values.** Strings, numbers, booleans, lists, dicts and `None` round-trip. Everything else does not. A component whose configuration is a model name and a `top_k` needs no serialization code at all, and most custom components are in that category. This is worth saying explicitly in an interview, because the reflexive answer — "always implement them" — is wrong and adds noise to every component you write. ## When you must intervene The common triggers: - **A `set`.** JSON has no set. Convert to a sorted list on the way out, back to a set on the way in. Sorting is not cosmetic: it makes the serialized form deterministic, so a re-export does not produce a spurious diff. - **A callable.** A streaming callback or a filter function cannot be embedded. Haystack ships `serialize_callable`/`deserialize_callable`, which record the callable's import path and resolve it again. A lambda or a closure has no import path and cannot be serialized at all — that is a real design constraint on your component's API. - **A `Secret`.** Credentials serialize as a reference; on the way back, `deserialize_secrets_inplace(data["init_parameters"], keys=["api_key"])` rebuilds the object before construction. - **Nested components or tools.** Anything that itself has a serialized form has to be serialized recursively rather than embedded raw. - **Custom objects generally** — a config dataclass, a device descriptor, an enum — need a chosen wire form and a conversion each way. ## The pattern ``` def to_dict(self): return default_to_dict(self, allowed=sorted(self.allowed)) @classmethod def from_dict(cls, data): data["init_parameters"]["allowed"] = set(data["init_parameters"]["allowed"]) return default_from_dict(cls, data) ``` Note the asymmetry that trips people: `default_to_dict(self, **params)` takes the parameters as keyword arguments, while `from_dict` mutates `data["init_parameters"]` in place and then hands the *whole* `data` dict to `default_from_dict`. Reaching for `data["allowed"]` instead of `data["init_parameters"]["allowed"]` is the single most common mistake in this code. ## The import-path trap Because `type` is an import path, the loading process must be able to import that exact path. A component class defined in a Jupyter notebook, in a `__main__` script, or inside a function body will serialize happily and then fail to load in a worker, a test process, or another service. The fix is not clever — move the class into an installable module — but the failure is confusing enough that it is worth naming before it costs an afternoon. A related discipline: renaming or moving a component class breaks every stored pipeline file that references it. Once definitions are checked in, class paths are part of your contract with those files. ## Configuration, not state `to_dict` must describe how to rebuild the component, which means it reports the values it was constructed with. If a component loads a model in `warm_up()` and stores it on an attribute, that attribute must not appear in the serialized form — and if `warm_up()` mutates one of the attributes `to_dict` reads, a warmed instance will serialize differently from a cold one and your round-trip stops being faithful. Keeping the serialized surface identical to the constructed surface is the invariant to hold, and a good test for a custom component is exactly that: construct, serialize, deserialize, serialize again, and assert the two dicts are equal.
- Why does a component class defined in a notebook break pipeline loading?The serialized `type` is a fully qualified import path, and a class living in a notebook or under `__main__` has no path another process can import. Serialization succeeds, then loading fails in a worker, a test runner or a service with an import error. The fix is to move the class into an installable module — and to remember that renaming or relocating a component afterwards invalidates every stored pipeline file referencing it.
- How do you serialize a component that takes a callback function as an init argument?Use Haystack's `serialize_callable`/`deserialize_callable`, which store the callable's import path and resolve it again on load. That constrains your API: a module-level function serializes, a lambda or a closure has no import path and cannot. If you need behaviour that is genuinely per-instance, take a name or an enum in the constructor and map it to a function internally, so the serialized form stays a plain string.
- What is a good test for a custom component's serialization?A round-trip equality test: construct the component, call `to_dict()`, rebuild with `from_dict()`, serialize again, and assert both dicts are equal. It catches asymmetric conversions, missing parameters, non-deterministic ordering from sets, and — if you warm the component before the second serialization — any leakage of run-time state into what should be pure configuration.
saying these in an interview costs you the question
- Assumes every custom component needs explicit to_dict
- Stores init arguments under differently named private attributes
- Reads data["key"] instead of data["init_parameters"]["key"] in from_dict
- Puts loaded models or run-time state into to_dict
- Expects a class defined in __main__ to reload elsewhere