skip to content

When must a custom Haystack component implement to_dict and from_dict?

level: seniorimportance: should knowfreq 30%

answer

  1. Default works until JSON cannot express it
  2. One-to-one init parameter to attribute
  3. Convert out, convert back, then delegate
  4. Type key is an import path
  5. Configuration only, never warmed state

basics

~20 s

Only when default serialization cannot round-trip your constructor arguments. Haystack's default maps each init parameter to a same-named attribute and needs JSON-friendly values, so sets, callables, custom objects and secrets require explicit conversion in to_dict and reconstruction in from_dict.

solid answer

~40 s

A component serializes to `{"type": "<module path>.<ClassName>", "init_parameters": {...}}`. If you write no `to_dict`, Haystack falls back to `default_to_dict`, which has two requirements: a 1:1 mapping between `__init__` parameter names and attributes of the same name, and values JSON can express. Satisfy both and you need nothing. The moment an init argument is a `set`, a callable, a device handle, a `Secret` or any custom object, you implement `to_dict()` to convert it to a serializable form and `from_dict()` to convert it back before delegating to `default_from_dict(cls, data)`. Two traps follow. The `type` string is an import path, so a class defined in a notebook or under `__main__` serializes fine and then fails to load elsewhere. And `to_dict` must report configuration, never run-time state, so a warmed instance and a fresh one serialize identically.

code

python · 24 lines
python
from haystack import component, default_from_dict, default_to_dict


@component
class TagFilter:
    def __init__(self, allowed: set[str]):
        self.allowed = allowed

    def to_dict(self) -> dict:
        return default_to_dict(self, allowed=sorted(self.allowed))

    @classmethod
    def from_dict(cls, data: dict) -> "TagFilter":
        data["init_parameters"]["allowed"] = set(data["init_parameters"]["allowed"])
        return default_from_dict(cls, data)

    @component.output_types(kept=list[str])
    def run(self, tags: list[str]):
        return {"kept": [tag for tag in tags if tag in self.allowed]}


original = TagFilter(allowed={"rag", "search"})
print(original.to_dict())
print(TagFilter.from_dict(original.to_dict()).run(tags=["rag", "other"]))

go deeper

for a junior

Know that a component can be saved as a dict of its type and its constructor arguments, and that simple components with plain string and number configuration need no extra code for that to work.

for a middle

Explain the two requirements of the default path — matching init parameter and attribute names, and JSON-expressible values — and show the convert-then-delegate shape of to_dict and from_dict.

for a senior

Name the traps you have hit: import paths that do not exist outside a notebook, callables and secrets needing their own helpers, non-deterministic set ordering, and warmed state leaking into the serialized form. Test the round trip.

for a principal

Own the consequence for the platform: once pipeline definitions are versioned artifacts, component class paths and init-parameter names become a compatibility contract, and renaming a component is a migration rather than a refactor.

## What serialization is for Haystack lets you turn a pipeline into a dict or a YAML file and load it back somewhere else. That is what makes a pipeline a reviewable artifact rather than a pile of glue code: you can diff it, promote it between environments, or let a non-Python tool render it. Component serialization is the unit that makes it work — a pipeline's serialized form is essentially its components' serialized forms plus the connections between them. ## The serialized shape Every component becomes: ``` { "type": "my_package.filters.TagFilter", "init_parameters": {"allowed": ["a", "b"]} } ``` `type` is the fully qualified import path of the class. `init_parameters` is what will be passed to `__init__` on the way back. Note what is *not* there: no run-time state, no loaded model, no outputs. Serialization captures how the component was configured, and reconstruction rebuilds it from scratch. ## The default path and its two requirements If you implement nothing, Haystack uses `default_to_dict`/`default_from_dict`. For that to work: 1. **1:1 naming.** Each `__init__` parameter must be stored on an attribute of the same name. If your constructor takes `allowed` and you store `self._allowed`, discovery breaks. 2. **JSON-expressible values.** Strings, numbers, booleans, lists, dicts and `None` round-trip. Everything else does not. A component whose configuration is a model name and a `top_k` needs no serialization code at all, and most custom components are in that category. This is worth saying explicitly in an interview, because the reflexive answer — "always implement them" — is wrong and adds noise to every component you write. ## When you must intervene The common triggers: - **A `set`.** JSON has no set. Convert to a sorted list on the way out, back to a set on the way in. Sorting is not cosmetic: it makes the serialized form deterministic, so a re-export does not produce a spurious diff. - **A callable.** A streaming callback or a filter function cannot be embedded. Haystack ships `serialize_callable`/`deserialize_callable`, which record the callable's import path and resolve it again. A lambda or a closure has no import path and cannot be serialized at all — that is a real design constraint on your component's API. - **A `Secret`.** Credentials serialize as a reference; on the way back, `deserialize_secrets_inplace(data["init_parameters"], keys=["api_key"])` rebuilds the object before construction. - **Nested components or tools.** Anything that itself has a serialized form has to be serialized recursively rather than embedded raw. - **Custom objects generally** — a config dataclass, a device descriptor, an enum — need a chosen wire form and a conversion each way. ## The pattern ``` def to_dict(self): return default_to_dict(self, allowed=sorted(self.allowed)) @classmethod def from_dict(cls, data): data["init_parameters"]["allowed"] = set(data["init_parameters"]["allowed"]) return default_from_dict(cls, data) ``` Note the asymmetry that trips people: `default_to_dict(self, **params)` takes the parameters as keyword arguments, while `from_dict` mutates `data["init_parameters"]` in place and then hands the *whole* `data` dict to `default_from_dict`. Reaching for `data["allowed"]` instead of `data["init_parameters"]["allowed"]` is the single most common mistake in this code. ## The import-path trap Because `type` is an import path, the loading process must be able to import that exact path. A component class defined in a Jupyter notebook, in a `__main__` script, or inside a function body will serialize happily and then fail to load in a worker, a test process, or another service. The fix is not clever — move the class into an installable module — but the failure is confusing enough that it is worth naming before it costs an afternoon. A related discipline: renaming or moving a component class breaks every stored pipeline file that references it. Once definitions are checked in, class paths are part of your contract with those files. ## Configuration, not state `to_dict` must describe how to rebuild the component, which means it reports the values it was constructed with. If a component loads a model in `warm_up()` and stores it on an attribute, that attribute must not appear in the serialized form — and if `warm_up()` mutates one of the attributes `to_dict` reads, a warmed instance will serialize differently from a cold one and your round-trip stops being faithful. Keeping the serialized surface identical to the constructed surface is the invariant to hold, and a good test for a custom component is exactly that: construct, serialize, deserialize, serialize again, and assert the two dicts are equal.

  • Why does a component class defined in a notebook break pipeline loading?
    The serialized `type` is a fully qualified import path, and a class living in a notebook or under `__main__` has no path another process can import. Serialization succeeds, then loading fails in a worker, a test runner or a service with an import error. The fix is to move the class into an installable module — and to remember that renaming or relocating a component afterwards invalidates every stored pipeline file referencing it.
  • How do you serialize a component that takes a callback function as an init argument?
    Use Haystack's `serialize_callable`/`deserialize_callable`, which store the callable's import path and resolve it again on load. That constrains your API: a module-level function serializes, a lambda or a closure has no import path and cannot. If you need behaviour that is genuinely per-instance, take a name or an enum in the constructor and map it to a function internally, so the serialized form stays a plain string.
  • What is a good test for a custom component's serialization?
    A round-trip equality test: construct the component, call `to_dict()`, rebuild with `from_dict()`, serialize again, and assert both dicts are equal. It catches asymmetric conversions, missing parameters, non-deterministic ordering from sets, and — if you warm the component before the second serialization — any leakage of run-time state into what should be pure configuration.

saying these in an interview costs you the question

  • Assumes every custom component needs explicit to_dict
  • Stores init arguments under differently named private attributes
  • Reads data["key"] instead of data["init_parameters"]["key"] in from_dict
  • Puts loaded models or run-time state into to_dict
  • Expects a class defined in __main__ to reload elsewhere

context