What does Haystack's ComponentTool wrap, and where does its schema come from?
answer
- a component becomes a callable tool
- introspection, not hand-written JSON Schema
- type hints on run() do the work
- the description is what the model reads
- the result must reach the model as text
basics
~20 sComponentTool turns any Haystack component into a tool the model can call. It derives the JSON parameter schema from the component's run() signature and type hints, uses the docstring or an explicit description as the tool description, and converts the component's output to a string for the model.
solid answer
~50 s`ComponentTool(component=..., name=..., description=...)` adapts a Haystack component to the `Tool` interface, so a retriever, a web-search component or a custom `@component` becomes something a chat generator can invoke. The parameter schema handed to the model is generated by inspecting the component's `run()` signature and type hints — including Pydantic models and dataclasses — so you do not hand-write JSON Schema; you can still pass an explicit `parameters` dict to override it. `name` and `description` default from the component, but the description is what the model actually reasons over, so writing it deliberately is the single highest-leverage thing you do. Because the model can only receive text, the component's output dict is converted to a string; `outputs_to_string` lets you supply a handler that renders, say, retrieved `Document` objects as readable snippets instead of a raw dump. This is why Haystack teams rarely write bespoke tool functions: the pipeline component they already test is the tool.
code
python · 17 linesfrom haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.components.retrievers.in_memory import InMemoryBM25Retriever
from haystack.tools import ComponentTool
store = InMemoryDocumentStore()
def render(documents):
return "\n\n".join(d.content for d in documents)
search_tool = ComponentTool(
component=InMemoryBM25Retriever(document_store=store),
name="search_docs",
description="Search the internal handbook for passages answering a question.",
outputs_to_string={"source": "documents", "handler": render},
)go deeper
Know that ComponentTool makes an existing Haystack component callable by the model, and that its argument schema comes from the component's run() signature rather than hand-written JSON.
Explain schema derivation from type hints, the role of name and description, the explicit parameters override, and the fact that output is stringified before the model ever sees it.
Show that you tune outputs_to_string for token cost and readability, keep descriptions written for tool selection, and understand warm_up and serialization implications of wrapping a heavy component.
Own the reuse argument: the same tested component serves both deterministic pipelines and agents, so tool surface area is governed by component design standards — typed run signatures, meaningful docstrings, serializable configuration.
## The idea Haystack already models units of work as components: a retriever, a splitter, a web search, a custom `@component` you wrote. A tool, from the model's point of view, is also a unit of work with a name, a description and a typed argument list. `ComponentTool` is the adapter between the two, so the same tested object serves both a deterministic pipeline and an agent. ``` search_tool = ComponentTool( component=InMemoryBM25Retriever(document_store=store), name="search_docs", description="Search the internal knowledge base and return matching documents.", ) ``` ## Where the schema comes from The JSON parameter schema advertised to the model is generated by introspecting the wrapped component's `run()` method: parameter names, type hints, defaults and whether each is required. Nested structures are handled — dataclasses and Pydantic models expand into object schemas — so a component whose `run()` takes `filters: dict[str, Any]` or a typed config object still produces a usable schema. Two consequences follow. First, **type hints are load-bearing**: an untyped or `Any`-typed parameter gives the model no guidance and produces vague or wrong arguments. Second, if the generated schema is not what you want — you want to hide a parameter, rename it, or constrain it with an enum — pass `parameters` explicitly and it takes precedence. You can also mark a parameter as *not* the model's business: `inputs_from_state` binds an argument to a value carried in the agent's state, and the framework supplies it at invocation time rather than asking the model to guess it. ## Name and description `name` defaults to a snake_case form of the component's class name, and `description` falls back to the component's docstring. Both defaults are conveniences, not recommendations. The description is the only prose the model sees when deciding whether this tool is the right one, so it should say what the tool *is for* and when to reach for it, not what class it wraps. Wrong-tool-selection failures in an agent are, more often than not, description failures. Names matter mechanically too: the invoker resolves calls by name, names must be unique within the tool list, and providers constrain the allowed character set. ## Getting the output back to the model A component returns a dict of Python objects — `{"documents": [Document, ...]}`. The model can only read text. `ComponentTool` therefore converts the result before it becomes a tool message. By default it serializes the output structure; `outputs_to_string` lets you take control with `{"source": "documents", "handler": my_render_fn}` so that, for instance, ten `Document` objects become ten titled snippets instead of a wall of JSON containing embeddings and scores. This is not cosmetic: the serialized result lands in the prompt of every subsequent step, so a fat default rendering multiplies cost across the whole run. Separately, `outputs_to_state` can route parts of the same result into the agent's state, so structured data survives the run without being squeezed through text. ## Warm-up and lifecycle Components that need `warm_up()` — model-loading ones especially — are warmed as part of the agent's own `warm_up()`, so an agent placed in a pipeline gets its tools warmed when the pipeline warms. Building a `ComponentTool` around a heavyweight component still means that component is instantiated once and shared across the agent's steps, not rebuilt per call. ## Serialization `ComponentTool` supports `to_dict`/`from_dict`, which is what allows an agent that uses it to survive a pipeline `to_dict()`/YAML round trip — provided the wrapped component is itself serializable. A tool built from a closure over a live connection object is where that breaks. ## When not to use it If the work is a plain function with no component lifecycle — a unit conversion, a date lookup — the `@tool` decorator on the function is lighter and produces the same schema-from-signature behaviour. `ComponentTool` earns its place when the callee is genuinely a component: it has `warm_up`, configuration, serialization, or you want the exact same object used elsewhere in your pipelines.
- Why is outputs_to_string worth configuring rather than accepting the default?Because the stringified tool result is appended to the conversation and re-sent on every subsequent step. A default serialization of retrieved documents can carry scores, ids and metadata the model does not need, inflating every later prompt. A handler that renders just title and text cuts token cost across the whole run and usually improves the model's reasoning as well.
- What happens if the wrapped component's run() parameters are untyped?The generated schema degrades — the model gets a parameter with no meaningful type, so it guesses, and you see malformed-argument failures. Either add proper type hints to the component's run(), or pass an explicit `parameters` schema to ComponentTool. Typed run signatures are the cheapest reliability win in an agent setup.
- When would you write a plain @tool function instead of a ComponentTool?When the callee is just a function with no lifecycle: no warm_up, no configuration to serialize, no reuse elsewhere. `@tool` derives the same schema from the signature and docstring with far less ceremony. Reach for ComponentTool when you want the very component your pipelines already use, with its warm-up and serialization behaviour intact.
- How do you keep the model from having to supply an argument the agent already knows?Bind it with `inputs_from_state` so the value is taken from the agent's state at invocation time instead of being requested from the model. This is how you inject things like a tenant id, a filter, or documents produced by an earlier step — values the model should not be inventing and should not even see as a parameter.
saying these in an interview costs you the question
- Hand-writing JSON Schema when the run() signature already provides it
- Leaving the auto-derived description and blaming the model for wrong tool picks
- Assuming raw Python objects reach the model unserialized
- Ignoring that stringified tool output is re-sent on every later step
- Wrapping an unserializable component and expecting pipeline YAML to round-trip