skip to content

LangChain

You will learn the framework that popularised LLM application plumbing: prompts, model wrappers, memory, retrievers, tools, and the LCEL Runnable model that composes them. Interviewers ask about LangChain because it is the shared vocabulary for these apps, and they want to hear where its abstractions earn their keep and where you would drop to a raw client.

part ofAI agent & RAG frameworksoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

In LangChain, what does the @tool decorator expose to the model from a function?

level: juniorimportance: must knowfreq 80%

answer

  1. three fields the model actually sees
  2. docstring is not documentation here
  3. type hints become a schema
  4. hide runtime-only arguments from the model

basics

~20 s

The @tool decorator wraps a Python function as a Tool: the function name becomes the tool name, the docstring becomes the description, and the type hints become a JSON argument schema. The model sees only those three things.

solid answer

~50 s

`@tool` from `langchain_core.tools` converts a plain function into a `BaseTool`. By default it infers three pieces of metadata that are serialised into the model's tool specification: `name` (the function name, or a string you pass as `@tool("search_orders")`), `description` (the docstring), and `args_schema` (a Pydantic model built from the parameter names and type hints). The function body is never sent anywhere — the model only ever emits a name plus an arguments object, and your process runs the code. That means tool quality is prompt engineering: an untyped `**kwargs` signature or a missing docstring leaves the model guessing. You can tighten it by passing an explicit `args_schema`, by using `parse_docstring=True` so per-argument docstring lines become field descriptions, or by using `Field(description=...)` on a Pydantic schema. `return_direct=True` makes the agent stop and return the tool output verbatim instead of looping back to the model.

code

python · 22 lines
python
from typing import Annotated, Literal
from langchain_core.tools import tool, InjectedToolArg


@tool(parse_docstring=True)
def lookup_order(
    order_id: str,
    detail: Literal["summary", "full"] = "summary",
    tenant_id: Annotated[str, InjectedToolArg] = "",
) -> str:
    """Look up a customer order by its id.

    Args:
        order_id: The public order identifier, e.g. ORD-1234.
        detail: How much of the order to return.
    """
    return f"{order_id} ({detail}) for {tenant_id}"


print(lookup_order.name)
print(lookup_order.description)
print(lookup_order.args_schema.model_json_schema()["properties"].keys())

go deeper

for a junior

Be able to write a tool with @tool, and say plainly that the model sees only the name, the docstring and the argument schema derived from type hints.

for a middle

Explain how args_schema is inferred, when to pass an explicit Pydantic model with Field descriptions, and what parse_docstring and return_direct change.

for a senior

Show judgment about the model-facing surface: constrained enums over free strings, InjectedToolArg for anything authorisation-relevant, and content_and_artifact to keep large payloads out of the context window.

for a principal

Own tool descriptions as a versioned interface. Renaming a tool or editing a docstring changes model behaviour, so treat the catalogue as a product surface with evals, naming conventions and a cap on how many tools any one agent is offered.

## What a tool actually is A LangChain tool is an object implementing `BaseTool` with three pieces of public metadata — `name`, `description`, `args_schema` — plus a `_run` (and optionally `_arun`) implementation. Everything an agent does with tools flows through those four things. The metadata is serialised into the provider's function/tool-calling format when you call `model.bind_tools([...])`; the implementation is invoked locally when the model asks for it by name. `@tool` is the ergonomic way to produce that object from a function you already have. ## Where each field comes from - **name** — the Python function name by default. `@tool("lookup_order")` overrides it. The name is what the model emits in a tool call, so it should be a stable, descriptive identifier; renaming a function silently renames the tool and invalidates any prompt or few-shot example that referenced it. - **description** — the function's docstring. This is the single highest-leverage field: it is the only natural-language guidance the model has about *when* to pick this tool over its siblings. A docstring that says what the tool returns, what it costs, and when *not* to use it measurably reduces wrong-tool selection. - **args_schema** — inferred from the signature via Pydantic when `infer_schema=True` (the default). `def get_weather(city: str, units: str = "celsius")` becomes a schema with a required string `city` and an optional string `units`. Untyped parameters degrade to permissive schemas, and the model starts inventing shapes. With `parse_docstring=True`, a Google-style docstring's `Args:` section is parsed so each argument gets its own description in the schema — usually better than cramming argument semantics into the top-level description. ## Taking control of the schema When inference is not enough, pass an explicit Pydantic model: `@tool(args_schema=SearchInput)`. This lets you attach `Field(description=...)`, constraints (`ge`, `max_length`), enums for closed vocabularies, and defaults. Constraining an argument to a `Literal` of five valid values is far more reliable than asking the model in prose to pick one of five. For cases where the decorator is awkward — a partially applied function, a closure over a client — `StructuredTool.from_function(func=..., name=..., description=..., args_schema=...)` builds the same object explicitly. Subclassing `BaseTool` and implementing `_run` is the fullest-control option and is what you reach for when the tool needs its own state, custom error handling, or a non-trivial async path. ## Arguments the model must not see Some parameters are runtime plumbing, not model choices: a tenant id, a database session, an auth token. Annotating them with `InjectedToolArg` (`Annotated[str, InjectedToolArg]`) keeps them out of the schema sent to the model while still requiring them at invocation time. This is the correct pattern for anything security-relevant — never let the model choose the tenant it queries. `InjectedToolCallId` similarly injects the current tool-call id when a tool needs to construct its own `ToolMessage`. ## Invocation and outputs A decorated tool is a Runnable: `tool.invoke({"city": "Berlin"})` runs it with a dict of arguments, and `await tool.ainvoke(...)` runs the async path (decorating an `async def` gives you a real coroutine implementation rather than a thread-pool wrapper). Passing a whole tool-call dict instead — `tool.invoke(tool_call)` — returns a `ToolMessage` already stamped with the matching `tool_call_id`, which is exactly what the agent loop needs to feed back to the model. Two output modifiers matter. `return_direct=True` tells the agent to return this tool's output as the final answer without another model call — useful for a tool whose output is already the deliverable, dangerous when you actually wanted the model to summarise. `response_format="content_and_artifact"` lets the function return a `(content, artifact)` tuple, where only `content` goes back into the conversation and the artifact stays available to your code — the standard way to keep a 5 MB dataframe out of the prompt while still using it downstream. ## Common failure modes No docstring means an empty description and near-random tool selection. Overlapping descriptions across two tools produce oscillation between them. A schema with a free-form `str` where you meant an enum produces values you then have to defensively validate. And a tool whose docstring describes behaviour the implementation no longer has is a silent, permanent bug — the model believes the docstring.

  • How would you keep a tenant id out of the model's reach while the tool still needs it?
    Annotate the parameter with `InjectedToolArg` (`Annotated[str, InjectedToolArg]`). It is stripped from the schema the model sees, so the model cannot choose or spoof it, and it is supplied by your code at invocation time. Any authorisation-relevant argument belongs in this category — session identity, tenant, user id — because anything in the visible schema is ultimately attacker-influenced through the prompt.
  • What does return_direct=True change about the agent loop?
    With `return_direct=True`, once the tool runs the agent returns its output as the final result instead of sending the `ToolMessage` back to the model for another turn. It saves a model call and preserves the tool output verbatim, which is right for a tool whose output *is* the answer. It is wrong whenever you wanted the model to interpret, filter or combine the result, and it also means the model never gets to recover from a bad tool result.
  • When would you subclass BaseTool instead of using @tool?
    When the tool needs more than a function: injected clients or connection pools held as fields, custom `handle_tool_error` behaviour, distinct sync and async implementations, or access to the callback manager inside `_run` for progress reporting. `StructuredTool.from_function` covers the middle ground where you need explicit metadata but no custom class.

The docstring and type hints are the tool's job advert. The model never reads the code, only the advert — a vague advert gets you the wrong applicant.

saying these in an interview costs you the question

  • Thinking the model executes the function body itself
  • Leaving the docstring empty and blaming the model for wrong tool choice
  • Assuming untyped parameters still produce a strict schema
  • Putting a tenant or auth token in the visible argument schema
  • Believing the tool name can change freely without affecting behaviour

context

open as a page

In LangChain, what do a Runnable's invoke, batch and stream methods each do?

level: juniorimportance: must knowfreq 78%

basics

~20 s

invoke runs the runnable once on one input and returns the final output. batch runs many inputs, concurrently by default. stream yields output in chunks as it is produced. Each has an async twin: ainvoke, abatch, astream.

open as a page

In LangChain, what does a BaseChatMessageHistory store, and when does in-memory storage break?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A chat message history holds one conversation's ordered messages, keyed by a session id, through the messages property plus add_messages and clear. InMemoryChatMessageHistory keeps that list in process RAM, so a restart or a second replica loses the conversation.

open as a page

In LangChain, how does a ChatModel differ from the older text-completion LLM interface?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A ChatModel takes a list of role-tagged messages and returns an AIMessage object; the LLM interface takes a plain string and returns a plain string. Tool calling, structured output and multimodal content exist only on the ChatModel path.

open as a page

In LangChain, when do you use PromptTemplate versus ChatPromptTemplate?

level: juniorimportance: must knowfreq 78%

basics

~20 s

PromptTemplate formats one string with declared variables and suits plain text-completion models. ChatPromptTemplate formats an ordered list of role-tagged messages (system, human, ai) and is what you use with chat models, which is nearly always today.

open as a page

Why did LangChain v1 replace AgentExecutor with create_agent?

level: middleimportance: must knowfreq 70%

basics

~20 s

AgentExecutor was an opaque while-loop: you could not pause it, resume it, inspect its state or extend it. LangChain v1's create_agent returns a LangGraph graph, so the same loop gains persistence, human-in-the-loop pausing, streaming of intermediate steps and middleware hooks.

open as a page

How does a LangChain agent turn an AIMessage tool call into a ToolMessage?

level: middleimportance: must knowfreq 65%

basics

~20 s

The model returns an AIMessage carrying tool_calls, each with a name, an args dict and an id. The agent looks the name up in its tool list, invokes it with args, and appends a ToolMessage whose tool_call_id matches, then calls the model again.

open as a page

What does the | operator build in LangChain's LCEL, and what gets coerced?

level: middleimportance: must knowfreq 74%

basics

~20 s

The pipe operator builds a RunnableSequence — itself a Runnable — that feeds each step's output into the next. Along the way LCEL coerces plain values: a dict becomes a RunnableParallel and a callable becomes a RunnableLambda.

open as a page

In LangChain, how do buffer, summary and token-buffer conversation memories differ?

level: middleimportance: must knowfreq 60%

basics

~20 s

Buffer memory replays every turn verbatim, so fidelity is perfect and tokens grow without limit. Summary memory keeps a running LLM-written summary, so the prompt stays flat but each turn costs an extra model call and loses detail. Token-buffer memory keeps the newest messages that fit a token limit.

open as a page

How does RunnableWithMessageHistory feed prior turns into a LangChain LCEL chain?

level: middleimportance: must knowfreq 68%

basics

~20 s

RunnableWithMessageHistory wraps a runnable. On each call it uses the session id from config['configurable'] to fetch that conversation's history, prepends the stored messages to the input, invokes the inner chain, then appends the new input and output messages back to the store.

open as a page

In LangChain, when should you use with_structured_output() instead of PydanticOutputParser?

level: middleimportance: must knowfreq 72%

basics

~20 s

Prefer with_structured_output() whenever the provider supports tool calling or JSON-schema mode: the schema is enforced at the API layer. Fall back to PydanticOutputParser only for models with no native structured output, since it merely asks nicely in the prompt.

open as a page

What does MessagesPlaceholder do in a LangChain ChatPromptTemplate?

level: middleimportance: must knowfreq 66%

basics

~20 s

MessagesPlaceholder reserves a slot inside a ChatPromptTemplate that is filled at format time with a whole list of message objects — typically prior turns — instead of a formatted string, preserving each message's role and metadata.

open as a page

In LangChain, what does as_retriever(search_type=...) control on a vectorstore?

level: middleimportance: must knowfreq 66%

basics

~20 s

as_retriever() wraps a vectorstore as a retriever you call with invoke(query). search_type selects the algorithm — "similarity" (default), "mmr" for diversity, or "similarity_score_threshold" — and search_kwargs passes k, fetch_k, lambda_mult, score_threshold and store-specific metadata filters.

open as a page

How does LangChain's RecursiveCharacterTextSplitter decide where to cut a document?

level: middleimportance: must knowfreq 72%

basics

~20 s

It tries a separator list in order — blank lines, then newlines, then spaces, then bare characters — recursing into any piece still over chunk_size, then merges neighbouring pieces up to chunk_size while repeating chunk_overlap units of the previous chunk.

open as a page

In LangChain v1, how do you wire a retriever into a question-answering chain?

level: middleimportance: must knowfreq 62%

basics

~20 s

Retrieve, format, prompt: call retriever.invoke(question), join the returned Documents' page_content into a context string, and render it into a prompt template beside the question. The pre-1.0 create_retrieval_chain helper packaged that shape and belongs to the legacy compatibility surface, not to v1's slim core.

open as a page

Your LangChain chat sessions overflow the context window; how do you bound history?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Insert trim_messages from langchain_core.messages as a step before the prompt, so every invocation cuts the loaded history to a token budget. Keep the system message pinned, trim from the oldest end, and start the kept window on a human message so tool replies are never orphaned.

open as a page

In LangChain, what does a document loader return, and when do you use lazy_load()?

level: juniorimportance: should knowfreq 48%

basics

~20 s

LangChain document loaders return Document objects, each with page_content text and a metadata dict (source, page). load() builds the full list in memory; lazy_load() yields Documents one at a time, so large corpora stream into splitting and embedding.

open as a page

In LCEL, when do you use RunnableParallel versus RunnablePassthrough.assign?

level: middleimportance: should knowfreq 60%

basics

~20 s

RunnableParallel runs several runnables on the same input and returns a dict of exactly their outputs, replacing the input. RunnablePassthrough.assign runs them the same way but merges the results into the incoming dict, keeping the original keys.

open as a page

Why can LangChain's JsonOutputParser stream partial results while PydanticOutputParser cannot?

level: middleimportance: should knowfreq 42%

basics

~20 s

JsonOutputParser parses incomplete JSON on every chunk and emits progressively fuller dicts. PydanticOutputParser must construct and validate a complete model instance, and a half-finished object cannot satisfy required fields, so it only parses once the text is complete.

open as a page

What do partial variables do in a LangChain PromptTemplate?

level: middleimportance: should knowfreq 52%

basics

~20 s

Partial variables pre-bind some of a template's inputs, returning a new template that expects only the remaining ones. Values may be plain strings or zero-argument callables, which are evaluated each time the prompt is formatted.

open as a page

What does invoking a LangChain ChatPromptTemplate return, and why?

level: middleimportance: should knowfreq 38%

basics

~20 s

It returns a PromptValue — a ChatPromptValue wrapping the rendered messages — not a string and not a raw list. The wrapper exposes to_messages() and to_string(), so the same rendered prompt can feed either a chat model or a text model.

open as a page

What stops a LangChain create_agent loop from running forever?

level: seniorimportance: should knowfreq 55%

basics

~20 s

An agent from create_agent normally stops when the model returns a message with no tool calls. The hard backstop is the graph recursion limit, default 25 super-steps, which raises GraphRecursionError rather than returning a partial answer.

open as a page

In LangChain, what happens when a tool raises or the model sends invalid args?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Two different failures. Arguments that do not match the schema never reach your code: they land in AIMessage.invalid_tool_calls with an error. A tool that runs and raises propagates unless you handle it and return a ToolMessage with status set to error so the model can retry.

open as a page

In LangChain, how do bind, with_config and configurable_fields differ on a Runnable?

level: seniorimportance: should knowfreq 46%

basics

~10 s

bind fixes call-time arguments passed into the runnable itself. with_config attaches framework-level config such as tags, callbacks and run_name. configurable_fields leaves a parameter open so the caller can override it per invocation through config['configurable'].

open as a page

Why does an LCEL chain stop streaming token-by-token, and how do you fix it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

A step that only implements invoke must consume the whole upstream output before producing anything, so the chain emits one big chunk. Fix it by making that step a generator-based RunnableLambda, or move it out of the streaming path.

open as a page

How do you configure a LangChain chat model for timeouts, retries and rate limits?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Set timeout and max_retries on the model constructor, and pass rate_limiter=InMemoryRateLimiter(...) to throttle outbound requests. Keep API keys in environment variables or a SecretStr rather than literals, and remember the limiter is per-process and counts requests, not tokens.

open as a page

How do you get accurate per-call token usage from a LangChain chat model?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Read usage_metadata on the returned AIMessage — it carries input_tokens, output_tokens and total_tokens reported by the provider. When streaming, request usage explicitly (ChatOpenAI's stream_usage=True) and sum the chunks, because usage arrives only on the final chunk.

open as a page

Why do curly braces in user text break a LangChain f-string prompt template?

level: seniorimportance: should knowfreq 47%

basics

~20 s

By default a LangChain template uses f-string formatting, so every {...} in the template body is parsed as a variable slot. Literal braces — JSON examples, code, regexes — must be doubled as {{ and }}, or formatting fails with a missing-key error.

open as a page

How do LangChain few-shot prompt templates select which examples to include?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A few-shot template either holds a fixed examples list, rendered in full every call, or an example_selector that picks a subset per request — for instance SemanticSimilarityExampleSelector, which embeds the incoming input and returns the k nearest stored examples.

open as a page

When would you wrap a LangChain retriever in ContextualCompressionRetriever?

level: seniorimportance: should knowfreq 46%

basics

~20 s

When you need to retrieve wide but prompt narrow. ContextualCompressionRetriever runs a base retriever, then passes its documents plus the query through a compressor that filters, reranks or trims them — so you can raise k for recall without paying for all of it in prompt tokens.

open as a page

showing 1–30 of 36