skip to content

In LangChain, what do a Runnable's invoke, batch and stream methods each do?

level: juniorimportance: must knowfreq 78%

answer

  1. One interface, six calling conventions
  2. Sync trio and async trio
  3. Batch is concurrent, not a loop
  4. Streaming yields chunks, invoke yields one result
  5. Composition inherits all of them free

basics

~20 s

invoke runs the runnable once on one input and returns the final output. batch runs many inputs, concurrently by default. stream yields output in chunks as it is produced. Each has an async twin: ainvoke, abatch, astream.

solid answer

~50 s

Every LangChain component that participates in LCEL implements the `Runnable` interface, and that interface is the reason a chain has a uniform calling convention. `invoke(input)` runs it once and returns the complete output. `batch(inputs)` takes a list and runs the entries concurrently — the default implementation uses a thread pool, capped by `max_concurrency` in the config — returning results in input order. `stream(input)` returns an iterator of output chunks so you can render tokens as they arrive rather than waiting for the whole response. Each has an async counterpart: `ainvoke`, `abatch`, `astream`, plus `batch_as_completed`/`abatch_as_completed` when you want results as they finish instead of in order. The important part is that these are inherited by composition: pipe several runnables together and the resulting chain exposes all of them without you writing any batching or streaming code.

go deeper

for a junior

Know the six names and what each returns: invoke gives one result, batch gives a list, stream gives chunks, and a-prefixed versions are async. Say that any composed chain exposes all of them.

for a middle

Explain that batch is concurrent by default with max_concurrency as the throttle, and that the default stream falls back to one whole-output chunk when a component has nothing incremental to give.

for a senior

Talk about running async end to end behind a server, capping concurrency to respect provider rate limits, and using return_exceptions so one bad element does not lose a bulk job.

for a principal

Frame the interface as the actual product of LCEL: a uniform calling convention plus config and callback propagation is what makes components swappable and traceable, and it is the yardstick when judging whether a framework's abstraction earns its cost.

## What a Runnable is `Runnable` is the single protocol that LangChain Expression Language (LCEL) is built on. Prompt objects, chat models, output parsers, retrievers, and arbitrary functions all implement it, which is what makes them interchangeable pieces in a pipeline. Implementing the protocol means supporting one required method — `invoke` — and inheriting sensible defaults for everything else. ## The six core methods **`invoke(input, config=None)`** — the fundamental operation. One input in, one complete output out, synchronously. Every other method is defined in terms of it unless a component overrides them with something better. **`ainvoke`** — the async version. For components backed by network I/O (models, retrievers, HTTP tools) this is a genuinely non-blocking implementation, not a thread wrapper, which matters when you are serving many concurrent requests from an async web server. **`batch(inputs, config=None, *, return_exceptions=False)`** — takes a *list* of inputs and returns a list of outputs in the same order. The default implementation does not loop serially; it fans the inputs out across a thread pool. You control the width with `max_concurrency` in the config, which is how you stay under a provider's rate limit. `return_exceptions=True` makes a failing element return the exception object in its slot instead of blowing up the whole batch — useful for bulk offline jobs where one bad row should not lose the other 999. **`abatch`** — the async equivalent, built on `asyncio.gather` rather than threads. **`stream(input, config=None)`** — returns an iterator of chunks. For a chat model these are token-sized message chunks; for a chain they are whatever the last streaming-capable step emits. The default implementation for a component that has no incremental behaviour simply yields one chunk containing the entire output, so `stream` always *works* even when it does not actually stream. **`astream`** — async iterator version, the one you normally use behind a web endpoint that pushes server-sent events. There are two more worth knowing: `batch_as_completed` and `abatch_as_completed` yield `(index, output)` pairs as each finishes, so a slow element does not hold up faster ones. ## Why the interface exists The payoff is compositional. When you write `a | b | c`, the resulting object is itself a `Runnable`, so it also has `invoke`, `batch`, `stream` and the async trio. You did not implement concurrency, you did not implement chunk propagation, and you did not implement async — the sequence delegates to its steps. This is the honest answer to "why write chains this way at all": the pipe syntax is cosmetic, but the uniform interface it produces is not. The same uniformity carries the `RunnableConfig` — callbacks, tags, metadata, `run_name`, `max_concurrency`, `recursion_limit` — down through every nested step, which is what makes end-to-end tracing work without threading a context object through your own code by hand. ## Practical notes - **Don't hand-roll a loop over `invoke` when you mean `batch`.** A `for` loop is serial; `batch` is concurrent and rate-limit aware. - **Async is contagious.** If your call site is async, use `ainvoke`/`astream` end to end. Calling the sync method from inside an event loop blocks it. - **`stream` succeeding is not proof of streaming.** A chain returns an iterator whether or not anything inside it produces incremental output; if you get exactly one chunk, some step in the middle is buffering. - **Batch is not a provider batch API.** It is client-side concurrency over ordinary requests; it does not get you a discounted asynchronous batch endpoint. ## What it does not give you The interface gives you calling conventions, not durability. There is no built-in retry (that is `.with_retry()`), no checkpointing of intermediate results, and no cross-process state. A crashed `batch` starts over from scratch.

  • How do you stop a batch of 500 inputs from tripping the provider's rate limit?
    Pass `max_concurrency` in the `RunnableConfig` — `chain.batch(inputs, config={"max_concurrency": 5})` — which caps how many elements run at once in the default thread-pool implementation. For long jobs also add `.with_retry()` so the occasional 429 is retried rather than failing the element, and consider `return_exceptions=True` so one bad input does not abort the run.
  • If a step in the middle of a chain has no async implementation, what happens when you call astream on the chain?
    It still works. The default async methods fall back to running the sync implementation in a thread pool, so correctness is preserved but that step occupies a worker thread and gives up the concurrency benefit of the event loop. Under load that fallback is a real bottleneck, so prefer components with native async implementations on hot paths.
  • What is the difference between batch and batch_as_completed?
    `batch` returns a list aligned with the input order, so it waits for every element before returning anything. `batch_as_completed` yields `(index, output)` pairs as each element finishes, in completion order. Use the latter when you want to start processing or displaying early results and the slowest element should not gate the fast ones.

saying these in an interview costs you the question

  • Thinking batch just loops over invoke serially
  • Assuming stream returning an iterator proves real token streaming
  • Believing batch calls a provider batch/discount endpoint
  • Calling sync invoke inside an async event loop
  • Thinking you must implement streaming yourself for each chain

context