skip to content

When would you stream an OpenAI Assistants run instead of polling it?

level: seniorimportance: should knowfreq 38%

answer

  1. who is waiting for the text?
  2. one is resumable, the other is not
  3. a paused job still needs your answer
  4. no offset to rewind to
  5. one conversation, one execution at a time

basics

~20 s

Stream when a human is waiting on the text, because polling only reveals the answer after the run finishes. Streaming does not remove the tool-output round trip, the one-active-run-per-thread rule, or the need to re-read the run after a dropped connection.

solid answer

~50 s

Polling means creating the run and re-fetching it until it reaches a terminal state — the SDK's `create_and_poll` helper wraps exactly that loop. It is simple, restart-safe, and fine for background jobs, but the user sees nothing until the run completes, and a tight loop burns request quota, so back off between checks. Streaming instead opens an SSE connection that emits typed events — `thread.run.created`, `thread.message.delta`, `thread.run.requires_action`, `thread.run.completed` — so you can render tokens as they arrive. The catch is that streaming changes delivery, not semantics: a run that wants one of your own functions still stops at `requires_action` and needs outputs submitted before it continues, and there is no resume offset, so a dropped connection means retrieving the run by id and reading the thread's messages rather than restarting the run and paying twice.

code

python · 12 lines
python
with client.beta.threads.runs.stream(
    thread_id=thread.id,
    assistant_id=assistant.id,
) as stream:
    for event in stream:
        if event.event == "thread.message.delta":
            for block in event.data.delta.content or []:
                print(block.text.value, end="", flush=True)
        elif event.event == "thread.run.requires_action":
            pending_run_id = event.data.id  # submit tool outputs to resume
        elif event.event == "thread.run.failed":
            print(event.data.last_error)

go deeper

for a junior

Know both options exist: you can wait for the run to finish and read the answer, or subscribe to events and print text as it arrives.

for a middle

Explain the event types a streamed run emits and that polling helpers simply re-fetch the run until its status is terminal. Say which you would pick for a chat UI versus a batch job.

for a senior

Demonstrate failure handling: persist the run id before streaming, recover a dropped connection by retrieving the run, back off between polls, and respect the one-active-run-per-thread rule.

for a principal

Frame it as a delivery decision layered over an asynchronous job, and set the standard that every interactive stream has a poll-based recovery path so a network blip never costs a duplicate run.

## Two ways to watch the same state machine A run is an asynchronous job. The API gives you two ways to observe it, and they answer different questions. **Polling** — create the run, then re-fetch it on an interval until `status` is terminal. The SDK ships helpers that do this for you (`create_and_poll`, and `submit_tool_outputs_and_poll` after the tool round trip), so most polling code is one call. When the run reaches `completed` you list the thread's messages and take the reply. **Streaming** — create the run with streaming enabled and consume Server-Sent Events. The event stream is typed: run lifecycle events (`thread.run.created`, `thread.run.in_progress`, `thread.run.requires_action`, `thread.run.completed`, `thread.run.failed`), message events (`thread.message.created`, `thread.message.delta`, `thread.message.completed`) and run-step events (`thread.run.step.created`, `thread.run.step.delta`). The delta events carry incremental text you append to what you have rendered so far. ## Choosing Stream when perceived latency matters — a chat UI where a five-second wait feels broken but streaming tokens feel instant. Poll when nothing is watching: a scheduled enrichment job, a webhook handler, a queue consumer. Polling is also the more robust choice when your process may be killed and resumed by a different worker, because the run id is all the state you need; a half-consumed SSE connection is not resumable. An honest interview answer names a hybrid: stream for the interactive path, but persist the run id first, so if the stream dies you can fall back to retrieving the run. ## What streaming does not fix **The tool round trip.** If the assistant calls one of your own functions, the stream surfaces `thread.run.requires_action` and then goes quiet. You must submit outputs — optionally streaming that submission too, which yields a fresh event stream for the continuation. Built-in tools such as file search and the code interpreter execute server-side and never require this. **Thread concurrency.** A thread may have only one active run, and while a run is active you cannot append messages to that thread — the API rejects the write. If a user sends a second message before the first answer lands, you queue it yourself, cancel the run, or use a different thread. Streaming does not relax this. **Connection loss.** There is no resume-from-offset. If the SSE connection drops mid-answer, the run keeps executing server-side. The correct recovery is to retrieve the run by id and, once it is terminal, list the thread messages to get the final text. Creating a new run instead duplicates the work and the bill, and appends a second assistant message to the thread. **Run expiry.** The run's overall window still applies, most sharply when it is parked at `requires_action`. Streaming gives you the pause faster; it does not widen the window. ## Polling hygiene Each retrieve is an ordinary API request and counts against your request-per-minute allowance. A tight `while True` loop with no delay can rate-limit your own application while producing no faster an answer, since the run finishes when it finishes. Use a short initial interval with exponential backoff and a hard ceiling tied to the run's expiry, and treat a run stuck in `queued` far longer than expected as a capacity signal worth logging rather than something to hammer. ## Debugging with run steps Whichever mode you choose, the run object alone rarely explains a bad answer. Run steps show the sequence — which tool was called, with what arguments, what the model did with the result. When retrieval returns nothing useful or a function is invoked with malformed arguments, that is where you see it. Streaming exposes the same information live through step events, which is why it doubles as a good development-time trace even for workloads that will poll in production. ## Migration note The Assistants surface is beta and being retired in favour of the Responses API, whose streaming is a flat event stream over a single request rather than a job you watch. The reasoning transfers, though: interactive paths stream, background paths do not, and neither mode excuses you from persisting an identifier you can recover from.

  • Your SSE stream drops halfway through an answer. What is the correct recovery?
    Retrieve the run by the id you persisted before opening the stream. The run continues executing server-side regardless of your connection, so once it reaches a terminal status you list the thread's messages and take the completed reply. There is no resume offset to reconnect at. Creating a fresh run instead re-executes the work, bills you twice, and leaves two assistant messages in the thread.
  • How do you handle a second user message arriving while a run is still active on that thread?
    You cannot append it — the API rejects writes to a thread with an active run. Pick a policy: queue the message and post it once the run reaches a terminal state, or cancel the in-flight run and start a new one that includes both turns. Cancelling suits chat UIs where the user has clearly changed direction; queuing suits pipelines where every message must be answered.
  • Does streaming change how many tokens you are billed for?
    No. Streaming is a delivery mechanism; the run consumes the same prompt and completion tokens either way, and the usage figures on the terminal run object are what you are charged. What streaming changes is perceived latency and your ability to abort early — cancelling a run partway does stop further generation, which is a real saving on long answers the user has already dismissed.

saying these in an interview costs you the question

  • Thinking streaming removes the tool-output round trip
  • Reconnecting to a dropped stream expecting it to resume
  • Creating a new run after a connection drop instead of retrieving it
  • Polling in a tight loop with no backoff
  • Appending messages to a thread that has an active run

context