skip to content

In OpenAI's Assistants API, what states does a run pass through before finishing?

level: middleimportance: should knowfreq 48%

answer

  1. queued, then working, then done
  2. one non-terminal pause in the middle
  3. the model is waiting on you
  4. the pause has a deadline
  5. answer lands in the conversation, not the job

basics

~10 s

A run is created as queued, moves to in_progress, and normally ends completed. It can pause at requires_action while it waits for tool outputs, and can instead end as failed, cancelled, incomplete, or expired.

solid answer

~50 s

The Assistants API splits state across four objects: an **assistant** (model, instructions, tools), a **thread** (the stored message history), **messages**, and a **run** — one execution of an assistant against a thread, itemised into run steps. Creating a run returns it as `queued`; it becomes `in_progress` once the model starts, and terminates as `completed`, `failed`, `cancelled`, `incomplete`, or `expired`. The interesting non-terminal state is `requires_action`: the model asked for a tool, and the run parks until you POST the results to the submit-tool-outputs endpoint, after which it resumes. That pause is time-boxed — a run has an expiry, and letting it lapse moves it to `expired` and rejects the late outputs, so the whole loop has to be driven by something that survives your process restarting. Note that this API is in beta and OpenAI has published its retirement in favour of the Responses API.

code

python · 16 lines
python
run = client.beta.threads.runs.create_and_poll(
    thread_id=thread.id,
    assistant_id=assistant.id,
)

if run.status == "requires_action":
    outputs = [
        {"tool_call_id": call.id, "output": "42"}
        for call in run.required_action.submit_tool_outputs.tool_calls
    ]
    run = client.beta.threads.runs.submit_tool_outputs_and_poll(
        thread_id=thread.id, run_id=run.id, tool_outputs=outputs
    )

print(run.status)
messages = client.beta.threads.messages.list(thread_id=thread.id)

go deeper

for a junior

Be able to name the four objects — assistant, thread, message, run — and say that a run is one execution whose status you check until it is completed.

for a middle

Walk the state machine out loud: queued, in_progress, the requires_action pause for your own functions, and the terminal states completed, failed, cancelled, incomplete, expired.

for a senior

Show that you have operated it: runs expire while paused, a thread allows only one active run, and the assistant's reply must be read from the thread. Design the tool loop so a process restart can resume it.

for a principal

Argue about placement — a run is a short-lived execution, not a durable workflow engine, and human-in-the-loop steps do not fit inside its expiry window. Factor that into a migration plan off the retiring Assistants surface.

## The object model Before the lifecycle makes sense, the four objects have to be clear. - **Assistant** — a stored configuration: which model, what instructions, which tools, and which resources (a vector store for file search, files for the code interpreter). It holds no conversation. - **Thread** — a server-side conversation. You append messages to it; it grows indefinitely and the platform decides what fits into the model's context window at execution time. - **Message** — one user or assistant turn inside a thread, made of content parts plus optional attachments. - **Run** — one execution of a given assistant against a given thread. It is the only object with a lifecycle, and it is what you poll or stream. - **Run step** — the itemised trace of what the run did: a message creation, a tool call. Useful for debugging why an answer took the shape it did. All of this sits behind the beta surface (the SDK reaches it as `client.beta.threads...`, sending the assistants beta header for you). ## The state machine Creating a run returns it immediately in `queued`. From there: - **queued → in_progress** — the model started producing tokens or calling tools. - **in_progress → requires_action** — the model emitted tool calls the platform cannot execute itself. `required_action.type` is `submit_tool_outputs`, and `required_action.submit_tool_outputs.tool_calls` lists what to run. The run makes no further progress until you post outputs back for **every** call in that list, keyed by `tool_call_id`. - **requires_action → in_progress** — after your submission is accepted. - **→ completed** — the assistant's reply has been appended to the thread as a new message. Note that the answer is not in the run object; you list the thread's messages to read it. - **→ failed** — an error occurred; `last_error` carries a code and message (rate limit, invalid request, server error). - **→ cancelled / cancelling** — you called cancel; there is a brief `cancelling` state while the platform stops the work. - **→ incomplete** — the run hit a limit you set, such as maximum prompt or completion tokens, and stopped early with partial output. - **→ expired** — the run exceeded its window, most commonly because it sat in `requires_action` waiting for tool outputs that never came. Built-in tools do **not** produce `requires_action`. File search and the code interpreter run inside the platform, so the run stays `in_progress` while they execute; only your own functions require a round trip. ## Why expiry is the trap A run's pause at `requires_action` is bounded — roughly ten minutes in practice. That is fine for a fast HTTP lookup and fatal for anything human-in-the-loop or slow: a tool that needs an approval click, a nightly batch, a job queue with backpressure. When the window lapses the run goes `expired`, submitting outputs afterwards is rejected, and the work must be redone from a fresh run. The design lesson is that the run is a **short-lived execution**, not a durable workflow. Long-running work belongs in your own orchestrator, with the model call as a step inside it, not the other way round. ## Concurrency A thread can have only one active run at a time, and while a run is active you cannot append messages to that thread. Two user turns arriving concurrently for the same conversation therefore need queuing on your side; there is no server-side interleaving. This is the single most common production surprise for teams that map one thread to one long-lived user session. ## Reading the result When the run completes, the assistant's message is in the thread. You list thread messages (newest first) and take the messages created by that run. Run steps let you see the intermediate tool calls, which matters when the answer is wrong and you need to know whether retrieval or reasoning failed. ## Status of the API The Assistants API never left beta. OpenAI has stated it is superseded by the Responses API — which offers the same built-in tools plus hosted state without the thread/run ceremony — and published a shutdown date for the Assistants surface. Interviews still ask about the run lifecycle because plenty of production systems were built on it and are being migrated, and because the state machine is a good proxy for whether a candidate understands the difference between a request and a job. If you are starting fresh, build on Responses; if you are asked about Assistants, be ready to describe both the lifecycle and the migration.

  • Which tools cause a run to enter requires_action, and which do not?
    Only functions you declared yourself. The platform cannot execute your code, so it parks the run and hands you the tool calls to fulfil. Built-in tools — file search and the code interpreter — execute inside OpenAI's infrastructure, so the run simply stays in_progress while they work and you see them afterwards as run steps. A run that never touches your own functions therefore goes queued → in_progress → completed with no round trip.
  • Where do you read the assistant's actual answer after a run completes?
    From the thread, not the run. A completed run appends one or more assistant messages to the thread; you list the thread's messages and take the ones created by that run id. The run object itself carries status, usage, and last_error but not the reply text. Run steps sit alongside and show the intermediate tool calls, which is where you look when the answer is wrong and you need to know whether retrieval or reasoning failed.
  • A run ended with status incomplete rather than completed. What does that mean?
    It stopped against a limit rather than an error. Typically the run hit a maximum prompt or completion token budget you configured, so the assistant produced partial output and terminated cleanly. Treat it like a truncated answer: inspect the incomplete details and usage, then either raise the budget, shorten the thread, or split the task. It is not a failure — last_error is empty — so error-handling code keyed only on failed will silently ship half an answer.

saying these in an interview costs you the question

  • Saying requires_action means the run failed
  • Expecting file search or code interpreter to require tool outputs
  • Reading the answer off the run object instead of the thread
  • Assuming a paused run waits indefinitely for tool outputs
  • Running two concurrent runs on the same thread

context