skip to content

How can OpenAI's Images API show progress before a gpt-image render finishes?

level: seniorimportance: nice to knowfreq 27%

answer

  1. the spinner problem, not the throughput problem
  2. the response becomes an event stream
  3. previews arrive before the finished frame
  4. you can ask for a handful of intermediate frames
  5. the previews are not free

basics

~20 s

Set stream to true on the generations or edits call and ask for partial images. The API then emits a Server-Sent Events stream of progressively refined previews before the final image, which the UI can render as the result takes shape.

solid answer

~50 s

Image renders take seconds, so a blocking call leaves a spinner on screen. The Images API supports streaming: pass `stream: true` and request partial images (up to three), and the response becomes a Server-Sent Events stream instead of one JSON body. Each partial event carries a base64 preview of the image at an intermediate stage of refinement; a final event carries the completed image, which is the one you persist. Two caveats decide whether it is worth it. First, partials are not free — they are additional generated output and add to what you are billed, so a low-value surface may not justify them. Second, streaming changes your plumbing: the connection is held open for the duration, so proxies, load balancers and serverless timeouts must tolerate it, and you need a path from the server-side stream to the browser. For batch or background generation, skip streaming entirely and just render at low quality when speed matters.

go deeper

for a junior

Know that the call can stream, that setting stream to true turns the response into an event stream, and that intermediate previews arrive before the finished image. Say that the final event carries the image you keep.

for a middle

Explain that partial images are progressively refined previews, that you can request a small number of them, and that the final event is authoritative. Note that partials add to the billed output.

for a senior

Reason about the trade: streaming buys perceived latency at extra token cost while lowering quality buys real speed at lower cost. Cover the infrastructure consequences — proxy buffering, timeouts, serverless limits, and relaying the stream to the browser.

for a principal

Decide per surface where streaming belongs at all, and set the architecture: foreground streaming for high-value single renders, queued background generation with object storage and notification for everything else, with concurrency and spend controls at the queue.

## Why streaming exists for images A text completion streams because tokens arrive one at a time and reading can start immediately. An image is different: there is no partial answer to read, only a picture that becomes less noisy. Streaming here delivers *preview frames* — snapshots of the image at intermediate stages of refinement — so the user watches the result emerge rather than staring at an indeterminate spinner for several seconds. It is a perceived-latency feature, not a throughput one: the final image does not arrive sooner. ## The mechanics On a generation or edit call you set `stream: true` and ask for a number of partial images (the parameter accepts up to three). The HTTP response is then a Server-Sent Events stream rather than a single JSON document. Events arrive in order: each partial event carries a base64-encoded preview of the image so far, and a final event carries the completed image along with the call's usage accounting. Your client renders each partial as it arrives, replacing the previous one, and treats the last event as authoritative — that is the image you decode and persist. A correct consumer has to handle the ordinary SSE hazards: the stream can terminate mid-flight, in which case you have previews but no finished image and must treat the call as failed; an error can be delivered as an event rather than an HTTP status, so parsing must not assume success once headers are received; and the client must not treat the last *received* partial as the result when the stream broke early, or users end up with blurry half-rendered assets saved to their library. ## The cost trade Each partial is generated output and adds to the call's billed tokens. That makes streaming a per-surface decision rather than a global default: - **Worth it**: a foreground, user-initiated generation where the person is watching and waiting. The engagement benefit is real and the surcharge applies to a call the user cares about. - **Not worth it**: batch pipelines, background jobs, server-to-server generation, and any call whose result nobody is watching in real time. There is no one to show the previews to, so you are paying for frames that go straight to the bit bucket. - **Consider the alternative**: dropping `quality` to `low` genuinely makes the render finish sooner and costs *less*, whereas streaming costs more and only changes the waiting experience. For an exploratory grid, low quality beats streaming. ## Infrastructure implications Streaming holds a connection open for the whole render. That interacts with everything between your client and the API: - **Timeouts**: idle and total-request timeouts on load balancers, reverse proxies and API gateways must exceed the render duration. Buffering proxies are worse than slow ones — a proxy that buffers the whole body defeats streaming entirely without erroring. - **Serverless**: short function execution limits and response buffering make some serverless platforms a poor fit for holding an image stream open. - **Fan-out to the browser**: the SSE stream terminates at your backend, which holds your API key. You then need a second hop — your own SSE endpoint or a WebSocket — to relay previews to the browser, plus a decision about what happens when the user navigates away mid-render. - **Payload volume**: every partial is a full base64 image. Three partials plus the final frame means roughly four images of bytes over the wire for one result. ## The alternative pattern For anything that is not a foreground render, the standard architecture is a job queue: accept the user's request, enqueue it, generate without streaming, store the bytes in object storage, and notify the client when the asset is ready. That decouples render latency from HTTP request lifetime entirely, survives client disconnects, gives you a natural place to enforce per-user concurrency against rate limits, and lets retries be a queue concern. Streaming and queueing are not competitors — one improves a foreground experience, the other makes background generation reliable — and a mature product usually has both. ## What interviewers are checking This is a differentiator question rather than a screener. What earns credit is the trade-off reasoning: knowing that streaming buys perceived latency and costs extra output, that lowering quality is the cheaper lever when actual speed is the goal, and that holding a connection open for seconds has real consequences for proxies, timeouts and serverless runtimes. Reciting the parameter name alone is a thin answer.

  • Why is lowering quality often a better latency fix than streaming?
    Because it changes the actual render rather than the waiting experience, and it costs less rather than more. A low-quality render consumes far fewer image output tokens and finishes sooner, which is exactly right for exploration grids and previews. Streaming leaves the total time unchanged and adds billed output for the partial frames; it is worth it only when the user must watch a single high-value render.
  • What breaks when you put an SSE image stream behind a buffering reverse proxy?
    Nothing errors, which is what makes it insidious: the proxy accumulates the whole response and forwards it once complete, so the client receives every partial at the moment the final image arrives. The feature silently stops working while still costing extra tokens for the partials. Verify streaming end to end through the real ingress, not just against the API directly.
  • When would you avoid streaming entirely and use a job queue instead?
    For any generation nobody is watching live: batch asset production, scheduled work, server-to-server calls, and requests that must survive a client disconnect. Enqueue the request, render without streaming, write the bytes to object storage, and notify when ready. The queue also gives you per-user concurrency control against rate limits and a clean place to implement retries and spend caps.

saying these in an interview costs you the question

  • Thinks streaming makes the image finish faster
  • Assumes partial previews are free
  • Saves the last partial received when the stream breaks early
  • Ignores proxy buffering and connection timeouts
  • Streams background and batch generations nobody is watching

context