skip to content

OpenAI Image Generation

OpenAI generates images through the Images API with the gpt-image models, which accept a prompt plus optional input images for edits. Interviewers ask it as the current path after DALL-E was retired.

on this pageshow

questions

6

In OpenAI's Images API, how does a gpt-image generation return the image data?

level: juniorimportance: must knowfreq 62%

answer

  1. no link comes back from the API
  2. the image rides inside the JSON body
  3. look for a long encoded string field
  4. response_format was a DALL-E-era parameter
  5. data[].b64_json is the whole file

basics

~20 s

gpt-image models always return base64-encoded bytes in the response's data[].b64_json field. There is no hosted URL to fetch and no response_format choice, so the caller decodes the string and stores or serves the bytes itself.

solid answer

~40 s

A call to `/v1/images/generations` with a gpt-image model comes back as JSON containing a `created` timestamp, a `data` array, and a `usage` object. Each `data` entry carries `b64_json` — the whole image, base64-encoded, inline in the response body. The `response_format` parameter that let the older DALL-E endpoints ask for a temporary hosted URL is not supported for gpt-image models; base64 is the only delivery mode. Practically that means three things: your response payloads are large (a high-quality PNG can be several megabytes before encoding overhead), you must `base64` decode before writing a file, and you own persistence — if you want a URL for a browser to load, you upload the decoded bytes to your own object storage or CDN. `output_format` (png, jpeg, webp) and `output_compression` control what those bytes actually are.

code

python · 19 lines
python
import base64
from openai import OpenAI

client = OpenAI()

result = client.images.generate(
    model="gpt-image-1",
    prompt="A flat-design lighthouse on a cliff at dawn",
    size="1024x1024",
    quality="low",
    output_format="png",
    n=1,
)

image_bytes = base64.b64decode(result.data[0].b64_json)
with open("lighthouse.png", "wb") as f:
    f.write(image_bytes)

print(result.usage)

go deeper

for a junior

Know the field name and say it plainly: the image comes back as base64 in data[].b64_json, and you decode it before writing a file. Mentioning that there is no URL option shows you have actually made the call.

for a middle

Explain that base64 is the only delivery mode for gpt-image and that response_format belonged to the retired DALL-E endpoints. Connect output_format and output_compression to what the decoded bytes actually are.

for a senior

Show the operational consequences: response payloads inflate by about a third, body-size limits and logging need attention, and persistence is your responsibility. Describe the decode-then-upload-to-object-storage path as the default.

for a principal

Frame image bytes as an asset-lifecycle problem, not an API detail — where generated media is stored, how it is keyed and deduplicated, retention and cost of that store, and how the API tier stays stateless while multi-megabyte artefacts flow past it.

## The response shape An image generation request to OpenAI's Images API (`POST /v1/images/generations`) with a gpt-image model returns a JSON object roughly like: - `created` — a Unix timestamp - `data` — an array with one entry per requested image (`n` controls how many) - `usage` — token accounting for the call Each element of `data` holds `b64_json`: the complete image file, base64-encoded, embedded in the response body. There is no `url` field and no `revised_prompt` field for gpt-image models. ## Why there is no URL The older DALL-E 2 and DALL-E 3 endpoints accepted a `response_format` parameter with two values, `url` and `b64_json`. `url` returned a short-lived link to an OpenAI-hosted copy of the image, which expired after a bounded window — you had to download it promptly or lose it. That option does not exist for the gpt-image family: `response_format` is not a supported parameter for these models, and the output is always base64. Since DALL-E 2/3 were retired in May 2026, the URL path is gone from the product entirely, so any integration written against it has to be reworked to decode bytes. This is not a cosmetic difference. It changes where images live in your architecture. With URL output it was tempting (and fragile) to hand the OpenAI link straight to a browser. With base64 output, the bytes land in your process, and you make an explicit decision: write to disk, push to S3/GCS, store in a blob column, or return a `data:` URI to the client for a one-shot preview. ## Working with the bytes The mechanical steps are always the same: 1. Read `response.data[i].b64_json`. 2. Base64-decode it into raw bytes. 3. Write those bytes to a file or object store with the right extension. The extension must match `output_format`. gpt-image models accept `output_format` of `png`, `jpeg` or `webp`, defaulting to PNG. For `jpeg` and `webp` you can also pass `output_compression` (a 0–100 quality knob) to trade file size against artefacts — useful when you are storing thousands of generated assets and PNG is wastefully large. Transparency is the one coupling to remember: `background: "transparent"` only makes sense with a format that has an alpha channel, so pair it with PNG or WebP, not JPEG. ## Payload-size consequences Base64 inflates binary by roughly a third. A high-quality 1024×1536 PNG can be several megabytes raw, so the HTTP response can be large, and requesting `n` images multiplies it. Consequences worth naming in an interview: - Do not log the raw response body; you will fill your log store with megabytes of base64 per call. - Watch client, proxy and gateway body-size limits; a request that "works in curl" can be truncated by an API gateway with a small response cap. - Streaming JSON parsers or `n=1` per call keep memory predictable for large batches. - Serialising the base64 string into a queue message or a database row is usually the wrong move; store bytes in object storage and pass a key. ## The same rule applies to edits `POST /v1/images/edits` — the endpoint you use to modify an existing image with a prompt, with or without a mask — returns the identical envelope: `data[].b64_json` plus `usage`. So the storage plumbing you build for generation is reused unchanged for edits, and any code branching on "URL or base64" is dead code you should delete. ## Common mistakes The frequent bugs are all decoding-adjacent. Writing the base64 string to a `.png` file without decoding produces a file that no viewer opens. Prefixing bytes with a `data:` URI header when writing to disk corrupts the file, while *forgetting* the `data:image/png;base64,` prefix when injecting into an `<img src>` breaks the browser render — the same string needs opposite treatment in the two destinations. And assuming OpenAI keeps a copy you can re-fetch later is simply wrong: if you drop the bytes, the image is gone and regenerating costs another call, with a different result because generation is not deterministic. ## What to say in an interview State the field name, state that base64 is the only mode for gpt-image, and then show you have thought about ownership: your service persists the asset, your CDN serves it, and the API response is transient. That sequence — field, constraint, architectural consequence — is what separates a recall answer from an engineering one.

  • How would you serve those images to a browser without ballooning your API responses?
    Decode once on the server, upload the bytes to object storage (S3, GCS, Azure Blob) under a content-addressed key, and return that URL or a signed URL to the client. The browser then fetches the binary directly from a CDN instead of pulling multi-megabyte base64 through your own API layer, and you get caching, range requests and lifecycle expiry for free.
  • What does output_compression change, and when would you use it?
    It sets the quality/size trade-off for lossy outputs, so it applies when `output_format` is `jpeg` or `webp` and is meaningless for PNG. Lower values shrink the stored asset substantially at the cost of visible artefacts. It is worth tuning when you generate assets at volume and store them long-term; for a one-off hero image, keep PNG and skip it.
  • Does the edits endpoint return anything different from generations?
    No. `POST /v1/images/edits` returns the same envelope — `created`, a `data` array whose entries carry `b64_json`, and a `usage` object. The difference is on the request side: edits is a multipart upload carrying one or more input images and an optional mask. The response-handling code is shared.

saying these in an interview costs you the question

  • Says the API returns a CDN URL you can hotlink
  • Believes response_format: 'url' still works on gpt-image models
  • Assumes OpenAI stores the image so you can re-fetch it later
  • Writes the base64 string to a .png file without decoding it
  • Logs the full response body, including megabytes of base64

context

open as a page

Migrating OpenAI image calls from DALL-E 3 to gpt-image: which parameters change?

level: middleimportance: must knowfreq 54%

basics

~20 s

The endpoint path stays the same, but the parameter vocabulary changes: new size values, a low/medium/high quality scale instead of standard/hd, no style and no response_format, no revised_prompt in the response, and new background, output_format and moderation options.

open as a page

In OpenAI's /v1/images/edits endpoint, what must the mask file look like?

level: middleimportance: should knowfreq 44%

basics

~20 s

The mask is an optional PNG with an alpha channel, uploaded alongside the image and matching its dimensions. Fully transparent pixels mark the region the model may repaint; opaque pixels are preserved. Omit it and the whole image is re-rendered.

open as a page

How do you handle a moderation refusal from OpenAI's Images API in production?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A blocked image request returns HTTP 400 with a moderation_blocked error, not a rate-limit or server error. It is not retryable: surface a clear refusal, invite the user to rephrase, log the attempt for abuse review, and never loop on backoff.

open as a page

How does OpenAI bill a gpt-image generation, and what drives the cost?

level: seniorimportance: should knowfreq 42%

basics

~20 s

gpt-image calls are billed in tokens, not at a flat per-image price. Text prompt tokens, any input image tokens, and image output tokens are counted separately; output tokens scale with the requested size and quality tier.

open as a page

How can OpenAI's Images API show progress before a gpt-image render finishes?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Set stream to true on the generations or edits call and ask for partial images. The API then emits a Server-Sent Events stream of progressively refined previews before the final image, which the UI can render as the result takes shape.

open as a page