skip to content

Migrating OpenAI image calls from DALL-E 3 to gpt-image: which parameters change?

level: middleimportance: must knowfreq 54%

answer

  1. same endpoint, different vocabulary
  2. size and quality enums both moved
  3. two DALL-E-only parameters simply vanish
  4. one response field disappears, one appears
  5. variations has no successor model

basics

~20 s

The endpoint path stays the same, but the parameter vocabulary changes: new size values, a low/medium/high quality scale instead of standard/hd, no style and no response_format, no revised_prompt in the response, and new background, output_format and moderation options.

solid answer

~40 s

You still POST to `/v1/images/generations`, so the URL and auth are unchanged — everything else in the request body needs review. `size` moves to the gpt-image set (square, portrait, landscape and `auto`), so a hard-coded DALL-E 3 `1792x1024` becomes `1536x1024`. `quality` moves from `standard`/`hd` to `low`/`medium`/`high`/`auto`. The DALL-E 3-only `style` parameter (`vivid`/`natural`) is gone, and `response_format` is gone with it because gpt-image always returns base64. On the response side, `revised_prompt` disappears and a `usage` object appears, because gpt-image is billed in tokens rather than at a flat per-image price. New knobs arrive: `background`, `output_format`, `output_compression` and `moderation`. The variations endpoint has no gpt-image equivalent — re-express that use case as an edit call.

go deeper

for a junior

Be able to say the endpoint is the same but several values changed: new size options, a low/medium/high quality scale, and no style or response_format parameter. Naming two concrete changes is enough at this level.

for a middle

Walk the request and response field by field — size, quality, style, response_format, revised_prompt, usage — and explain that base64 is now the only output. Mention the new background and output_format knobs.

for a senior

Emphasise the silent breakages: cost shifts because billing moved to tokens and quality drives them, logging loses revised_prompt, and variations has no replacement. Describe how you would stage and verify the port.

for a principal

Treat it as a supplier-change exercise: pin an eval set of prompts and score outputs before and after, budget the token-based cost model against the old flat price, and design the abstraction so the next image-model turnover is a config change rather than a rewrite.

## What stays the same The transport is unchanged. It is still `POST /v1/images/generations` with a bearer API key, still a JSON body with `model` and `prompt`, and the SDK entry point is still `client.images.generate(...)`. That surface stability is what makes people underestimate the migration: the call compiles and the shape looks familiar, but several parameter *values* are now rejected, and one response field your code reads has vanished. ## Size DALL-E 3 accepted `1024x1024`, `1792x1024` and `1024x1792`. The gpt-image family uses `1024x1024`, `1536x1024` (landscape), `1024x1536` (portrait) and `auto`, which lets the model pick an aspect ratio from the prompt. Any code that stored `1792x1024` in a config file or a database of user presets is now sending a value the model does not accept. `auto` is genuinely useful for user-driven generation, but it means you cannot assume output dimensions — if your UI reserves a fixed box, pin the size explicitly instead. ## Quality DALL-E 3 had two tiers, `standard` and `hd`. gpt-image has `low`, `medium`, `high` and `auto`. This is more than a rename, because quality is now a direct cost lever: the number of billed image output tokens rises sharply from low to high. The migration decision is therefore a product decision — thumbnails, drafts and previews belong at `low`, and only the asset a user commits to needs `high`. A one-to-one mapping of `hd` → `high` across every call path is the most common way a migration quietly triples the image bill. ## Style and response_format: deleted, not renamed `style: "vivid" | "natural"` was specific to DALL-E 3 and has no successor; steer the look through the prompt instead. `response_format` is also gone, because gpt-image always returns base64 in `data[].b64_json`. Any branch in your code that handled `url` output — download the link, follow the expiry, retry the fetch — is dead and should be removed outright rather than left behind a flag. ## Response-side changes DALL-E 3 rewrote your prompt internally and returned the rewritten text as `revised_prompt` on each data entry; teams logged it for debugging and sometimes displayed it. gpt-image does not return that field, so any code reading it must be deleted, and any dashboard built on it loses its source. In exchange the response carries a `usage` object reporting input and output tokens, which is what you now feed into cost telemetry. ## New parameters worth adopting - `background`: `transparent`, `opaque` or `auto`. Transparent output requires an alpha-capable `output_format` (PNG or WebP), and it removes a whole background-removal step from asset pipelines producing icons, stickers or product cut-outs. - `output_format`: `png`, `jpeg` or `webp`, with `output_compression` tuning the lossy formats. DALL-E always gave you PNG; now storage cost is a parameter. - `moderation`: `auto` or `low`, adjusting how strictly the model's own content filtering applies. - `n` greater than 1 is available; DALL-E 3 only supported a single image per call, so batching patterns that worked around that limit can be simplified. ## The variations endpoint `/v1/images/variations` existed only for DALL-E 2 and produced alternates of an uploaded image with no prompt at all. There is no gpt-image variations model, so that capability moves to `/v1/images/edits`: upload the source image and describe the change you want in a prompt. The result is more controllable, but it is not a drop-in — the caller must now supply intent in words, which usually means a product change, not just a code change. ## A migration checklist 1. Grep for hard-coded `1792`, `1024x1792`, `"hd"`, `"standard"`, `"vivid"`, `"natural"`, `response_format` and `revised_prompt`. 2. Delete the URL-download path; keep only base64 decoding. 3. Re-decide quality per call site rather than mapping tiers mechanically. 4. Wire `usage` into cost tracking, since the per-image flat price no longer exists. 5. Re-test moderation behaviour — the filtering and its error surface are not identical to the old endpoints, so prompts that used to pass may not, and vice versa. 6. Replace variations calls with edit calls, and write the prompt that expresses the intent. ## What interviewers are checking The question rewards someone who has actually done the port. The tell of experience is naming the *silent* breakages — quality tiers that changed cost, `revised_prompt` disappearing from logs, and a variations capability that has no direct replacement — rather than reciting that the endpoint path is unchanged.

  • Your DALL-E 3 code logged revised_prompt for debugging. What replaces it under gpt-image?
    Nothing returns the model's internal rewriting, so that observability is gone. You compensate by logging the prompt you sent along with the generation parameters and a stable identifier for the stored asset, and by keeping the `usage` object so cost and prompt are correlated. If prompt rewriting matters to your product, do it yourself in an earlier LLM step where the rewritten text is yours to log.
  • Why is mapping quality 'hd' straight to 'high' across every call site a mistake?
    Because billing changed shape. gpt-image charges image output tokens that scale steeply with quality and size, so a blanket `high` applies premium cost to previews, thumbnails and abandoned drafts that no user ever keeps. The right migration re-decides quality per call site: `low` or `medium` for exploration and previews, `high` only for the asset a user actually commits to.
  • How do you replicate the old variations endpoint?
    Send the source image to `/v1/images/edits` with a prompt describing the variation you want — for example a different colourway, angle or background. There is no promptless variations model in the gpt-image family. The trade-off is that you must express intent in words, which gives far more control but requires a product surface where that intent exists.

saying these in an interview costs you the question

  • Thinks it is a pure model-string swap with no body changes
  • Maps hd to high everywhere without re-costing the call sites
  • Keeps reading revised_prompt from the response
  • Assumes 1792x1024 is still a valid size
  • Looks for a gpt-image variations endpoint

context