In OpenAI's /v1/images/edits endpoint, what must the mask file look like?
answer
- it is a multipart upload, not JSON
- the file needs an alpha channel
- colour in the mask is ignored
- erased pixels are the editable region
- same dimensions as the source image
basics
~20 sThe mask is an optional PNG with an alpha channel, uploaded alongside the image and matching its dimensions. Fully transparent pixels mark the region the model may repaint; opaque pixels are preserved. Omit it and the whole image is re-rendered.
solid answer
~50 s`POST /v1/images/edits` is a multipart form upload, not a JSON body: you send the source `image` file, a `prompt`, and optionally a `mask`. The mask must be a PNG carrying an alpha channel, with the same width and height as the image it applies to. Transparency is the signal — pixels with alpha zero are the area the model is allowed to replace, and opaque pixels tell it to leave that region alone. Colour in the mask is irrelevant; only alpha is read, which is why a black-and-white JPEG mask does not work. Two practical points trip people up: the prompt should describe the **whole desired result**, not just the patch, because the model reasons about the complete image; and the mask is optional — with no mask at all, the model re-renders the entire image guided by the prompt and the input, which is often what you actually want.
code
python · 16 linesimport base64
from openai import OpenAI
client = OpenAI()
with open("porch.png", "rb") as image, open("mask.png", "rb") as mask:
result = client.images.edit(
model="gpt-image-1",
image=image,
mask=mask,
prompt="A golden retriever wearing a red knitted beanie, sitting on the same wooden porch",
size="1024x1024",
)
with open("edited.png", "wb") as f:
f.write(base64.b64decode(result.data[0].b64_json))go deeper
Remember the two hard requirements: the mask is a PNG with transparency, and the transparent pixels are the part the model may change. Say that it is uploaded as a file alongside the image, not embedded in JSON.
Explain that only the alpha channel is read, that dimensions must match the source, and that the prompt should describe the finished image rather than the patch. Note that the mask is optional and what happens without one.
Show judgement about when to mask at all: spatial containment for composed assets where surrounding pixels must survive, unmasked regeneration for global changes. Mention that input images are billed and moderated, which matters for user uploads.
Own the pipeline questions — validating and screening user-supplied uploads before spending a call, deciding where masks are authored, keeping generated assets and their source prompts auditable, and budgeting a workflow where each edit round trip costs input plus output tokens.
## The request shape Editing is a different HTTP shape from generation. `POST /v1/images/generations` takes JSON; `POST /v1/images/edits` takes `multipart/form-data`, because it carries binary file parts. The fields that matter are: - `model` — a gpt-image model - `image` — the source image file (gpt-image models accept multiple input images, used as references) - `prompt` — what the result should be - `mask` — optional, the editable-region file - the usual `size`, `quality`, `background`, `output_format` knobs The response is the same envelope as generation: a `data` array whose entries carry `b64_json`, plus `usage`. ## Alpha is the signal, not colour The single most common bug is authoring the mask as a black-and-white bitmap, the way many image-editing tools express masks. The API does not read luminance. It reads the alpha channel: pixels that are **fully transparent** delimit the region the model is free to repaint, and pixels that are opaque mark the region to keep. That is why the file must be a PNG (or another alpha-capable format) — a JPEG has no alpha channel at all and cannot express a mask. In practice you produce a mask by copying the source image, selecting the region to change, and erasing it to transparency — or by generating a same-size RGBA canvas where the edit region is `(0,0,0,0)` and everything else is opaque. The dimensions must match the image; a mismatched mask is rejected rather than being scaled to fit. ## Describe the whole picture in the prompt A second recurring mistake is writing the prompt as an instruction about the hole: "put a hat here". The model is generating a coherent image, so the prompt should describe the desired end state of the picture — for example, "a golden retriever wearing a red knitted beanie, sitting on the same wooden porch". The mask constrains *where* pixels may change; the prompt describes *what the image is*. Prompts written as patch instructions tend to produce edits that clash with the surrounding lighting, perspective and style, because the model was never told what the surroundings are supposed to be. ## Masked vs. unmasked editing The mask is optional, and the two modes serve different needs. - **With a mask**, you get spatial containment: everything outside the transparent region is intended to survive. That is what you want when a specific object must change and the rest of a composed layout — a logo, a product shot, text you cannot afford to have re-drawn — must not. - **Without a mask**, the model takes the input image (or images) as reference and regenerates the whole frame from the prompt. This produces more natural global changes — a season, a lighting mood, a stylistic transformation — because the model is free to adjust the entire composition coherently. A useful mental rule: mask when the constraint is spatial, skip the mask when the change is global. Reaching for a mask reflexively often gives worse results than simply describing the target image, because a masked edit forces the model to blend new content against pixels it did not generate. ## Multiple input images gpt-image models accept more than one input image on the edits endpoint. The extra images act as references — a product to place into a scene, a style to echo, characters to combine. Note that a mask, when supplied, applies to the first image; the references are inputs, not canvases. Every input image is also billed, as image input tokens, and every input image passes moderation, so a user-supplied upload can cause the call to be refused even when the prompt is innocuous. ## Operational notes Input images must be within the endpoint's accepted file types and size limits, so validate uploads before you spend a round trip. Because the request is multipart, your HTTP client, proxy and any API gateway in front of the service must all tolerate the body size — a base64-in-JSON habit from other endpoints does not carry over here. And because edits, like generation, are non-deterministic, an edit you like must be persisted; re-running the same request does not reproduce it. ## What interviewers are checking They want to know whether you have actually built a masked edit, because the alpha-versus-black-and-white detail and the whole-image prompting rule are both things you learn by getting them wrong once. A strong answer names the alpha requirement, the dimension match, and the fact that the mask is optional — and then explains when you would deliberately skip it.
- Why does a black-and-white JPEG mask fail on this endpoint?Because the API reads the alpha channel, not luminance, and JPEG has no alpha channel to read. Black pixels carry no meaning to the endpoint. The mask has to be an alpha-capable format such as PNG in which the editable region is fully transparent and everything to preserve is opaque, at exactly the source image's dimensions.
- How should the prompt differ between a masked edit and a full regeneration?Barely at all, and that surprises people. In both cases the prompt should describe the complete desired image rather than the patch, because the model composes a whole frame. The mask adds a spatial constraint on which pixels may change; it does not change the register of the prompt. Patch-style prompts like 'add a hat here' produce edits that clash with the surrounding lighting and perspective.
- What happens to moderation when the input image comes from an end user?The uploaded image is screened as well as the prompt, so a perfectly benign prompt can still be refused because of what the user uploaded — photographs of real people are a common trigger. Design the flow so a refusal on an edit is surfaced as a clear, non-retryable message about the upload, and never silently retried, since the same bytes will be refused again.
saying these in an interview costs you the question
- Supplies a black-and-white mask and expects black to mark the region
- Uploads a JPEG mask, which has no alpha channel
- Writes the prompt as a patch instruction instead of describing the whole image
- Assumes the mask is required for any edit
- Sends a mask whose dimensions differ from the source image