In image generation, when do you use mask inpainting instead of conversational editing?
answer
- ask what each one guarantees
- one freezes pixels, one regenerates all
- masks cost labour and show seams
- identity across a set, not one region
- explore conversationally, finalise masked
basics
~20 sUse a mask when a region of an approved photograph must be guaranteed byte-identical outside the edit — legal, brand or evidentiary constraints. Use conversational multi-reference editing when the change is global or identity must carry across several outputs, accepting that the whole frame is regenerated.
solid answer
~50 sMask inpainting re-noises and regenerates only the masked region, conditioned on the surrounding pixels and the prompt, then composites the untouched area straight back. That gives a hard pixel-level guarantee: everything outside the mask is the original file. It costs you a mask — hand-drawn or from a segmentation model — and it fails at boundaries, where seams, lighting mismatch or context bleed show up. Conversational multi-reference editing instead takes the source plus a handful of reference images and an instruction, and regenerates the whole image while preserving subject identity; current hosted models accept on the order of a dozen references and hold identity across a multi-turn thread. Nothing is guaranteed unchanged, and small drift accumulates over turns. Practically: masks for legally approved photography where only the rug may change; conversational editing when you need the same armchair rendered coherently in a new room, new lighting or new material.
go deeper
Know the basic distinction: a mask edits only the region you draw and leaves the rest of the file alone, while conversational editing rewrites the whole image from an instruction and reference pictures.
Explain the mechanics — re-noising inside the mask and compositing outside it, versus full regeneration conditioned on references — and name the characteristic failures of each, seams and context bleed versus drift.
Choose by guarantee under a real constraint: approved photography, legal captions or regulated packaging force masks, while relighting and identity across a set force conversational editing. Be ready to describe the explore-then-freeze hybrid and how you would audit which pixels changed.
Own the workflow policy: where sign-off happens, which stage is allowed to regenerate approved assets, and what evidence the pipeline must retain about provenance of each edit. That decision shapes review cost and legal exposure far more than model choice does.
## Two mechanisms, two guarantees The decision is not about which produces a nicer picture. It is about what each one *promises*, and the promises are categorically different. **Mask inpainting.** You supply the source image, a binary or feathered mask, and a prompt. The pipeline encodes the image, re-noises the masked region, denoises it with the surrounding context and the prompt as conditioning, decodes, and composites the result so that pixels outside the mask come from the original file. The guarantee is exact: the untouched area is unchanged, bit for bit, because it was never regenerated. Outpainting is the same mechanism with the mask outside the original frame. **Conversational multi-reference editing.** You supply the source, a set of reference images, and an instruction in natural language across a turn-by-turn thread. The model regenerates the entire frame while carrying subject identity and scene context forward from the references and previous turns. As of mid-2026 hosted models in this class accept on the order of a dozen reference images, support localized instructions in words rather than masks, and offer camera and lighting direction. The guarantee is *semantic* — that the chair is recognisably the same chair — not *pixel-level*. ## When the mask is mandatory Any time some part of the image is a matter of record. Approved product photography where the product geometry has been signed off. A photograph with a legal caption. A medical or insurance image where the untouched region is evidence. Regulated packaging where the label must be the real one. Compliance workflows where you must be able to say precisely which pixels changed, and diff the file to prove it. Masks are also right for small, well-localized swaps in a large frame — replace the rug, remove a light switch, change the vase — where regenerating the whole image would be a large amount of compute and a large amount of risk to buy a small change. ## When conversational editing wins When the change is inherently global. Relighting a scene, changing the season visible through a window, moving the camera, restyling a room around a fixed product: none of these decompose into a region you could draw a boundary around, and a mask would produce a hard edge between two mutually inconsistent lighting models. Also when *identity across a set* is the actual requirement. A catalogue that needs one armchair in six room styles is not six independent generations and not six masked edits — it is one subject carried through a series, which is exactly what reference-image conditioning and multi-turn threads are built for. Asking for 'the same chair, same room, now in oak' and getting a coherent variant is a different capability from region replacement. And when there is no one to author masks. Mask production is real labour: either a person drawing them, or a segmentation model plus a review step for its mistakes. For high-volume or self-serve workflows that pipeline may cost more than the generation. ## The failure modes to name *Mask inpainting.* Boundary seams where the regenerated region does not match the original in lighting, grain, colour temperature or focus. Context bleed, where the model continues the surroundings into the mask instead of following the prompt. Scale confusion when the mask is small — the model has little context and produces an object at the wrong size. Semantic incoherence, because nothing outside the mask can adapt: put a lamp in and no shadow appears anywhere else in the room, because those pixels are frozen by construction. Feathering the mask edge softens seams but blurs the guarantee. *Conversational editing.* Drift: each turn is a fresh generation, so details you never mentioned wander — a cushion changes weave, a wall shifts a shade, a logo deforms slightly. Over a long thread this compounds until the output no longer matches the approved original. Identity preservation is good but not exact, which is fatal for a product whose geometry is the point. Resolution and framing can shift between turns. And the fact that nothing is pinned means you cannot prove which parts are original. ## The hybrid, which is usually the right production answer Explore conversationally, finalise with masks. Use multi-turn reference editing to find the composition and lighting you want, then, once an image is approved, restrict all further changes to masked regions so the approved content is frozen. For a catalogue: generate the six room styles conversationally around a fixed product reference set, get sign-off, and thereafter make seasonal swaps — different rug, different throw — as masked edits on the approved frames. This keeps the exploratory flexibility where it is cheap and the guarantee where it is required. A second hybrid worth knowing: generate conversationally, then composite the approved product photograph back in over the generated scene, with a masked pass only to reconcile shadow and contact edges. That gives literal product fidelity with a generated environment. ## How to answer the question Lead with the guarantee, not the tooling. Mask inpainting is the only one of the two that promises untouched pixels, so it is the answer whenever some region must be provably unchanged; conversational multi-reference editing is the answer when the change is global or identity must carry across a set, and its cost is drift and the absence of any pixel-level promise.
- What causes visible seams at a mask boundary, and how do you reduce them?The regenerated region is sampled independently of the frozen surroundings, so grain, colour temperature, focus and lighting direction can all differ slightly, and the composite makes the discontinuity obvious. Mitigations: feather the mask, dilate it so the model sees more context to blend into, run a low-strength pass over the seam band, and match noise and colour after compositing. Each softens the exactness of the untouched-pixels guarantee a little.
- Why does drift accumulate over a long conversational editing thread?Every turn regenerates the entire frame, and nothing is pinned, so any detail the instruction does not mention is re-sampled and may land slightly differently. Small deviations feed the next turn as input. After a dozen turns the cumulative change can be substantial. The practical discipline is to branch from an approved frame rather than continuing a long chain, and to re-anchor with the original reference images.
- How would you handle a catalogue where the product geometry must be exact but the room may be generated?Do not ask the model to draw the product. Generate the environment conversationally around reference images so the lighting and perspective are right, then composite the approved product photograph into the scene, and use a narrow masked pass only on contact shadows and edges where the two must meet. Exact product pixels, generated surroundings, and the reconciliation confined to a region you can inspect.
A mask is a stencil taped over a finished print — only the exposed area is repainted. Conversational editing is asking the studio to shoot the scene again with the same props: recognisably the same, never the same negative.
saying these in an interview costs you the question
- Assumes conversational editing leaves unmentioned areas untouched
- Thinks masks are only about convenience, not guarantees
- Ignores that an inpainted lamp casts no shadow outside the mask
- Treats reference-image identity preservation as pixel-exact
- Runs a long edit thread without re-anchoring to the approved frame