How does OpenAI bill a gpt-image generation, and what drives the cost?
answer
- not a flat price per picture any more
- three separate token streams
- two request parameters move the number most
- the response tells you what you spent
- cost per kept asset, not per call
basics
~20 sgpt-image calls are billed in tokens, not at a flat per-image price. Text prompt tokens, any input image tokens, and image output tokens are counted separately; output tokens scale with the requested size and quality tier.
solid answer
~50 sUnlike the retired DALL-E endpoints, which charged a fixed price per image per size, gpt-image models are token-billed like the text models. Three streams count: text input tokens for the prompt, image input tokens when you upload source images to `/v1/images/edits`, and image output tokens for what is rendered. Output tokens are a function of `size` and `quality` — a high-quality portrait costs many times a low-quality square — so those two parameters are your primary cost lever, not an aesthetic afterthought. The response carries a `usage` object reporting input and output token counts, which is what you feed into per-tenant cost telemetry rather than counting requests. Cost control in practice means generating exploration and preview renders at `low`, reserving `high` for the asset a user commits to, keeping `n` honest, and caching or persisting results so a re-render is never accidental.
go deeper
Know that gpt-image is billed in tokens rather than a flat per-image price, and that larger sizes and higher quality cost more. Mentioning that the response reports usage is a good sign.
Break the bill into text input, image input and image output tokens, and explain that image output tokens are driven by size and quality. Show that you would read the usage object rather than assume a price.
Demonstrate operating experience: set quality per call site, persist generated bytes so re-renders never happen accidentally, watch edit loops and abandoned generations, and track cost per kept asset alongside rate-limit-driven queuing.
Own the economics — per-tenant budgets and chargeback, the unit-cost model behind pricing the feature, guardrails that fail closed on spend, and the strategic call about which surfaces justify premium renders at all.
## Flat price to token metering The older DALL-E endpoints had a price list: pick a model and a size, pay a fixed amount per image. Capacity planning was arithmetic on request counts. gpt-image changed the model to token metering, which aligns image generation with the rest of the platform's billing but breaks every dashboard that assumed one call equals one known price. Three token streams appear on an image call: 1. **Text input tokens** — your prompt. Usually small, but not free, and prompts assembled from long templates or user-supplied descriptions add up at volume. 2. **Image input tokens** — charged when you send images in, which happens on `/v1/images/edits`. Every reference image you attach is counted, so a multi-reference edit is materially more expensive than a plain generation. 3. **Image output tokens** — the dominant term. This is what the render costs, and it is determined by the requested `size` and `quality`. ## Size and quality are the cost dial The accepted `quality` values for gpt-image are `low`, `medium`, `high` and `auto`, and `size` covers a square, a portrait, a landscape preset and `auto`. Output token counts climb steeply as you move up the quality tiers and as the pixel area grows. The engineering consequence is that quality is a *budget decision made per call site*, not a global default: - A grid of exploration thumbnails a user will scroll past belongs at `low`. - An in-progress preview inside an editing loop belongs at `low` or `medium`. - The final asset a user exports, prints or publishes belongs at `high`. Teams that migrated by mapping the old `hd` tier to `high` everywhere often discovered the bill after the fact, because the expensive tier was now applied to drafts nobody kept. ## Read usage, do not count requests The response includes a `usage` object with input and output token counts for the call (with input broken down between text and image contributions). This is the number to persist. Concretely, a workable telemetry design records, per generation: tenant or user id, the prompt hash, `model`, `size`, `quality`, `n`, and the `usage` numbers, keyed to the stored asset. That gives you three things request counting cannot: accurate per-tenant chargeback, the ability to attribute a cost spike to a specific call site, and evidence for whether a quality downgrade actually saved money. ## Where the money leaks - **Silent re-renders.** Generation is non-deterministic, so there is no cache-by-prompt on the provider side. If your UI regenerates on every page load, navigation or retry, you pay every time. Persist the bytes and key the asset by the request parameters yourself. - **`n` greater than one.** Each image in the batch is billed; `n=4` is four renders, not a bulk discount. It is a good product choice for a chooser UI and a bad default. - **Abandoned work.** Users generate far more than they keep. Cost per *kept* asset, not cost per call, is the metric that reflects the product. - **Edit loops.** Each turn of an iterative edit re-uploads image inputs and produces a new output, so a ten-step refinement costs roughly ten generations plus the input images each time. - **Streaming partials.** Requesting intermediate preview frames adds billed output; it buys perceived latency, and you should decide whether that trade is worth it per surface. ## Rate limits are a separate constraint Cost and throughput are different ceilings. Image endpoints have their own request-rate limits, and an image render takes seconds, not milliseconds, so a burst of user-triggered generations queues. The architectural answer is usually a job queue with per-user concurrency caps: it protects the rate limit, gives you a natural place to enforce a spend budget, and lets the UI show progress rather than block a request thread. ## Guardrails worth building For any multi-tenant product, a per-tenant token budget checked before the call, with a hard stop and a soft warning, is cheap to build and prevents the runaway case. Pair it with a default quality set by surface rather than by config, and an alert on tokens-per-kept-asset rather than tokens-per-day — the ratio moves first when a UI change starts burning renders. ## What interviewers are checking The weak answer is "you pay per image". The strong answer names the token model, identifies size and quality as the levers, points at the `usage` object as the source of truth, and then talks about the non-obvious leaks: re-renders, abandoned generations and edit loops. That progression shows someone who has operated an image feature rather than merely called the endpoint.
- Why does caching by prompt hash matter more for images than for chat completions?Because a single image render is expensive and non-deterministic, so a repeat call neither reproduces the previous result nor costs less. Text responses are often cheap enough to regenerate, and providers offer prompt caching for long shared prefixes. For images, the durable artefact is the bytes: persist them keyed by model, prompt and parameters, and serve the stored asset on repeat views instead of re-rendering.
- How do you keep an interactive edit loop from becoming the most expensive surface in the product?Run the iteration at low or medium quality and render high quality only once, on the user's explicit commit. Cap the number of refinement turns, and avoid re-uploading unnecessary reference images each round, since every input image is billed. Track cost per finished asset for the loop specifically — it is usually the first place a per-call view of spend hides a problem.
- What should per-tenant image cost telemetry actually record?Per call: tenant, call site, model, size, quality, n, the usage token counts, and the id of the stored asset. That set lets you compute cost per tenant, per feature and per kept asset, and to attribute a spike to a specific surface. Counting requests alone is useless once quality and size vary, since two calls can differ in cost by an order of magnitude.
- Does requesting n=4 cost less per image than four separate calls?No. Each generated image is billed as its own set of image output tokens, so n=4 costs about four times n=1 at the same size and quality. The parameter saves round trips and gives a chooser UI a coherent batch, but it is not a bulk-pricing mechanism, and defaulting to it multiplies spend on every generation a user never keeps.
saying these in an interview costs you the question
- Says images are billed at a flat price per image
- Assumes quality only affects looks, not cost
- Counts API requests as the cost metric
- Thinks n=4 is cheaper per image than four calls
- Believes identical prompts are cached and re-served for free