skip to content

When sending images to a vision model, how do you choose between base64, a URL, and an uploaded file?

level: middleimportance: must knowfreq 62%

answer

  1. Same pixels either way; only plumbing differs
  2. Ask: how many times is this sent?
  3. Base64 inflates payload about a third
  4. URLs add a fetch you cannot instrument
  5. Upload once, reference many times

basics

~20 s

Inline base64 suits one-off images but inflates the request by roughly a third. A URL keeps requests small when the image is already hosted and reachable by the provider. An uploaded file reference wins when the same image is sent many times across requests.

solid answer

~50 s

All three end with the same pixels in the model's context; they differ in who stores the bytes and how often they cross the wire. **Base64** embeds the image in the request body — no external dependency, no hosting, but about 33% size overhead and a fresh upload on every call. **A URL** keeps the request tiny, at the price of the provider needing network access to your host and of a fetch that can be slow, rate-limited, or fail on a private or signed-URL endpoint. **An uploaded file reference** costs one upload and then a small identifier per request, which is the right shape when an asset is reused. For a QA pipeline reading failure screenshots: base64 for the one-off screenshot captured seconds ago, a URL for CDN-hosted reference images, and uploaded files for the fixed set of screenshots replayed across every regression run.

go deeper

for a junior

Know the three ways to supply an image — inline base64, a URL, or a previously uploaded file reference — and that base64 makes the request noticeably bigger.

for a middle

Explain the reuse question as the deciding factor, quantify base64's roughly one-third overhead, and name the extra failure modes a URL fetch introduces.

for a senior

Demonstrate operational judgment: failure isolation in logs, retry behaviour against expiring signed URLs, and where retention policy pushes you off uploaded files.

for a principal

Own the policy across the platform — which asset classes are uploaded and cached, what the retention and deletion contract is, and how image plumbing interacts with data-residency commitments.

## The three transports A vision request needs the image bytes to end up on the provider's side. There are three standard ways to get them there, and the choice is a plumbing decision, not a quality one — the model sees the same picture regardless. **Inline base64.** The bytes are encoded into the request body, usually as a data URL carrying the MIME type. Nothing external is involved: no bucket, no signed URL, no prior call. The costs are that base64 expands binary by about a third (4 bytes out for every 3 in, plus padding), that the whole payload travels on every call, and that a large image can push you against the provider's request-size limit long before it hits the per-image pixel limit. **A URL.** You send a link and the provider fetches it. Requests stay small and you avoid re-uploading, but you have introduced a second network hop that you do not control the timing of. Everything that can go wrong with a fetch now can go wrong inside your model call: the host is slow, the object is private, a signed URL expired between generation and fetch, an egress firewall blocks the provider, or a rate limiter on your CDN sees a burst of identical requests. URLs also require the image to be reachable from the public internet, which is a non-starter for many internal systems. **An uploaded file reference.** You upload once, get an identifier back, and pass that identifier in subsequent requests. Now the bytes cross the wire exactly once and every later request carries a short string. The costs are lifecycle: something has to track which identifiers exist, files expire or need deleting, and an identifier is scoped to the account and region you uploaded under. ## The deciding question: how many times will this image be sent? That single question resolves most cases. - **Once, and it was just produced.** A screenshot captured by a test runner two seconds ago exists only in memory. Base64 is right: hosting it merely to link it adds a storage round-trip and a lifecycle problem for a byte string you will never use again. - **Once, but it already lives somewhere public.** A CDN-hosted reference design or a product image. A URL is right: the bytes are already at rest and re-uploading them buys nothing — provided the provider can actually reach the host. - **Many times.** A regression suite replaying 400 baseline screenshots on every run would otherwise base64-encode and re-transmit the same megabytes on every run, for every model call. Upload once, reference by id. This is also where request-size pressure disappears. ## Second-order concerns **Failure isolation.** With base64, a failed request fails for one reason: the model call failed. With URLs, a failed request may be your storage, your DNS, an expired signature, or the model. If you cannot tell those apart in your logs, you will misattribute outages. Uploads separate the concerns cleanly: the upload either succeeded or it did not, before any inference is attempted. **Latency shape.** Base64 pays its cost in upload bandwidth on the request itself, which matters on slow client links and for large screenshots. URLs pay it as an opaque server-side fetch you cannot instrument. Uploads front-load the cost, so steady-state calls are fast — the right shape for an interactive loop where the same screenshot is examined repeatedly. **Privacy and retention.** Uploading creates a stored object with a retention policy; base64 does not, beyond ordinary request handling. For regulated screenshots — anything containing customer data on screen — that difference is worth checking against the provider's data policy, not assumed. **Limits.** Every provider caps per-image size, pixel dimensions and images per request, and oversized images are downscaled or rejected rather than silently accepted at full fidelity. Check the specific numbers for the model you are calling rather than carrying over a figure from another provider; the practical habit is to resize client-side to the resolution you actually need before sending, which reduces cost on every transport at once. ## What does not change Transport does not affect what the model can see. If a 4K screenshot's 11px error text is unreadable, sending it as a URL instead of base64 will not help — that is a resolution and detail-setting problem. Keep the two decisions separate: transport controls how bytes arrive, detail settings control how much of the image survives into the model's context.

  • Your provider fetches images by URL and you use short-lived signed links. What can go wrong?
    The signature can expire between generating the link and the provider fetching it, especially if the request is queued or retried. Retries are the classic trap: the retry re-sends the same expired URL and fails identically. Either give links a lifetime well beyond your retry window, regenerate the URL on each retry, or switch to an uploaded file reference so the bytes are already on the provider's side.
  • Does base64 versus URL change the token cost of the image?
    No. Image cost is driven by the pixels the model ends up processing — resolution and any detail setting — not by how the bytes were transported. Base64 inflates the HTTP payload you upload, which affects bandwidth and request-size limits, but the billed image tokens are identical across all three transports.
  • When would you deliberately avoid uploaded file references despite heavy reuse?
    When retention is the problem rather than the cost: screenshots containing customer data create a stored object with its own lifecycle and deletion obligation. If the compliance cost of managing that store outweighs the bandwidth saved, inline the bytes per request and keep nothing at rest on the provider side.

saying these in an interview costs you the question

  • Claiming URLs make images cheaper in tokens
  • Ignoring that the provider must reach the URL host
  • Re-encoding the same asset as base64 on every call
  • Assuming base64 has no size overhead
  • Treating a failed image fetch as a model failure

context