skip to content

In MCP, how does resources/read return binary data such as a PDF or an image?

level: middleimportance: should knowfreq 48%

answer

  1. two shapes for one contents entry
  2. not every resource is characters
  3. binary travels encoded, not raw
  4. the field is named blob
  5. mimeType tells the host how to handle it

basics

~20 s

Binary data travels as a base64 string in the blob field of a contents entry, while textual data uses the text field. An entry carries one or the other, never both, and mimeType tells the host how to interpret it.

solid answer

~40 s

A `resources/read` result is a `contents` array, and each entry is one of two shapes. A text entry carries a `text` field holding the characters directly. A binary entry carries a `blob` field holding the bytes **base64-encoded** as a JSON string. Both shapes repeat the entry's own `uri` and may carry a `mimeType` — `text/markdown`, `application/pdf`, `image/png` — which is what the host actually uses to decide whether to render it, attach it, or refuse it. An entry never has both fields: the server picks the representation that matches the data, and a server that base64-encodes plain text just makes it opaque and roughly a third larger on the wire. A single read may mix shapes across entries, so a client must branch per entry rather than assuming the whole result is text.

code

json · 15 lines
json
{
  "resultType": "complete",
  "contents": [
    {
      "uri": "file:///srv/reports/q3.md",
      "mimeType": "text/markdown",
      "text": "# Q3 summary\nRevenue up 4%.\n"
    },
    {
      "uri": "file:///srv/reports/q3-chart.png",
      "mimeType": "image/png",
      "blob": "iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mP8z8BQDwAEhQGAhKmMIQAAAABJRU5ErkJggg=="
    }
  ]
}

go deeper

for a junior

Remember that a contents entry has either text or blob, and that blob means base64-encoded bytes. Name mimeType as the field that says what kind of content it is.

for a middle

Explain why JSON forces base64 for binary, the roughly 33% size cost, and why a read result must be handled per entry because shapes can be mixed within one array.

for a senior

Show judgment on payload size: what you refuse to ship as a raw blob, what derived representation you publish instead, and how you keep mimeType accurate enough for hosts to act on.

for a principal

Own the representation policy across a server fleet — which artefacts are exposed byte-exact, which are exposed as extracted text, and how that choice bounds context cost and memory pressure for every host that reads them.

## Two content shapes, one array `resources/read` answers with a `contents` array. Each element is one of two content shapes: - **Text contents** — carries `text`, a normal JSON string holding the characters. - **Blob contents** — carries `blob`, a JSON string holding the bytes **base64-encoded**. Both shapes also repeat the entry's own `uri` and may carry a `mimeType`. The two fields are alternatives, not options on one object: an entry has `text` or `blob`, never both and never neither. The array is plural on purpose, so one URI may expand into several entries, and those entries may be of different shapes — a report resource might return a markdown summary as text and its chart as a PNG blob. A client that inspects only the first entry, or that assumes the whole result is text, will drop content or crash. ## Why base64 at all MCP is JSON-RPC, and JSON strings hold Unicode text, not arbitrary octets. A PNG's bytes are not valid UTF-8, so there is no way to put them in a JSON string unencoded. Base64 is the standard escape hatch: it maps arbitrary bytes onto a safe 64-character alphabet. The cost is size — base64 inflates payloads by roughly a third — plus the encode and decode work on both ends. That cost is the main reason servers should not reflexively base64 everything. ## Choosing the shape The rule is simply: does the data have a faithful text representation? Source code, markdown, JSON, CSV, logs and configuration all do, so they go in `text`. Images, PDFs, audio, compiled artefacts, archives and anything with a byte-exact meaning go in `blob`. Two failure modes come from getting this wrong. Base64-encoding text makes the resource unreadable to anything that inspects the payload, defeats diffing and caching heuristics, and costs 33% more bandwidth for no benefit. Conversely, forcing binary through `text` corrupts it: the encoder either throws on invalid UTF-8 or silently substitutes replacement characters, so the bytes that come out are not the bytes that went in. A subtle case is text in an unknown or legacy encoding — if the server cannot decode it reliably, returning it as a blob with a precise `mimeType` is more honest than guessing. ## What mimeType is for `mimeType` is a media type such as `text/markdown`, `application/json`, `application/pdf` or `image/png`. It appears in two places: optionally on the descriptor in `resources/list`, telling a client what a read will probably produce, and optionally on each entry in the read result, telling it what this entry actually is. The host uses it to decide handling: render markdown, show an image inline, attach a PDF, or decline content the model cannot consume. It is a declaration, not a guarantee — a client that must be safe should still validate rather than trust the label blindly, exactly as it must treat annotations and self-reported identity as untrusted. Omitting `mimeType` is legal but pushes the host into sniffing, which is worse for everyone. ## Size and practicality Nothing in the protocol caps resource size, but a read returns the whole thing in one JSON result. Large blobs mean a large in-memory JSON document on both sides, a 33% base64 surcharge, and a body that may be worthless anyway because the model cannot consume raw PDF or PNG bytes without host-side handling. Servers exposing large binary artefacts usually publish a descriptor plus an `https://` URI, or expose a narrower derived resource — an extracted text layer, a thumbnail, a summary — instead of the raw artefact. ## Envelope details in 2026-07-28 The content shapes themselves are unchanged by revision 2026-07-28, but the surrounding result is not: every result now carries a required `resultType`, `"complete"` for an ordinary answer. A client written against an older server must treat a missing `resultType` as `"complete"`. Nothing about `text` or `blob` depends on the transport — the same shapes ride stdio and Streamable HTTP identically, though on stdio the whole JSON message must still be a single line with no embedded newlines, which base64 satisfies naturally. ## Common mistakes The recurring ones are: assuming `contents[0]` is the answer; assuming every entry is text; expecting `blob` to be a URL or a file path rather than base64 data; and sending a data URI inside `text` instead of using `blob`, which works by accident in some hosts and nowhere else.

  • What is the cost of returning a large PDF as a blob, and what would you do instead?
    The whole file is base64-encoded into one JSON string, inflating it by about a third and forcing both sides to hold the full document in memory. Raw PDF bytes are also useless to the model without host-side handling. A better server publishes a narrower derived resource — extracted text, a page range, or a summary — or points at an https:// URI, keeping the read result small enough to pack into context.
  • May a single resources/read result mix text and blob entries?
    Yes. contents is an array and each entry independently carries either text or blob, so one URI can legitimately return a markdown document alongside an image. Clients must branch per entry on which field is present rather than testing the first entry and assuming the rest match.
  • If a server omits mimeType on a contents entry, what should the client do?
    mimeType is optional, so a client must cope. In practice it falls back to sniffing — the presence of text versus blob already tells it whether the payload is characters or bytes — or treats the content as opaque. That guesswork is exactly why servers should set an accurate mimeType: it is the only reliable signal the host has for rendering or refusing the content.

saying these in an interview costs you the question

  • Saying blob holds a URL or file path instead of base64
  • Base64-encoding plain text because it is safer
  • Assuming the whole contents array is one shape
  • Believing an entry can carry both text and blob
  • Treating mimeType as a guarantee rather than a declaration

context