In MCP, how does resources/read return binary data such as a PDF or an image?
answer
- two shapes for one contents entry
- not every resource is characters
- binary travels encoded, not raw
- the field is named blob
- mimeType tells the host how to handle it
basics
~20 sBinary data travels as a base64 string in the blob field of a contents entry, while textual data uses the text field. An entry carries one or the other, never both, and mimeType tells the host how to interpret it.
solid answer
~40 sA `resources/read` result is a `contents` array, and each entry is one of two shapes. A text entry carries a `text` field holding the characters directly. A binary entry carries a `blob` field holding the bytes **base64-encoded** as a JSON string. Both shapes repeat the entry's own `uri` and may carry a `mimeType` — `text/markdown`, `application/pdf`, `image/png` — which is what the host actually uses to decide whether to render it, attach it, or refuse it. An entry never has both fields: the server picks the representation that matches the data, and a server that base64-encodes plain text just makes it opaque and roughly a third larger on the wire. A single read may mix shapes across entries, so a client must branch per entry rather than assuming the whole result is text.
code
json · 15 lines{
"resultType": "complete",
"contents": [
{
"uri": "file:///srv/reports/q3.md",
"mimeType": "text/markdown",
"text": "# Q3 summary\nRevenue up 4%.\n"
},
{
"uri": "file:///srv/reports/q3-chart.png",
"mimeType": "image/png",
"blob": "iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mP8z8BQDwAEhQGAhKmMIQAAAABJRU5ErkJggg=="
}
]
}go deeper
Remember that a contents entry has either text or blob, and that blob means base64-encoded bytes. Name mimeType as the field that says what kind of content it is.
Explain why JSON forces base64 for binary, the roughly 33% size cost, and why a read result must be handled per entry because shapes can be mixed within one array.
Show judgment on payload size: what you refuse to ship as a raw blob, what derived representation you publish instead, and how you keep mimeType accurate enough for hosts to act on.
Own the representation policy across a server fleet — which artefacts are exposed byte-exact, which are exposed as extracted text, and how that choice bounds context cost and memory pressure for every host that reads them.
## Two content shapes, one array `resources/read` answers with a `contents` array. Each element is one of two content shapes: - **Text contents** — carries `text`, a normal JSON string holding the characters. - **Blob contents** — carries `blob`, a JSON string holding the bytes **base64-encoded**. Both shapes also repeat the entry's own `uri` and may carry a `mimeType`. The two fields are alternatives, not options on one object: an entry has `text` or `blob`, never both and never neither. The array is plural on purpose, so one URI may expand into several entries, and those entries may be of different shapes — a report resource might return a markdown summary as text and its chart as a PNG blob. A client that inspects only the first entry, or that assumes the whole result is text, will drop content or crash. ## Why base64 at all MCP is JSON-RPC, and JSON strings hold Unicode text, not arbitrary octets. A PNG's bytes are not valid UTF-8, so there is no way to put them in a JSON string unencoded. Base64 is the standard escape hatch: it maps arbitrary bytes onto a safe 64-character alphabet. The cost is size — base64 inflates payloads by roughly a third — plus the encode and decode work on both ends. That cost is the main reason servers should not reflexively base64 everything. ## Choosing the shape The rule is simply: does the data have a faithful text representation? Source code, markdown, JSON, CSV, logs and configuration all do, so they go in `text`. Images, PDFs, audio, compiled artefacts, archives and anything with a byte-exact meaning go in `blob`. Two failure modes come from getting this wrong. Base64-encoding text makes the resource unreadable to anything that inspects the payload, defeats diffing and caching heuristics, and costs 33% more bandwidth for no benefit. Conversely, forcing binary through `text` corrupts it: the encoder either throws on invalid UTF-8 or silently substitutes replacement characters, so the bytes that come out are not the bytes that went in. A subtle case is text in an unknown or legacy encoding — if the server cannot decode it reliably, returning it as a blob with a precise `mimeType` is more honest than guessing. ## What mimeType is for `mimeType` is a media type such as `text/markdown`, `application/json`, `application/pdf` or `image/png`. It appears in two places: optionally on the descriptor in `resources/list`, telling a client what a read will probably produce, and optionally on each entry in the read result, telling it what this entry actually is. The host uses it to decide handling: render markdown, show an image inline, attach a PDF, or decline content the model cannot consume. It is a declaration, not a guarantee — a client that must be safe should still validate rather than trust the label blindly, exactly as it must treat annotations and self-reported identity as untrusted. Omitting `mimeType` is legal but pushes the host into sniffing, which is worse for everyone. ## Size and practicality Nothing in the protocol caps resource size, but a read returns the whole thing in one JSON result. Large blobs mean a large in-memory JSON document on both sides, a 33% base64 surcharge, and a body that may be worthless anyway because the model cannot consume raw PDF or PNG bytes without host-side handling. Servers exposing large binary artefacts usually publish a descriptor plus an `https://` URI, or expose a narrower derived resource — an extracted text layer, a thumbnail, a summary — instead of the raw artefact. ## Envelope details in 2026-07-28 The content shapes themselves are unchanged by revision 2026-07-28, but the surrounding result is not: every result now carries a required `resultType`, `"complete"` for an ordinary answer. A client written against an older server must treat a missing `resultType` as `"complete"`. Nothing about `text` or `blob` depends on the transport — the same shapes ride stdio and Streamable HTTP identically, though on stdio the whole JSON message must still be a single line with no embedded newlines, which base64 satisfies naturally. ## Common mistakes The recurring ones are: assuming `contents[0]` is the answer; assuming every entry is text; expecting `blob` to be a URL or a file path rather than base64 data; and sending a data URI inside `text` instead of using `blob`, which works by accident in some hosts and nowhere else.
- What is the cost of returning a large PDF as a blob, and what would you do instead?The whole file is base64-encoded into one JSON string, inflating it by about a third and forcing both sides to hold the full document in memory. Raw PDF bytes are also useless to the model without host-side handling. A better server publishes a narrower derived resource — extracted text, a page range, or a summary — or points at an https:// URI, keeping the read result small enough to pack into context.
- May a single resources/read result mix text and blob entries?Yes. contents is an array and each entry independently carries either text or blob, so one URI can legitimately return a markdown document alongside an image. Clients must branch per entry on which field is present rather than testing the first entry and assuming the rest match.
- If a server omits mimeType on a contents entry, what should the client do?mimeType is optional, so a client must cope. In practice it falls back to sniffing — the presence of text versus blob already tells it whether the payload is characters or bytes — or treats the content as opaque. That guesswork is exactly why servers should set an accurate mimeType: it is the only reliable signal the host has for rendering or refusing the content.
saying these in an interview costs you the question
- Saying blob holds a URL or file path instead of base64
- Base64-encoding plain text because it is safer
- Assuming the whole contents array is one shape
- Believing an entry can carry both text and blob
- Treating mimeType as a guarantee rather than a declaration