skip to content

GraphQL defines no file type, so how does a client send a file with an operation?

level: juniorimportance: must knowfreq 52%

answer

  1. Nothing in the spec carries bytes
  2. The JSON body is the constraint
  3. Three answers, none of them specified
  4. One request, or bytes sent elsewhere
  5. Upload is a server convention

basics

~20 s

The GraphQL specification defines no binary scalar and no file transport. Two conventions fill the gap: a multipart request that carries the operation and the bytes together, or uploading to storage separately and passing a reference back through a mutation.

solid answer

~50 s

Nothing in GraphQL carries bytes. The built-in scalars are Int, Float, String, Boolean and ID, and the GraphQL over HTTP working draft describes a JSON request body, which has no byte type. So there are three real answers. **Base64 into a String variable** works anywhere but inflates the payload by roughly a third and buffers the whole file as a string on both sides — acceptable for a few kilobytes, not for a scan. **The GraphQL multipart request specification**, a community document rather than a Foundation one, sends the request as `multipart/form-data`: an `operations` field holding the JSON operation with `null` where each file goes, a `map` field binding numbered file parts to variable paths, then the file parts. Servers that implement it conventionally declare an input-only `Upload` scalar. **A signed upload URL** takes the bytes out of the graph entirely: one mutation mints a short-lived URL, the client uploads straight to storage, a second mutation attaches the stored object by reference.

code

graphql · 5 lines
graphql
scalar Upload

type Mutation {
  attachConditionPhoto(accessionId: ID!, photo: Upload!): ConditionReport!
}

go deeper

for a junior

Recall that no GraphQL scalar holds binary and that uploads are a convention layered on top. Be able to name the two mainstream routes — a multipart request, or a signed URL plus a reference — without needing the wire details.

for a middle

Explain why the JSON body forces the choice, what base64 costs in bytes and memory, and that the Upload scalar is an input-only server convention whose runtime value differs between servers. Expect a follow-up on returning files.

for a senior

Show that you pick by file size and lifecycle rather than by habit, and that you know the multipart route makes the request body non-JSON, which every intermediary in the path must tolerate. Be ready to say when you would refuse uploads on the graph at all.

for a principal

Own the policy: which asset classes are allowed through the graph, which must go to storage directly, and how that rule is enforced rather than merely documented. Consider what a per-team exception costs when the upload path is a special case at every hop.

## Nothing in the specification carries bytes The GraphQL specification defines five built-in scalar types — `Int`, `Float`, `String`, `Boolean` and `ID` — and none of them is binary. The specification also says nothing about HTTP at all. The GraphQL over HTTP working draft, which does, describes a request whose body is JSON. JSON has no byte type. That is the whole reason this question exists, and it is the sentence to lead with in an interview: a file has to reach the server either **encoded into a JSON value**, or **outside the JSON body entirely**. Every real technique is one of those two, and there are three in common use. ## Route one: base64 into a String variable Declare the variable as `String!` and send the file base64-encoded. No server feature is required and no convention has to be agreed — it is an ordinary operation with an unusually long string in it. The costs are arithmetic. Base64 inflates the payload by about a third. The whole encoded file exists as a single string in the client's memory, in the request body, in the server's parsed variables, and in anything downstream that records variables — a log line, a trace attribute, an error report. Nothing streams: the server cannot begin doing anything with the file until the last byte of the JSON body has arrived and been parsed. For a 4 KB signature capture or a small SVG this is genuinely the right call, and saying so is a better answer than reflexively reaching for machinery. For a 380 MB master scan it is not. ## Route two: the GraphQL multipart request specification This is a **community specification**. It is not part of the GraphQL specification, it is not part of the GraphQL over HTTP working draft, and it is not published by the GraphQL Foundation — candidates who say "GraphQL supports file uploads" are usually thinking of this and are usually wrong about its status. It describes sending the request as `multipart/form-data` with three kinds of field: - `operations` — the JSON operation, with `null` at every variable position a file will occupy; - `map` — a JSON object binding each numbered file field to the variable path where its bytes belong; - the file fields themselves, one per file. The server decodes the two JSON fields, substitutes each uploaded file into the variables at the mapped paths, and then executes the operation exactly as it would any other. The response is an ordinary GraphQL JSON response with `data` and `errors` — nothing about the reply is special. Servers that implement the convention conventionally declare a custom scalar named `Upload`, used only in **input** positions. No specification names that scalar. It has no literal form — you cannot write a file into a document — so a file always arrives through a variable. And what the resolver actually receives when it reads that argument differs by server: a readable stream, a temporary file path, a byte buffer, a lazily-opened handle. The SDL looks portable across servers; the resolver body is not. The headline benefit is that the file and its metadata arrive in **one mutation**, authorized by the same field authorization as any other mutation, validated together, and either committed together or not at all. ## Route three: upload out of band and pass a reference The bytes never enter the graph's request path. A mutation returns a short-lived signed URL scoped to one storage key; the client sends the bytes directly to the storage service over ordinary HTTP; a second mutation attaches the stored object to the domain record by its identifier. Three round trips instead of one, and a two-phase state to manage — but the graph stays a JSON-only API, and file size stops being the graph's problem. ## Choosing between them A museum collection graph makes the split concrete. A 240 KB condition photo attached to an accession record rides comfortably with its mutation: one request, one authorization decision, one transaction, and the photo is never stored without the condition report that explains it. A 380 MB master scan from the digitisation bench does not: it wants to go straight to storage, it wants to be retryable, and no request path in the graph should be holding it. ## What an interviewer is listening for Three things. That you know the specification defines nothing here. That you can name the multipart convention as a **convention** and describe roughly how it binds bytes to a variable. And that you treat the choice as a size-and-lifecycle trade rather than a default. The weakest answer is "you just use the `Upload` scalar" — it names a server convention as though it were a language feature, and it does not survive the next question, which is what happens when the file is a gigabyte.

  • Is base64-encoding a file into a String variable ever the right call?
    For a few kilobytes, yes. A signature capture or a small icon costs one ordinary JSON request, needs no server convention and no special handling at any hop. The costs bite with size: roughly a third more bytes on the wire, the whole file held as one string in memory on both sides, no streaming, and the payload appearing anywhere variables are logged or traced. Past a few hundred kilobytes it stops being defensible.
  • Can a GraphQL field return file bytes the same way?
    No. The multipart request convention is request-only, and an `Upload` scalar is declared for input positions — there is no meaningful serialization of a file into a JSON response. The normal answer is to return a URL, often signed and time-limited, that the client fetches over ordinary HTTP. That also keeps the bytes off the graph's response path and lets a CDN serve them.
  • Does a multipart upload change what the client gets back?
    No. The reply is the ordinary GraphQL JSON response: `data`, and `errors` if anything failed. Only the request body differs. That asymmetry matters operationally — the special handling is confined to the request side, so response caching, error handling and client parsing are unchanged, while every hop that inspects or rewrites request bodies has to tolerate a body that is no longer JSON.

Either you fold the photograph into the same envelope as the form, or you post the form with a locker number and drop the photograph at the locker.

saying these in an interview costs you the question

  • Says the GraphQL specification defines an Upload scalar
  • Thinks multipart uploads are part of GraphQL over HTTP
  • Base64 for multi-hundred-megabyte files
  • Expects a query field to return file bytes
  • Assumes every GraphQL server accepts multipart requests
  • Calls the multipart convention a Foundation specification

context