Using json.Decoder.Token and More, how do you decode a huge JSON array one element at a time?
answer
- you must eat the brackets yourself
- one call asks: another element?
- Decode still works between token reads
- brackets come back as a delimiter value
basics
~20 sCall Token once to consume the opening bracket, then loop while More reports another element, calling Decode into one element value each time, and finally call Token again to consume the closing bracket. Only one element is ever in memory.
solid answer
~50 sCalling `Decode` into a `[]Record` would materialise the entire array, so you descend into it instead. `(*json.Decoder).Token()` returns the next syntactic token; the first call returns a `json.Delim` holding `[`. `(*json.Decoder).More()` then reports whether another element remains in the array currently being parsed, and inside that loop plain `Decode` reads one element at a time — mixing token reads and value decodes on one decoder is supported and is the point of the API. After the loop, a final `Token()` consumes the closing `]`. Peak memory is one element rather than the whole document. `Token` also gives you object keys, which is how you skip to a payload array nested inside a wrapper object, and `(*json.Decoder).InputOffset()` gives the byte position for error messages. For a top-level stream of concatenated values you need none of this — just `Decode` in a loop until `io.EOF`.
code
go · 14 linesdec := json.NewDecoder(r)
if _, err := dec.Token(); err != nil { // the opening '['
return err
}
for dec.More() {
var rec Record
if err := dec.Decode(&rec); err != nil {
return fmt.Errorf("at byte %d: %w", dec.InputOffset(), err)
}
process(rec)
}
if _, err := dec.Token(); err != nil { // the closing ']'
return err
}go deeper
Know that decoding a whole array into a slice loads every element at once, and that encoding/json offers a way to walk an array element by element instead. Recognising the loop is enough at this level.
Write the loop from memory and explain each piece: what a token is, that a bracket comes back as a delimiter value, and why More is asked before each element rather than after.
Bring the operational reality: a syntax error surfaces only when reached, so decide whether partially emitted output is acceptable, and attach the input byte offset to errors so a failure in a huge stream is diagnosable.
Push the question upstream. If you own the producer, newline-delimited records remove the need for hand-rolled bracket handling in every consumer, and that framing choice is a platform-wide contract worth settling once.
## The problem this API exists for A report export arrives as one JSON array containing several million objects. `json.NewDecoder(r).Decode(&records)` works, and it allocates the whole slice plus every element before your code sees the first one. On a large enough document that is the difference between a service that streams and a service that falls over. `encoding/json` solves it by letting you drop below the value level to the *token* level, and then climb back up for each element. ## The three calls `(*json.Decoder).Token() (json.Token, error)` returns the next token in the stream. A `json.Token` is an `any` holding one of: a `json.Delim` (a rune-typed value for `[`, `]`, `{` or `}`), a `string`, a `float64`, a `bool`, or `nil` for JSON null. `(*json.Decoder).More() bool` reports whether there is another element in the array or object currently being parsed. It is a *structural* question, not a question about the underlying reader; it answers "is the next thing an element, or the closing delimiter?". `(*json.Decoder).Decode(v any)` is the same call you already use, and it decodes the next complete value — which, positioned inside an array, is one element. The canonical loop is: one `Token()` to eat `[`; `for dec.More() { dec.Decode(&elem); ... }`; one `Token()` to eat `]`. Forgetting either bracket is the classic bug: skip the opening one and the first `Decode` tries to read the whole array; skip the closing one and any decoding you attempt after the loop starts at the wrong place. ## Descending into a wrapper object Real payloads are rarely a bare array. More often you get `{"cursor":"abc","rows":[ ... ]}`. Because `Token` returns object keys as plain strings, you can read tokens until you see the key you want, then read the `[` and run the element loop. That is also how you *skip* a section: `Decode` into a `json.RawMessage` captures a value without interpreting it, and `Token` walks past scalars one at a time. ## What this does not change Streaming the array does not make any single element cheaper. If one element is a hundred megabytes, decoding it costs a hundred megabytes; the technique bounds the number of elements alive at once, not the size of one. It also does not validate ahead: a syntax error a million elements in surfaces a million elements in, after your loop has already processed and possibly written out everything before it. For an export pipeline that means deciding, up front, whether partial output is acceptable or whether you must stage the work before committing it. ## Locating failures `(*json.Decoder).InputOffset() int64` returns the byte offset of the decoder's current position — the end of the most recently returned token and the start of the next. In a multi-gigabyte stream, an error message that says which byte offset failed is the difference between a five-minute diagnosis and an afternoon. `*json.SyntaxError` carries an `Offset` field for the same reason. ## When you do not need any of this The token loop is for descending *into* one large composite value. If the producer instead writes records with no enclosing array — one JSON value after another, conventionally one per line — the stream is already a sequence of top-level values and the plain `Decode`-until-`io.EOF` loop reads it with no `Token` or `More` at all. That is a good reason to prefer newline-delimited output when you control the producer: the reader stays trivial, a truncated stream is detectable, and no consumer has to hand-roll bracket handling. ## Practical shape Treat the token loop as an adapter at the edge of your system: convert the incoming array into a channel or a callback over elements as early as you can, so the rest of the code never knows the document was a single array. Keeping `Token` calls scattered through business logic is how these loops become unmaintainable, because the decoder's position is invisible global state to every function that touches it.
- What does More actually report?Whether another element remains in the array or object the decoder is currently inside — a structural question about the JSON, not about whether the underlying `io.Reader` has unread bytes. That is why it is the right loop condition: it goes false exactly when the next token is the closing delimiter.
- Do you need Token and More to read a file of records written one JSON value per line?No. Those are already top-level values, so a loop of plain `Decode` calls ending on `io.EOF` reads them, and whitespace including the newlines is skipped for you. `Token` and `More` are for descending into a single large array or object that wraps the records.
- How do you report where in a multi-gigabyte stream a decode failed?`(*json.Decoder).InputOffset()` returns the byte offset of the decoder's current position, so you can wrap the error with it. A `*json.SyntaxError` also carries an `Offset` field. Without one of those, an error from deep inside a huge document tells you nothing about where to look.
- Can you mix Token calls and Decode calls on the same decoder?Yes, and the element loop depends on it: `Token` positions you inside the array, then each `Decode` reads one complete element from that position. The two share one cursor, so the only rule is that you must keep track of where that cursor is.
saying these in an interview costs you the question
- Decodes the outer array into a slice for a huge document
- Forgets to consume the opening or closing delimiter
- Thinks More reports unread bytes on the reader
- Believes Token and Decode cannot be used on one decoder
- Assumes streaming makes an oversized single element cheap