In Dart, how do Codec.fuse, Converter chaining and LineSplitter let you process a large JSON Lines byte stream without building one giant string?
answer
- Converters are stream transformers
- bytes -> utf8.decoder -> LineSplitter
- LineSplitter handles \n, \r\n and \r
- json.fuse(utf8): objects to bytes
- fused JSON and UTF-8 skip the string
basics
~20 sConverters such as utf8.decoder implement StreamTransformer, so a byte stream can be piped through utf8.decoder and LineSplitter and each line decoded separately. fuse joins codecs or converters, and JSON with UTF-8 fused skips the intermediate string.
solid answer
~40 sEvery `Converter` in `dart:convert` implements `StreamTransformerBase`, so `bytes.transform(utf8.decoder)` decodes chunk by chunk, carrying a character split across two chunks correctly. `LineSplitter`, itself a `StreamTransformer` rather than a `Converter`, then emits one string per line, splitting on `\n`, `\r\n` or `\r` and removing the terminators, so each JSON Lines record can be passed to `jsonDecode` and memory stays proportional to one line. For whole documents, `fuse` composes: `json.fuse(utf8)` is a `Codec<Object?, List<int>>` that encodes objects straight to bytes and decodes bytes back, and `JsonEncoder.fuse(utf8.encoder)` returns a `JsonUtf8Encoder` that writes UTF-8 without an intermediate `String`. On the VM, `utf8.decoder.fuse(json.decoder)` likewise parses bytes directly. Where the byte stream comes from, a file or socket, is a separate concern.
code
dart · 22 linesimport 'dart:convert';
Stream<Map<String, dynamic>> readSettingsLog(Stream<List<int>> bytes) => bytes
.transform(utf8.decoder)
.transform(const LineSplitter())
.where((line) => line.isNotEmpty)
.map((line) => jsonDecode(line) as Map<String, dynamic>);
final jsonToBytes = json.fuse(utf8); // Codec<Object?, List<int>>
Future<void> main() async {
final bytes = jsonToBytes.encode({'theme': 'dark'});
print(jsonToBytes.decode(bytes)); // {theme: dark}
final chunks = Stream.fromIterable([
utf8.encode('{"theme":"dark"}\r\n{"theme":'),
utf8.encode('"light"}\n'),
]);
await for (final entry in readSettingsLog(chunks)) {
print(entry['theme']); // dark, then light
}
}go deeper
Know that utf8.decoder and LineSplitter can be applied to a stream with transform to read text line by line.
Explain chunk-safe decoding, LineSplitter's terminator rules, and what json.fuse(utf8) returns.
Design large imports as streamed, line-delimited records, use fused byte-level converters for big documents, and keep memory proportional to one record.
Choose data formats for exports and sync with streaming in mind, so clients never need a whole document in memory to process it.
## Why streaming matters Reading a large export, for example years of settings changes in a JSON Lines file with one JSON object per line, as a single `String` means holding every byte, the decoded text and the parsed objects in memory at once. `dart:convert` is built so the same converters work on **chunks**, letting memory stay proportional to one record. ## Converters are stream transformers `Converter<S, T>` implements **`StreamTransformerBase<S, T>`**. Every converter therefore has two faces: - `converter.convert(input)` for a whole value in memory; - `stream.transform(converter)` for a stream of chunks, using its chunked conversion support. `utf8.decoder` is a `Converter<List<int>, String>`. Used as a transformer, it keeps state between chunks, so a multi-byte character split across two network packets is decoded correctly, something naive per-chunk `utf8.decode` calls get wrong. ## Splitting lines `LineSplitter` extends `StreamTransformerBase<String, String>`; it is a **stream transformer, not a `Converter`**, so it cannot be fused, but it chains with `transform`. Its rules: - A line ends at **`\n`**, **`\r\n`** or a lone **`\r`**. - The emitted lines **do not contain** the terminators. - A final non-empty line may end at the end of input without a terminator. - For a string already in memory, `const LineSplitter().convert(text)` returns a `List<String>` and `LineSplitter.split(text)` a lazy `Iterable<String>`. ## The pipeline 1. Start from a `Stream<List<int>>` of bytes, from a file, a socket or an HTTP response. 2. `.transform(utf8.decoder)` produces text chunks. 3. `.transform(const LineSplitter())` produces whole lines. 4. `.where((line) => line.isNotEmpty)` skips blank lines. 5. `.map(jsonDecode)` turns each line into a map, which you then validate. Each stage processes data as it arrives, so only the current chunk and line are held in memory. ## Fusing `fuse` composes conversions: | Expression | Result | |---|---| | `json.fuse(utf8)` | `Codec<Object?, List<int>>`: `encode` goes object to bytes, `decode` bytes to object | | `utf8.fuse(base64)` | `Codec<String, String>`: text to base64 and back | | `JsonEncoder.withIndent(' ').fuse(utf8.encoder)` | a `JsonUtf8Encoder` that writes bytes directly | | `utf8.decoder.fuse(json.decoder)` | on the VM, a specialised decoder that parses bytes without first building a string | The `Codec` documentation notes that fused codecs generally attempt to optimize the operation and can be faster than running each step separately. `JsonUtf8Encoder` is documented as equivalent to encoding to a JSON string and then UTF-8 encoding it, but without creating the intermediate string. For decoding, a fused codec runs the second codec's decoder first, then the first's, so `json.fuse(utf8).decode(bytes)` decodes UTF-8 and then JSON. ## Errors inside a pipeline - A malformed line makes `jsonDecode` throw a `FormatException` inside `map`, which arrives as an error event on the stream. With `await for`, that error is thrown out of the loop and ends the import. - To skip bad records instead, decode inside a function that catches `FormatException` and returns `null`, then filter the `null`s out, and count the skipped lines for a report. - Invalid UTF-8 bytes make `utf8.decoder` emit a `FormatException` too; use `Utf8Decoder(allowMalformed: true)` only if replacement characters are acceptable in the data. - Always validate each decoded map's shape as well, because a syntactically valid line can still be the wrong record. ## Choosing the tool - **One small settings document:** `jsonDecode(utf8.decode(bytes))` or `json.fuse(utf8).decode(bytes)`; memory is not a concern. - **A large single JSON document:** fused byte-level converters avoid one full-size string, though the parsed objects still live in memory. - **Many records:** use a line-delimited format with `LineSplitter`, so each record is parsed and released. ## Mistakes to avoid - Calling `utf8.decode` on each chunk separately, which corrupts characters split across chunks. - Expecting `LineSplitter` to keep the newline characters, or to treat `\r\n` as two line breaks. - Trying to `fuse` a `LineSplitter`; chain it with `transform` instead.
- Why not call utf8.decode on each chunk of a byte stream?A multi-byte UTF-8 character can be split across two chunks. Decoding each chunk separately then sees a truncated sequence and throws a `FormatException`, or with `allowMalformed` inserts U+FFFD. `stream.transform(utf8.decoder)` keeps the partial bytes between chunks and decodes them correctly.
- In json.fuse(utf8), which codec runs first when decoding?The second one. A fused codec encodes with the first codec and then the second, and decodes in reverse: `utf8` turns the bytes into a string, then `json` parses it. On the VM, fusing `utf8.decoder` with `json.decoder` is specialised to parse the bytes directly.
saying these in an interview costs you the question
- LineSplitter is a Converter and can be fused with json.decoder
- LineSplitter keeps the \n at the end of each line
- Decoding each chunk with utf8.decode separately is equivalent
- json.fuse(utf8) decodes JSON before UTF-8
- Streaming conversion needs a third-party package in Dart