A Dart backup CLI hangs forever after a spawned hashing isolate hits an unreadable file — why, and how do onError, onExit and kill prevent it?
answer
- nobody told the spawner
- errorsAreFatal defaults to true
- register listeners at spawn time
- error arrives as two strings
- kill: beforeNextEvent or immediate
basics
~20 sThe worker threw, errors are fatal by default, so it died without sending a result, and the spawner's await on its port waits forever. Pass onError and onExit ports to Isolate.spawn so a crash or silent exit is reported, and kill workers you no longer need.
solid answer
~50 s`Isolate.spawn` defaults to `errorsAreFatal: true`, so an uncaught `FileSystemException` in the worker terminates it. Nothing reaches the main isolate unless you asked: `await port.first` waits for a result that will never come, and the CLI hangs. Pass `onError: port.sendPort` and `onExit: port.sendPort` **to `spawn` itself** — adding listeners afterwards races with a worker that may already be dead. Uncaught errors then arrive as a two-element list of strings (error and stack trace), which you can wrap in a `RemoteError`; an exit without a result arrives as the `onExit` response, `null` by default. Also catch per-file errors inside the worker so one bad file does not kill the batch, and stop workers you no longer need with `isolate.kill()` — `priority: Isolate.beforeNextEvent` by default, or `Isolate.immediate`. `Isolate.run` wires all of this for you, which is a strong reason to prefer it for one-shot work.
code
dart · 30 linesimport 'dart:io';
import 'dart:isolate';
int checksum(List<int> bytes) =>
bytes.fold(0, (hash, b) => (hash * 31 + b) & 0xffffffff);
void hashWorker((SendPort, List<String>) message) {
final (reply, paths) = message;
Isolate.exit(reply, <String, int>{
for (final p in paths) p: checksum(File(p).readAsBytesSync()),
});
}
Future<Map<String, int>> hashInWorker(List<String> paths) async {
final port = ReceivePort();
await Isolate.spawn(
hashWorker,
(port.sendPort, paths),
onError: port.sendPort, // uncaught errors as [error, stack]
onExit: port.sendPort, // null if it dies without a result
debugName: 'hasher',
);
final Object? first = await port.first;
return switch (first) {
Map<String, int> hashes => hashes,
[final String error, final String? stack] =>
throw RemoteError(error, stack ?? ''),
_ => throw StateError('hasher exited without a result'),
};
}go deeper
Know that a spawned isolate's errors do not reach you unless you listen for them.
Explain onError and onExit: what each sends, why they belong in the spawn call, and that errorsAreFatal defaults to true.
Diagnose a hung await on a worker port, make every worker report crashes and exits, catch per-item failures inside it, and kill workers on timeout or teardown.
Set a supervision policy for worker isolates — restart, fail the job, or skip items — and prefer Isolate.run where that policy is simply rethrow.
## The failure The backup tool spawns a hashing worker for a batch and waits for its result: ```dart final port = ReceivePort(); await Isolate.spawn(hashWorker, (port.sendPort, paths)); final hashes = await port.first; // hangs ``` One file in the batch is unreadable. `readAsBytesSync` throws a `FileSystemException`, the worker never reaches its `Isolate.exit(reply, hashes)`, and the tool waits forever with no error message. ## Why nothing is reported - **Errors are fatal by default.** `Isolate.spawn` takes `errorsAreFatal`, defaulting to `true`: an uncaught error shuts the isolate's event loop down. - **Isolates do not propagate errors on their own.** An error in a worker is that isolate's business. The spawner only hears about it if it registered an **error listener**. - **The port has no other sender.** `port.first` completes only when a message arrives; a dead worker sends nothing, so the future never completes. ## Listening properly `Isolate.spawn` accepts two ports that turn crashes into messages: | Parameter | Message sent | When | |---|---|---| | `onError: SendPort` | A two-element `List`: the error's `toString()` and the stack trace as a string (or `null`) | For each uncaught error | | `onExit: SendPort` | The exit response, `null` unless a response was set with `addOnExitListener` | As the isolate's last act before terminating | Key points: 1. **Register them in the `spawn` call.** The equivalent methods `addErrorListener` and `addOnExitListener` exist, but the SDK warns that a fast-failing isolate may be gone before they are processed. Passing the ports to `spawn` — or spawning with `paused: true`, adding listeners, then resuming — closes that race. 2. **Errors arrive as strings.** The original exception object cannot cross the isolate boundary through the error listener; rebuild it as `RemoteError(message, stack)` or log it. The `Isolate.errors` stream does this wrapping for you. 3. **Distinguish outcomes.** With one port for everything, the first message is either the result, an error list, or `null` from `onExit`. With `errorsAreFatal: true`, an error is followed by the exit message. 4. **Consider `errorsAreFatal: false`** for a long-lived worker that should survive a bad request and keep serving; you still need `onError` to learn about it. ## Stopping workers `isolate.kill({int priority = Isolate.beforeNextEvent})` requests shutdown using the isolate's terminate capability, which the `Isolate` object from `spawn` carries: - `Isolate.beforeNextEvent` (default) — shut down when the current event finishes; - `Isolate.immediate` — shut down as soon as possible, possibly during an event. A CLI that gives up after a timeout, or a pool that is being torn down, should kill its workers rather than wait for them. An isolate also ends on its own once its entry point returns and it has no open receive ports or pending work, or when it calls `Isolate.exit`. ## Hardening the worker itself - Catch per-file errors inside the loop and record them in the result (`path -> error`) so one unreadable file does not throw away the whole batch. - Give every isolate a `debugName` such as `hasher-3` so tooling and logs identify which worker failed. - For one computation per isolate, prefer `Isolate.run`: it registers `onError` and `onExit` on its result port, rethrows typed errors in the caller, and fails with a `RemoteError` if the isolate ends without a result — the hang above cannot happen.
- Why is calling addErrorListener right after Isolate.spawn returns not reliable?The new isolate starts running concurrently and may throw and terminate before the listener request is processed; then no error is ever sent. Pass `onError` and `onExit` to `spawn`, or spawn with `paused: true`, add the listeners, and call `resume` with the isolate's pause capability.
- When would you spawn a worker with errorsAreFatal set to false?For a long-lived worker serving many independent requests, where one bad request should not take the whole worker down. Uncaught errors are then reported to `onError` listeners while the isolate keeps processing events, so the spawner still needs to listen and decide how to answer the failed request.
saying these in an interview costs you the question
- An uncaught error in a spawned isolate is rethrown in the spawner automatically.
- Isolate.spawn defaults to errorsAreFatal: false.
- Adding an error listener after spawn returns never misses errors.
- The error listener receives the original exception object.
- kill() defaults to stopping the isolate mid-event.