In a Flutter app using dio, how would you write a retry interceptor that retries only safe failures with capped, jittered exponential backoff?
answer
- onError plus dio.fetch
- attempt counter in requestOptions.extra
- idempotent methods or explicit opt-in
- DioExceptionType decides retryable
- full jitter under a cap
basics
~10 sOverride Interceptor.onError: skip non-idempotent methods, cancels and 4xx; retry connection errors, timeouts and 502-504 by waiting a random delay under a capped exponential bound, then resolving with dio.fetch(requestOptions), counting attempts in requestOptions.extra.
solid answer
~40 sI extend dio's `Interceptor` and override `onError`. First I decide whether the failure is retryable: the method must be idempotent (GET, HEAD, PUT, DELETE, OPTIONS) or explicitly opted in through `requestOptions.extra`; the `DioExceptionType` must be `connectionError`, `connectionTimeout` or `receiveTimeout`, or `badResponse` with 502, 503 or 504. `cancel`, `badCertificate` and other 4xx are passed on with `handler.next`. I keep an attempt counter in `requestOptions.extra`, because `dio.fetch` sends the retry back through the whole interceptor chain, including this one. The delay is full jitter — a random value between zero and `min(cap, base * 2^attempt)` — or the server's `Retry-After` when present. Then `handler.resolve(await dio.fetch(options))`, and on a final failure `handler.next(error)`.
code
dart · 62 linesimport 'dart:math';
import 'package:dio/dio.dart';
class RetryInterceptor extends Interceptor {
RetryInterceptor(this._dio, {this.maxRetries = 3});
final Dio _dio;
final int maxRetries;
final _rng = Random();
static const _idempotent = {'GET', 'HEAD', 'OPTIONS', 'PUT', 'DELETE'};
static const _gateway = {502, 503, 504};
bool _retryable(DioException e) {
final o = e.requestOptions;
final safe = _idempotent.contains(o.method.toUpperCase()) ||
o.extra['idempotent'] == true;
if (!safe) return false;
return switch (e.type) {
DioExceptionType.connectionError ||
DioExceptionType.connectionTimeout ||
DioExceptionType.receiveTimeout =>
true,
DioExceptionType.badResponse =>
_gateway.contains(e.response?.statusCode),
_ => false,
};
}
@override
Future<void> onError(
DioException err,
ErrorInterceptorHandler handler,
) async {
final options = err.requestOptions;
final attempt = (options.extra['retryAttempt'] as int?) ?? 0;
if (attempt >= maxRetries || !_retryable(err)) {
handler.next(err);
return;
}
final boundMs = min(10000, 500 * (1 << attempt));
await Future<void>.delayed(
Duration(milliseconds: _rng.nextInt(boundMs + 1)),
);
if (options.cancelToken?.isCancelled ?? false) {
handler.next(err);
return;
}
options.extra['retryAttempt'] = attempt + 1;
if (options.data is FormData) {
options.data = (options.data as FormData).clone();
}
try {
handler.resolve(await _dio.fetch<dynamic>(options));
} on DioException catch (e) {
handler.next(e);
}
}
}go deeper
Know that dio interceptors can catch errors in onError and that retries should target temporary network failures, not every error.
Walk through onError, handler.resolve versus handler.next, dio.fetch, and the DioExceptionType values that indicate a transient failure.
Justify the idempotency gate, full jitter with a cap, Retry-After, recursion bounded by extra, FormData cloning and cancellation during the wait.
Discuss where retry policy should live across clients, how its budget interacts with user-facing latency, and what server guarantees make writes retryable.
## Where retries live in dio dio 5 runs every request through an **interceptor chain**. An `Interceptor` has three callbacks — `onRequest`, `onResponse` and `onError` — each receiving a handler that either passes the value on (`next`), completes the request with a response (`resolve`), or fails it (`reject`). A retry interceptor lives in `onError`: it receives a `DioException`, and if it decides to retry, it re-sends the original `RequestOptions` with `dio.fetch` and calls `handler.resolve` with the result, so the caller's `await dio.get(...)` simply gets the successful response. dio ships no retry interceptor of its own; community plugins exist, but writing one is a common interview exercise because every decision in it is a policy decision. ## Deciding what is retryable A retry is only correct when repeating the request cannot cause harm and the failure is plausibly transient. **Method first.** GET, HEAD, OPTIONS, PUT and DELETE are idempotent by HTTP semantics; POST and PATCH are not. A POST may be opted in only when the server de-duplicates it, which the caller signals through `requestOptions.extra`. **Then the failure type:** | `DioExceptionType` | Retry an idempotent call? | Why | |---|---|---| | `connectionTimeout` | yes | no connection was established | | `connectionError` | yes | socket failure or reset, often transient on mobile | | `receiveTimeout` | yes | the server may have processed it, which is fine only because the call is idempotent | | `badResponse` 502 / 503 / 504 | yes | gateway or overload, typically transient | | `badResponse` other 4xx, 500 | no | the same request will fail the same way | | `cancel` | never | the user or app asked to stop | | `badCertificate` | never | a security failure, not a flaky network | ## Backoff that does not stampede Plain exponential backoff (500 ms, 1 s, 2 s...) makes every device that failed at the same time retry at the same time. **Full jitter** picks a uniformly random delay between zero and the exponential bound, and a **cap** stops the bound growing past what a person holding a phone will tolerate. If a 503 carries a `Retry-After` header — `err.response?.headers.value('retry-after')` — prefer it, clamped to your cap. ## Mechanics that trip people up 1. **Recursion through the chain.** `dio.fetch` runs the full interceptor chain again, including this interceptor. Storing the attempt count in `requestOptions.extra` is what bounds the total: the nested call retries until the budget is spent, then its failure propagates to the outer catch, which passes it on. 2. **Which Dio instance.** Retrying through the same `Dio` keeps the auth, logging and base URL interceptors in play. A separate bare `Dio` for retries silently drops the auth header. 3. **Bodies that cannot be replayed.** A `FormData` body is finalised on first send, and sending it again throws a `StateError`; clone it with `FormData.clone()` before retrying. A one-shot `Stream` body cannot be replayed at all. 4. **Cancellation during the wait.** Check `requestOptions.cancelToken?.isCancelled` after the delay, so a user leaving the screen does not trigger a late retry. 5. **`Interceptor`, not `QueuedInterceptor`.** A `QueuedInterceptor` handles its `onError` callbacks one at a time, so one request's backoff would delay every other failing request, and a retried request failing again would queue behind the very callback awaiting it. ## In the field-survey app A surveyor's app loads assigned sites with GETs and uploads photos with multipart POSTs. The interceptor retries the GETs freely on connection errors. The upload POST is retried only if the server accepts an idempotency key and the caller marks the request with an `extra` flag; its `FormData` is cloned for each attempt. Attempts are capped at three with a 10-second ceiling, and every retry is logged with its attempt number so field complaints can be matched to the failure type.
- Why is receiveTimeout retryable for a GET but not for a plain POST?A receive timeout means the request was sent and the response did not arrive in time, so the server may already have applied it. Repeating a GET is harmless; repeating a POST can create a duplicate record unless the server de-duplicates it with an idempotency key.
- What goes wrong if the retry interceptor uses a fresh Dio() for the retried call?The fresh instance has none of the original interceptors or `BaseOptions`, so the retry loses the auth header, logging and any base URL resolution. Retrying through the same `Dio` keeps the chain intact, and the attempt counter in `extra` is what stops the recursion.
- Why not simply retry every DioException up to three times?A `cancel` would override the user's intent, a `badCertificate` is a security signal, and a 400 or 401 fails identically every time. Blanket retries also resend non-idempotent writes and triple the load on an already struggling server.
saying these in an interview costs you the question
- Retry every DioException, including cancel and 4xx responses.
- POST can be retried after a receiveTimeout because the request obviously failed.
- Fixed one-second delays are fine; jitter is only a server concern.
- Use a separate bare Dio for retries so the interceptor cannot recurse.
- The same FormData object can be resent without cloning.