skip to content

In Apollo Client 4, which failures does RetryLink retry with its default options, and should it sit before or after ErrorLink in the chain?

level: seniorimportance: nice to knowfreq 20%

answer

  1. the stream errored, or a result arrived
  2. an errors array is still a result
  3. five tries, counting the first
  4. every retry re-runs later links
  5. outer links see only the final outcome

basics

~20 s

RetryLink retries only when the request itself errors, such as a dropped connection or a ServerError, and passes GraphQL errors through. By default it tries five times in total. ErrorLink goes before it to see only the final outcome.

solid answer

~40 s

`RetryLink` re-subscribes to the rest of the chain when that Observable **errors**. That covers a rejected `fetch`, a `ServerError`, a `ServerParseError` and multipart protocol errors. A response carrying an `errors` array is a normal result, so it is passed up and not retried: a resolver that fails on the dashboard is not retried. With no options, `attempts.max` is 5 including the first request, with no `retryIf` any error is retried (4xx statuses included), and the delay starts around 300 ms, doubles each time and is randomized by default. Each retry calls `forward` again, so links after `RetryLink`, such as `SetContextLink`, run on every attempt. Put `ErrorLink` before `RetryLink` so that it reports once, after retries are exhausted; placed after, it sees every failed attempt.

code

ts · 19 lines
ts
import { ApolloLink, HttpLink, ServerError } from "@apollo/client";
import { ErrorLink } from "@apollo/client/link/error";
import { RetryLink } from "@apollo/client/link/retry";
import { OperationTypeNode } from "graphql";

const retryLink = new RetryLink({
  delay: { initial: 500, max: 8000, jitter: true },
  attempts: (attempt, operation, error) => {
    if (attempt >= 4) return false;
    if (operation.operationType === OperationTypeNode.MUTATION) return false;
    return !(ServerError.is(error) && error.statusCode < 500);
  },
});

const reportLink = new ErrorLink(({ error, operation }) => {
  console.error(`${operation.operationName ?? "anonymous"} failed after retries`, error);
});

export const link = ApolloLink.from([reportLink, retryLink, new HttpLink({ uri: "/graphql" })]);

go deeper

for a junior

Recall that RetryLink re-sends a request that failed on the network, and that it is added to the link chain like any other link.

for a middle

Explain that RetryLink acts on stream errors, not on GraphQL errors in a result, and give its defaults: five total attempts, retry on any error, and a randomized, growing delay.

for a senior

Show how placement decides what ErrorLink observes and whether retries re-read the token, and configure retryIf and attempts so 4xx answers and mutations are not blindly retried.

for a principal

Decide the client's retry budget as a policy, balancing perceived reliability on flaky networks against duplicated writes and extra load on a struggling backend.

## What RetryLink watches `RetryLink` is a non-terminating link from `@apollo/client/link/retry`. It calls `forward(operation)`, subscribes, and watches for the Observable to **error**. In Apollo Client that happens for **network-side failures**: - `fetch` rejecting, for example because the office Wi-Fi dropped mid-request; - a `ServerError` for a non-2xx status, such as a gateway `502` or `401`; - a `ServerParseError` for a body that is not JSON; - protocol errors inside a multipart response. A response whose body contains an `errors` array is **not** an error at the link level. It is a result, emitted with `next`. `RetryLink` passes it straight up. On an internal analytics dashboard where the `churnForecast` resolver fails, `RetryLink` does nothing, which is usually right: a resolver that failed deterministically would fail again. If a GraphQL error should be retried, an `ErrorLink` can do it once with `forward(operation)`. ## The defaults `new RetryLink()` with no options behaves as follows: | Option | Default | Meaning | |---|---|---| | `attempts.max` | `5` | total tries, **including** the first request | | `attempts.retryIf` | none | retry on any error | | `delay.initial` | `300` | first retry waits about 300 ms on average | | `delay.max` | `Infinity` | no cap on a single delay | | `delay.jitter` | `true` | delays are randomized | Two consequences come up in interviews: 1. **`max: 5` means four retries**, not five. `max: 1` disables retrying. 2. **Without `retryIf`, 4xx answers are retried too.** A `400` from a bad variable or a `401` from an expired token will not improve on a retry. Pass `attempts: { retryIf: (error) => ... }` that returns `false` for `ServerError`s below 500. Both `attempts` and `delay` also accept functions `(attempt, operation, error) => ...` for custom strategies. For example, the attempts function can refuse to retry mutations by checking `operation.operationType`, because a timed-out mutation may already have been applied. Whether an operation is safe to repeat is a general resilience question. The link only gives you the hook to encode the answer. ## Every retry re-runs the links after it A retry is another `forward(operation)` call, so **every link after `RetryLink` runs again on each attempt**. That shapes the order: - **`SetContextLink` after `RetryLink`**: each attempt reads the current token, so a token that rotates during a long backoff is still valid. - **`HttpLink` last**, as always. ## ErrorLink before or after Where `ErrorLink` sits decides what it sees: | Position | What `ErrorLink` sees on a flaky network | |---|---| | before `RetryLink` (outer) | only the final error, once retries are exhausted | | after `RetryLink` (inner, next to `HttpLink`) | every failed attempt, including ones a later retry recovers | For logging and user-facing reporting, the **outer** position is what you want: one report per operation, and no alarms for blips that `RetryLink` absorbed. The inner position is useful only for counting individual attempts, and it produces noisy logs. A token-refreshing `ErrorLink` also belongs outside, with `RetryLink` configured not to retry 401s, so the expiry reaches it immediately. ## A flaky connection, step by step With `ErrorLink → RetryLink → SetContextLink → HttpLink` and default `RetryLink` options, a panel query on an unstable connection behaves like this: 1. The first `fetch` rejects. `RetryLink` counts attempt 1, the default attempts check allows a retry, and it waits a randomized delay averaging about 300 ms. 2. The retry re-runs `SetContextLink` and `HttpLink`, and it fails again. The next wait averages about 600 ms. 3. The third try succeeds. The result flows up, and the outer `ErrorLink` never runs its handler. 4. Had all five tries failed, `RetryLink` would have passed the last error up, and the `ErrorLink` would have reported it once. ## Teardown and the 4.x details - When nothing subscribes to the operation any more, for example after the only panel using it unmounts, `RetryLink` cancels the pending timer and the in-flight attempt, so abandoned panels do not keep retrying. - In Apollo Client 4 the `error` passed to `retryIf`, `attempts` and `delay` is an `ErrorLike`, so the same `ServerError.is(...)` checks you use elsewhere apply. - Apollo Client 4 applies `errorPolicy` after the link chain, so `RetryLink` behaves the same whatever policy the query uses.

  • Why is a failing resolver on an Apollo Client 4 dashboard not retried by RetryLink, and what could retry it?
    The server still answered, and the `errors` array arrives as a normal result through `next`, not as an Observable error, so `RetryLink` passes it up. If the failure is known to be transient, an `ErrorLink` can check `CombinedGraphQLErrors.is(error)` and return `forward(operation)` to replay once. Most resolver failures are deterministic, though, so showing the error under `errorPolicy: "all"` is usually better.
  • In Apollo Client 4, how do you stop RetryLink from repeating a mutation that timed out?
    Pass a function as `attempts`. It receives `(attempt, operation, error)` and can return `false` when `operation.operationType` is `OperationTypeNode.MUTATION`. A timed-out mutation may already have been applied on the server, so an automatic resend can double-apply it unless the API makes the mutation idempotent. Queries keep their retries.

saying these in an interview costs you the question

  • RetryLink retries responses that contain GraphQL errors in the errors array.
  • attempts.max of 5 means five retries after the first request.
  • RetryLink skips 4xx responses by default, so no retryIf is needed.
  • Place ErrorLink next to HttpLink so it reports each operation once.
  • A retry reuses the original request, so later links do not run again.