skip to content

Flutter golden tests for a receipt widget pass on a developer's macOS laptop but fail in Linux CI with a sub-1% diff; why, and how do you fix it?

level: seniorimportance: should knowfreq 44%

answer

  1. the same widget, different pixels
  2. anti-aliasing and text rasterisation
  3. read the failures folder first
  4. one platform generates and compares
  5. tolerance is the last resort

basics

~20 s

Rendering is not guaranteed pixel-identical across hosts: real fonts, anti-aliasing and Flutter versions can differ, and LocalFileComparator demands an exact match. Generate and compare goldens on one platform - the CI one - and skip golden checks elsewhere.

solid answer

~50 s

`LocalFileComparator` passes only on an exact pixel match, and the same widget is not guaranteed to render identically on macOS and Linux - real fonts especially, and sometimes anti-aliased edges and gradients, rasterise slightly differently, so a golden recorded on a laptop fails on a Linux runner by a fraction of a percent. First read the `failures` output: if `isolatedDiff.png` shows scattered edge pixels, it is platform drift; if it shows a moved digit, it is a real regression. Then pick one reference platform: record goldens on it (a Linux container locally, or a CI job that runs `flutter test --update-goldens` and uploads the PNGs for review), and run golden tests only there - skip them on other hosts or put them behind a tag. Also pin the Flutter version, the view size and theme. A tolerant comparator is the last resort.

code

dart · 30 lines
dart
import 'dart:io';

import 'package:flutter/material.dart';
import 'package:flutter_test/flutter_test.dart';
import 'package:shop/receipt/receipt_card.dart';

void main() {
  testWidgets(
    'receipt card, dark theme',
    (WidgetTester tester) async {
      await tester.pumpWidget(
        MaterialApp(
          theme: ThemeData(brightness: Brightness.dark),
          home: Scaffold(
            body: RepaintBoundary(
              key: const Key('receipt'),
              child: ReceiptCard(receipt: sampleReceipt),
            ),
          ),
        ),
      );
      await expectLater(
        find.byKey(const Key('receipt')),
        matchesGoldenFile('goldens/receipt_dark.png'),
      );
    },
    // Goldens are generated and compared on Linux only.
    skip: !Platform.isLinux,
  );
}

go deeper

for a junior

Recall that golden images can differ slightly between operating systems, so goldens are recorded and checked on one agreed platform.

for a middle

Explain what differs across hosts - anti-aliasing, text rasterisation, Flutter version - and how to read the four failure images the comparator writes.

for a senior

Separate drift from regressions from the isolated diff, then enforce one reference platform, a pinned toolchain and reviewed regeneration after upgrades instead of loosening tolerances.

for a principal

Own the golden workflow: where baselines are generated, who approves them, and how upgrades are absorbed without eroding trust in the suite.

## Why the same widget produces different pixels A golden test compares a PNG rendered **on the machine running the test** with one committed from wherever it was generated. Under `flutter test` the default comparator, `LocalFileComparator`, reports failure if a single pixel differs. Several things legitimately differ between a macOS laptop and a Linux CI runner: - **Text rasterisation.** With real fonts loaded, glyph edges, hinting and sub-pixel positioning differ by platform - the most common cause. The default box-glyph test font removes most of this. - **Low-level rasterisation.** `flutter test` renders in software (the tool passes `--enable-software-rendering --skia-deterministic-rendering` unless Impeller is enabled for tests), which is repeatable on one kind of machine, but anti-aliased curves, blurs and gradients can still come out a colour step apart on a different OS or CPU architecture, such as an ARM laptop versus an x86-64 runner. - **Flutter version.** A different engine on each host - a laptop on a newer channel than CI - changes rendering independently of the OS. The flutter_test documentation says it plainly: a golden generated on Windows with fonts will likely differ from one produced by another operating system, and even the same platform on a different Flutter version may fail. ## Diagnose before you change anything 1. Open the `failures` folder the comparator wrote next to the golden: `*_masterImage.png`, `*_testImage.png`, `*_isolatedDiff.png`, `*_maskedDiff.png`. 2. Read the failure message: "Pixel test failed, 0.37%, 1790px diff detected." A tiny percentage with differences scattered along edges and text is **drift**. A compact cluster - a price moved, a line wrapped - is a **regression**, whatever the percentage. 3. Check the sizes: "image sizes do not match" means the view size, device pixel ratio or layout differ, not rasterisation. 4. Check `flutter --version` on both hosts. ## Fixes, in order of preference | Fix | What it does | Cost | |---|---|---| | One reference platform | goldens are generated and compared only on the CI OS | local runs skip goldens or use a container | | Pinned toolchain | same Flutter version everywhere (for example via a version manager) | discipline on upgrades | | Box test font for most goldens | removes most text rasterisation drift | real typefaces not covered | | Deterministic inputs | fixed view size, theme, locale, data; no animations mid-flight | test setup | | Tolerant comparator | accepts a small fraction of differing pixels | can hide real regressions | **One reference platform** is the standard answer. In practice: - Generate goldens with the same OS image as CI - run the tests in a Linux container locally, or add a CI job that runs `flutter test --update-goldens` on request and publishes the new PNGs for a reviewer to accept. - Guard golden tests so they only run there, for example `skip: !Platform.isLinux` on the golden `testWidgets`, or put them behind a test tag that only the Linux job selects. - After a **Flutter upgrade**, regenerate all goldens in one reviewed change rather than letting them fail one by one. ## Why not just raise a tolerance A tolerance is measured as the **fraction of pixels** that differ. On a 1080 x 1920 capture, 0.5% is over ten thousand pixels - more than enough to hide a changed digit in the receipt total. If you do add one, keep it tiny, scope it to the few goldens that need it, and still review failures visually. A tolerance treats the symptom; a single reference platform removes the cause. ## Preventing the next surprise - Keep light and dark receipt goldens in the same platform-guarded suite so both are generated together. - Put generation instructions in the repository so a new team member does not record goldens on their laptop. - Treat a golden diff in review as a design review: look at the images, not just the green check.

  • The diff is 0.02% but concentrated on one digit of the total; is it drift?
    No. Drift is spread thinly along edges and text across the image. A compact cluster on one glyph means the content or position changed - a real regression - however small the percentage. That is exactly why percentage thresholds are a poor substitute for looking at `isolatedDiff.png`.
  • Why regenerate every golden in one change after a Flutter upgrade?
    An engine upgrade can shift anti-aliasing or text shaping everywhere at once. Regenerating on the reference platform in one reviewed change separates 'the toolchain moved' from 'our UI changed', and a reviewer can skim the diffs for anything that is more than rendering noise.
  • Can the failure be caused by something other than the OS?
    Yes: a different Flutter version on the two hosts, a view size or device pixel ratio that differs (reported as 'image sizes do not match'), an animation captured mid-flight, or data that depends on the current date. Rule those out before blaming rasterisation.

saying these in an interview costs you the question

  • flutter test renders every golden pixel-identically on any host OS and CPU
  • Run --update-goldens in CI on every build so the job stays green
  • Raise the tolerance until CI passes and move on
  • A small diff percentage always means harmless rendering noise
  • Each developer should record goldens on their own machine