skip to content

Automated Verification

Flutter apps are checked by widget tests on a fake binding, golden image comparisons and integration tests on real devices, backed by mocks and fakes. Interviewers ask which level catches which bug.

part ofFlutteroverview, primer and where to startread it →
on this pageshow

explore

questions

21

In Flutter's flutter_test, how do you write a widget test proving a login form shows its validation messages when submitted empty?

level: juniorimportance: must knowfreq 72%

answer

  1. testWidgets hands you a WidgetTester
  2. pumpWidget inside a MaterialApp
  3. find.byKey, find.text
  4. tap does not rebuild by itself
  5. findsOneWidget vs findsNothing

basics

~10 s

Call testWidgets, build the form with tester.pumpWidget inside a MaterialApp, tap the submit button through a finder, call tester.pump() so the rebuild happens, then expect find.text('Email is required') to match findsOneWidget.

solid answer

~30 s

A widget test is a `testWidgets('...', (WidgetTester tester) async { ... })` body. I build the screen with `await tester.pumpWidget(const MaterialApp(home: Scaffold(body: LoginForm())))` - the `MaterialApp` supplies the `Material`, `Directionality` and theme a `TextFormField` needs. Then I act and assert: `await tester.tap(find.byKey(const Key('login-submit')))`, `await tester.pump()` because a tap only dispatches pointer events and the `setState` it triggers waits for the next frame, and finally `expect(find.text('Email is required'), findsOneWidget)`. For the happy path I type with `await tester.enterText(find.byKey(const Key('login-email')), '[email protected]')`, pump, and expect the message to be `findsNothing`. Every `tester` call returns a `Future` and must be awaited.

code

dart · 30 lines
dart
import 'package:flutter/material.dart';
import 'package:flutter_test/flutter_test.dart';
import 'package:my_app/login_form.dart';

void main() {
  testWidgets('empty submit shows both validation messages', (WidgetTester tester) async {
    await tester.pumpWidget(
      const MaterialApp(home: Scaffold(body: LoginForm())),
    );

    await tester.tap(find.byKey(const Key('login-submit')));
    await tester.pump();

    expect(find.text('Email is required'), findsOneWidget);
    expect(find.text('Password is required'), findsOneWidget);
  });

  testWidgets('a valid email clears the email message', (WidgetTester tester) async {
    await tester.pumpWidget(
      const MaterialApp(home: Scaffold(body: LoginForm())),
    );

    await tester.enterText(find.byKey(const Key('login-email')), '[email protected]');
    await tester.tap(find.byKey(const Key('login-submit')));
    await tester.pump();

    expect(find.text('Email is required'), findsNothing);
    expect(find.text('Password is required'), findsOneWidget);
  });
}

go deeper

for a junior

Recall the skeleton: testWidgets, pumpWidget in a MaterialApp, a finder, tap or enterText, pump, then expect with findsOneWidget or findsNothing.

for a middle

Explain why a pump is needed after a tap - the action only dispatches events and marks elements dirty - and why every tester call must be awaited.

for a senior

Show how you keep form tests stable: keys on controls, text only for user-visible messages, injected callbacks instead of reaching into State, one behaviour per test.

for a principal

Frame widget tests as the cheapest place to pin form behaviour, and argue which form rules deserve a widget test versus a plain validator unit test.

## What a widget test is A **widget test** runs a piece of Flutter UI inside the Dart VM on your machine, without a device or an emulator. The `flutter_test` package provides the `testWidgets` function, which registers a test and hands its body a **`WidgetTester`** - the object you use to build widgets, find them, interact with them and advance time. Under the hood `testWidgets` initialises a test binding (`TestWidgetsFlutterBinding`) that replaces the real engine with a fake one: frames are produced on demand, time is simulated, and the default test surface is **800 x 600 logical pixels**. A login form's validation is the textbook target: it is pure UI logic, it depends on user input, and a regression in it is exactly the kind of thing a unit test of a validator function would miss (for example, the validator is correct but nobody wired it to the `TextFormField`). ## The four steps: build, find, act, assert 1. **Build** - `await tester.pumpWidget(widget)` attaches the widget as the root of the tree and renders one frame. Wrap the screen in a `MaterialApp` (and usually a `Scaffold`): a `TextFormField` throws "No Material widget found" without a `Material` ancestor, and text needs a `Directionality`. 2. **Find** - a **finder** describes how to locate widgets in the current tree. The global `find` object offers `find.byKey`, `find.byType`, `find.text`, `find.widgetWithText`, `find.byIcon` and more. 3. **Act** - `tester.tap(finder)`, `tester.enterText(finder, text)` and `tester.drag(finder, offset)` simulate the user. 4. **Assert** - `expect(finder, matcher)` with a count matcher such as `findsOneWidget`. ## Why a pump follows every action `tester.tap` synthesises a pointer-down and pointer-up at the centre of the target. The button's `onPressed` runs, your code calls `formKey.currentState!.validate()`, and the `FormField`s mark themselves dirty - but **nothing is rebuilt until the next frame**. In a widget test, frames only happen when you ask for one, so the assertion must come after `await tester.pump()`. Forgetting that pump is the single most common reason a first widget test fails with "Found 0 widgets" for a message that clearly appears in the running app. `tester.enterText` is slightly different: it focuses the field (the finder must be, or contain, an `EditableText`, which `TextField` and `TextFormField` do) and replaces its content as if typed on the soft keyboard. A pump afterwards is still needed before asserting on anything that the new text causes to rebuild. ## The count matchers | Matcher | Passes when the finder locates | |---|---| | `findsNothing` | zero widgets | | `findsOneWidget` / `findsOne` | exactly one widget | | `findsWidgets` / `findsAny` | one or more | | `findsNWidgets(n)` / `findsExactly(n)` | exactly n | | `findsAtLeastNWidgets(n)` / `findsAtLeast(n)` | n or more | The shorter names (`findsOne`, `findsAny`, `findsExactly`, `findsAtLeast`) are the newer spellings; both sets exist in current `flutter_test`. Asserting `findsNothing` for the error text on the happy path is as important as `findsOneWidget` on the empty submit - it proves the message is conditional. ## Picking stable finders for a form - Give the fields and the submit button explicit `Key`s (`const Key('login-email')`) and find them with `find.byKey`; labels and copy change more often than keys. - Use `find.text` for the **thing the user reads** - the validation message itself - because that is the behaviour under test. - Remember that `find.text` also matches an `EditableText` whose controller holds that text, so a field containing the same string as a message will be counted. ## Awaiting everything `pumpWidget`, `pump`, `tap`, `enterText` and `drag` all return a `Future`. A missing `await` lets two tester operations overlap, and `flutter_test` detects it with a "Guarded function conflict" error that tells you to use `await` with all Future-returning test APIs. It is a common first-test mistake and a quick one to recognise. ## What this test does not cover It proves the form's wiring and messages. How the authentication call is replaced by a fake belongs to mocking; pixel-exact appearance belongs to golden tests; the same flow on a real device belongs to an `integration_test` run. Keeping the widget test to one screen and one behaviour is what keeps it fast and deterministic.

  • Why wrap the form in MaterialApp rather than pumping LoginForm alone?
    `pumpWidget` only wraps the widget in a `View`. A `TextFormField` needs a `Material` ancestor, a `Directionality`, localizations and a theme, all of which `MaterialApp` provides; without them the test fails with errors such as "No Material widget found" before it asserts anything.
  • How would you assert that the submit callback was not called when validation fails?
    Inject the callback as a constructor parameter, pass a closure that increments a counter or records the call, tap submit with empty fields, pump, and `expect(calls, 0)`. That checks behaviour without reaching into the form's `State`.
  • What does find.text match besides Text widgets?
    It also matches `Text.rich` by its plain text and, for an `EditableText`, compares against the controller's current value - so after `enterText`, `find.text('[email protected]')` finds the field itself.

saying these in an interview costs you the question

  • Asserting right after tester.tap without a pump and expecting the message
  • Pumping LoginForm with no MaterialApp and blaming the form for the crash
  • Dropping await on tester calls because the test seemed to pass
  • Finding the submit button by its label copy when a Key is available
  • Checking only the error case and never asserting findsNothing on valid input
open as a page

With mocktail in a Flutter test, how do you stub a dependency's async method and then verify the code under test called it?

level: middleimportance: must knowfreq 58%

basics

~10 s

Declare class MockWeatherApi extends Mock implements WeatherApi, stub with when(() => api.fetchForecast('Oslo')).thenAnswer((_) async => forecast), run the code, then assert with verify(() => api.fetchForecast('Oslo')).called(1).

open as a page

In a Flutter widget test, what is the difference between tester.pump() and tester.pumpAndSettle(), and when is each the right call?

level: middleimportance: must knowfreq 66%

basics

~20 s

tester.pump() renders one frame, optionally after advancing fake time by a duration; tester.pumpAndSettle() keeps pumping in 100 ms steps until no frame is scheduled. Use pump for precise control and pumpAndSettle to finish finite animations.

open as a page

In Flutter, how does a golden test with matchesGoldenFile work, and how do you create or refresh its baseline image?

level: juniorimportance: should knowfreq 42%

basics

~20 s

matchesGoldenFile renders the widget a finder locates, encodes it as PNG and compares it pixel for pixel with a stored file. Running flutter test --update-goldens writes the current rendering as the new baseline instead of comparing.

open as a page

In Flutter, how does an integration_test test differ from a widget test, and how do you write and run one on an Android emulator?

level: juniorimportance: should knowfreq 48%

basics

~20 s

An integration_test test uses the same testWidgets and finder API, but runs the real app on a device or emulator in real time with real plugins. Call IntegrationTestWidgetsFlutterBinding.ensureInitialized() first and run flutter test integration_test with -d.

open as a page

With Dart's package:http, how do you test Flutter code that makes HTTP requests without sending anything over the network?

level: juniorimportance: should knowfreq 38%

basics

~10 s

Inject an http.Client into the class that makes requests, and in tests pass MockClient from package:http/testing.dart, whose handler receives each Request and returns an http.Response, such as Response(body, 200).

open as a page

Why does text in a Flutter golden image render as solid boxes, and how do you make goldens show the app's real fonts?

level: middleimportance: should knowfreq 36%

basics

~20 s

Under flutter test, text uses a test font whose glyphs are boxes: FlutterTest since Flutter 3.10, Ahem before. Unregistered font families fall back to it. Load real fonts with FontLoader, often once in flutter_test_config.dart, accepting that goldens become platform-sensitive.

open as a page

In Flutter golden tests, how do you cover a receipt widget in light and dark themes and at larger text scales without duplicating tests?

level: middleimportance: should knowfreq 30%

basics

~20 s

Parameterise one testWidgets body with a ValueVariant (or a loop) over theme and text scale, build the MaterialApp from the current value, and include that value in the golden file name so each combination gets its own baseline.

open as a page

For Flutter integration tests, when do you run flutter drive with a driver script instead of flutter test integration_test, and what does the driver do?

level: middleimportance: should knowfreq 36%

basics

~20 s

flutter test integration_test is the default runner. Use flutter drive when the host must receive data from the test - screenshots, performance timelines, custom reportData - or when targeting the web; its driver script runs on the host and collects that data.

open as a page

How does mockito's @GenerateNiceMocks or @GenerateMocks code generation differ from mocktail, and what decides which one a Flutter team uses?

level: middleimportance: should knowfreq 45%

basics

~10 s

mockito generates mock classes into a .mocks.dart file from @GenerateNiceMocks or @GenerateMocks when build_runner runs; mocktail needs no generation and uses closures plus registerFallbackValue. The choice is mostly build step versus runtime setup.

open as a page

Why does a mocktail stub using any() on a custom parameter type throw a StateError, and how does registerFallbackValue fix it?

level: middleimportance: should knowfreq 42%

basics

~20 s

mocktail's any() must hand back a real value of the parameter's type while it captures the call, and it only knows fallbacks for primitives and core collections; registerFallbackValue(FakeForecastRequest()) in setUpAll supplies one for your type.

open as a page

In a Flutter test, how do you stub a MethodChannel so code that calls a plugin runs without the native side?

level: middleimportance: should knowfreq 36%

basics

~10 s

Register a handler on the test messenger: TestDefaultBinaryMessengerBinding.instance.defaultBinaryMessenger.setMockMethodCallHandler(channel, (call) async => ...), return a value per call.method, throw PlatformException for errors, and pass null to remove it.

open as a page

In a Flutter widget test, why can tester.tap(find.text('Sign in')) throw or print a hit-test warning, and how do you choose the finder?

level: middleimportance: should knowfreq 46%

basics

~20 s

tap needs a finder that matches exactly one widget and a centre that actually receives the hit. Duplicated text throws an ambiguity error; an off-screen or covered button prints a warning. Prefer find.byKey or find.widgetWithText, and scroll it into view first.

open as a page

Flutter golden tests for a receipt widget pass on a developer's macOS laptop but fail in Linux CI with a sub-1% diff; why, and how do you fix it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Rendering is not guaranteed pixel-identical across hosts: real fonts, anti-aliasing and Flutter versions can differ, and LocalFileComparator demands an exact match. Generate and compare goldens on one platform - the CI one - and skip golden checks elsewhere.

open as a page

How do you run Flutter integration_test tests on Firebase Test Lab for Android, and what must the uploaded builds contain?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Add an androidTest class run by FlutterTestRunner, build a debug app APK with -Ptarget pointing at the integration test plus a test APK with assembleAndroidTest, then upload both to Test Lab as an Instrumentation test.

open as a page

A Flutter integration test of checkout hangs in CI when the Android emulator shows a location permission dialog; why can't the test tap it, and how do you handle it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The permission dialog is native Android UI, outside Flutter's widget tree, so flutter_test finders cannot see or tap it. Pre-grant the permission before the run or inject a fake permission service, and add timeouts so a blocked test fails instead of hanging.

open as a page

In Flutter tests of a forecast view model, when does a hand-written fake WeatherApi serve better than a mocktail mock, and how do you build it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Prefer a hand-written class that implements WeatherApi with in-memory data when many tests share the dependency or assert on resulting state; keep mocktail mocks for one-off stubs and for interactions that are themselves the behaviour.

open as a page

In Flutter widget tests, why does real asynchronous work such as reading a file never complete, and when should you use tester.runAsync?

level: seniorimportance: should knowfreq 38%

basics

~20 s

testWidgets runs its body in a FakeAsync zone where time moves only when you pump, so work that depends on the OS or another isolate never gets real time to finish. Wrap such calls in tester.runAsync, then pump to render the result.

open as a page

A Flutter widget test fails with 'pumpAndSettle timed out' after tapping Sign in on a login screen that shows a CircularProgressIndicator; why, and how do you fix it?

level: seniorimportance: should knowfreq 52%

basics

~20 s

An indeterminate CircularProgressIndicator repeats its animation forever, so a frame is always scheduled and pumpAndSettle never settles until its ten-minute fake-time timeout throws. Assert the loading state with pump(), then complete the faked sign-in and settle.

open as a page

In Flutter, how do you install a custom goldenFileComparator that tolerates small pixel differences, and what risk does the tolerance carry?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Subclass LocalFileComparator, override compare to call GoldenFileComparator.compareLists and accept results whose diffPercent is under a threshold, then assign it to goldenFileComparator. The risk: diffPercent is the fraction of differing pixels, so a loose threshold hides real regressions.

open as a page

How do you measure a Flutter checkout screen's scrolling performance with integration_test's traceAction, and why run it with flutter drive --profile?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Wrap the scroll in binding.traceAction with a reportKey, let a driver script turn the returned timeline into a TimelineSummary and write it to build/, and run flutter drive --profile so the numbers reflect release-like performance.

open as a page