skip to content

How do you measure a Flutter checkout screen's scrolling performance with integration_test's traceAction, and why run it with flutter drive --profile?

level: seniorimportance: nice to knowfreq 26%

answer

  1. binding.traceAction around the scroll
  2. reportKey names the timeline
  3. TimelineSummary on the host
  4. profile build, not debug
  5. watchPerformance for frame timings

basics

~20 s

Wrap the scroll in binding.traceAction with a reportKey, let a driver script turn the returned timeline into a TimelineSummary and write it to build/, and run flutter drive --profile so the numbers reflect release-like performance.

solid answer

~40 s

In the test I keep the binding from `IntegrationTestWidgetsFlutterBinding.ensureInitialized()` and wrap the interaction: `await binding.traceAction(() async { await tester.fling(find.byType(ListView), const Offset(0, -500), 1000); await tester.pumpAndSettle(); }, reportKey: 'checkout_scroll')`. `traceAction` records the VM timeline during the action and stores it in `reportData` under that key (default `timeline`). A driver script calls `integrationDriver(responseDataCallback: ...)`, rebuilds a `Timeline` from the JSON, summarises it with `TimelineSummary.summarize` and writes files with `writeTimelineToFile(..., includeSummary: true)` into `build/`. I run `flutter drive --driver=test_driver/perf_driver.dart --target=integration_test/checkout_scroll_test.dart --profile` because debug builds are far slower than what users run, so their frame times mislead; the Flutter docs also add `--no-dds` on mobile devices. `watchPerformance` is the lighter option: it stores a `FrameTiming` summary under `performance`.

code

dart · 23 lines
dart
// integration_test/checkout_scroll_test.dart
import 'package:flutter/material.dart';
import 'package:flutter_test/flutter_test.dart';
import 'package:integration_test/integration_test.dart';
import 'package:shop/checkout/checkout_screen.dart';

void main() {
  final IntegrationTestWidgetsFlutterBinding binding =
      IntegrationTestWidgetsFlutterBinding.ensureInitialized();

  testWidgets('checkout list scrolls smoothly', (WidgetTester tester) async {
    await tester.pumpWidget(MaterialApp(home: CheckoutScreen(cart: largeDemoCart)));
    await tester.pumpAndSettle();

    final Finder list = find.byType(ListView);
    await binding.traceAction(() async {
      await tester.fling(list, const Offset(0, -600), 2000);
      await tester.pumpAndSettle();
      await tester.fling(list, const Offset(0, 600), 2000);
      await tester.pumpAndSettle();
    }, reportKey: 'checkout_scroll');
  });
}

go deeper

for a junior

Recall that integration_test can record performance with binding.traceAction and that the measurement is run with flutter drive in profile mode.

for a middle

Explain the path of the data: traceAction stores a timeline in reportData under a reportKey, the driver's responseDataCallback summarises it with TimelineSummary and writes files to build/.

for a senior

Build a repeatable perf check: warm-up passes, fixed data, a physical device, profile builds, and summary artefacts compared per commit to catch regressions.

for a principal

Decide which interactions get frame-time budgets and how regressions block merges, balancing device-lab cost against the risk of shipping jank.

## What you are measuring A **performance integration test** runs a real interaction - scrolling the checkout's list of items and delivery options - on a device, records what the engine did during it, and saves the numbers so you can compare runs. `integration_test` provides two tools on its binding: | Method | Records | Stored under | Typical use | |---|---|---|---| | `traceAction(action, {streams, retainPriorEvents, reportKey})` | the VM timeline during `action` | `reportKey`, default `timeline` | detailed build and raster timings, summarised on the host | | `watchPerformance(action, {reportKey})` | `FrameTiming` callbacks plus GC counts | `reportKey`, default `performance` | a quick frame-time summary without a timeline | Both write into the binding's **`reportData`**, which is sent to the host only when a **driver script** is running - that is why performance tests use `flutter drive`. ## Step 1: the on-device test - Keep the binding: `final binding = IntegrationTestWidgetsFlutterBinding.ensureInitialized();` - Build the checkout screen with realistic data (enough rows to scroll). - Wrap only the interaction you want to measure in `traceAction`. - Give every call a distinct `reportKey`; a second call with the same key overwrites the first. ## Step 2: the driver on the host The default `integrationDriver()` just writes all `reportData` into `integration_response_data.json`. For a readable result, pass a `responseDataCallback` that: 1. reads `data['checkout_scroll']` and rebuilds a `Timeline` with `Timeline.fromJson` from `package:flutter_driver`; 2. calls `TimelineSummary.summarize(timeline)`; 3. calls `summary.writeTimelineToFile('checkout_scroll', pretty: true, includeSummary: true)`. That produces, in `build/` (the test outputs directory), a `.timeline_summary.json` with aggregate numbers such as average and worst frame build and raster times, and a `.timeline.json` you can open in a trace viewer. ## Step 3: run it in profile mode ``` flutter drive \ --driver=test_driver/perf_driver.dart \ --target=integration_test/checkout_scroll_test.dart \ --profile ``` - **Why `--profile`**: `flutter drive` builds in debug by default. Debug builds run with assertions and without the optimisations of release builds, so frame times are much worse than users see and can point at problems that do not exist. Profile mode keeps release-like performance while leaving the tracing hooks the test needs. - The Flutter performance cookbook adds **`--no-dds`** when running on a mobile device or emulator, because the Dart Development Service would not be reachable from the host. - Measure on a **physical device** when the numbers matter; emulators share the host's CPU and GPU and vary run to run. ## Making the numbers useful - Warm up: scroll once outside `traceAction`, then measure, so first-use shader and image work does not dominate. - Keep data and device fixed between runs; compare trends, not single values. - Store the summary files as CI artefacts and plot the key numbers per commit. - `watchPerformance` waits for frame timings to be flushed (it delays in two-second steps), so it adds seconds to the run; that is expected, not a hang. ## What it does not tell you A timeline shows **that** frames were slow and which phase took the time; finding **why** - an expensive `build`, a `saveLayer`, a large image decode - is DevTools work on the same profile build. The integration test's job is to catch regressions automatically.

  • What happens if two traceAction calls in one test use the default reportKey?
    Both store their timeline under `timeline` in `reportData`, so the second overwrites the first and only one reaches the driver. Give each call its own `reportKey` and handle each key in the driver's `responseDataCallback`.
  • When would you choose watchPerformance over traceAction?
    When a frame-time summary is enough: `watchPerformance` collects `FrameTiming` data and GC counts and stores a summary under `performance`, with no timeline to post-process. Choose `traceAction` when you need per-phase detail or a trace to open in a viewer.

saying these in an interview costs you the question

  • Debug-mode frame times are close enough to what users see
  • traceAction results appear on the host without a driver script
  • Several traceAction calls can share the default reportKey safely
  • An emulator gives stable, representative frame timings
  • The timeline summary explains which widget caused the jank