Skip to main content
Back to Blog
Tutorial
2026-09-28

Flutter Golden Tests: Visual Regression with matchesGoldenFile

Flutter golden tests guide for matchesGoldenFile, stable fonts, custom comparators, CI baselines, and practical visual regression workflows.

Flutter Golden Tests: Visual Regression with matchesGoldenFile

Flutter golden tests compare a rendered widget or image against a committed reference image. The core matcher is matchesGoldenFile from flutter_test. It can receive a Finder, image, future image, bytes, or future bytes, and for widget finders it captures the first RepaintBoundary ancestor. The official API still documents flutter test --update-goldens as the way to create or refresh the reference images.

As of the current Flutter docs, the default flutter test comparator is LocalFileComparator. It resolves golden file keys relative to the test file, decodes PNG bytes, and performs pixel-for-pixel comparison. The docs for goldenFileComparator also show that you can replace the comparator in test setup, including with a tolerance comparator. Flutter 3.47 is the current stable documentation line on docs.flutter.dev, and the golden APIs discussed here remain active.

Use Flutter golden tests when you want fast visual regression coverage for widgets, component states, design-system variants, and carefully controlled screen slices. Do not use them as a full replacement for device-level mobile visual testing or browser visual testing. They are strongest when they run in a deterministic widget-test harness with fixed surface size, loaded fonts, mocked data, and a known Flutter SDK version.

The Exact Thing matchesGoldenFile Compares

matchesGoldenFile does not compare your intent. It compares pixels. When passed a finder, Flutter evaluates the finder, requires exactly one widget, captures the rendered image at the nearest useful repaint boundary, encodes that image, and passes the bytes to goldenFileComparator. With flutter test, that comparator normally reads a PNG from disk and returns success only for an exact match.

That simplicity is the strength of Flutter golden tests. It is also the trap. If the text style, font loader, device pixel ratio, test surface size, theme, locale, animation frame, or Flutter engine version changes, the pixels can change even when the product behavior did not.

Moving partWhy it affects pixelsHow to control it
Surface sizeLayout constraints change wrapping and overflowSet tester.view.physicalSize or use package helpers
Text scaleLarger text changes line breaks and clippingTest explicit text scales as separate cases
FontsAhem or host fonts alter glyph metricsLoad test fonts or use a package strategy
LocaleStrings and directionality change layoutWrap with MaterialApp and fixed locale
ThemeMaterial defaults can change across SDKsProvide explicit ThemeData
AnimationsCapturing a different frame changes pixelsPump to a known frame or settle intentionally
Flutter SDKRasterization and default widgets evolvePin CI SDK and review SDK upgrades separately

This is the main difference from browser visual regression testing. Browser tools often render pages in managed browser versions and compare screenshots across viewport widths. Flutter goldens run inside the Flutter test environment and are tightly coupled to the local SDK and host rendering details. For web-first visual workflows, see Playwright visual regression testing guide.

Start With A Small, Stable Widget Golden

A useful first golden test should prove one component contract in one fixed state. The code below avoids network data, provides a complete widget, sets a stable app wrapper, waits for rendering work to settle, checks that the expected widget exists, and then compares the golden.

import 'package:flutter/material.dart';
import 'package:flutter_test/flutter_test.dart';

class StatusBanner extends StatelessWidget {
  const StatusBanner({
    required this.title,
    required this.message,
    super.key,
  });

  final String title;
  final String message;

  @override
  Widget build(BuildContext context) {
    return RepaintBoundary(
      child: Material(
        color: const Color(0xfff8fafc),
        child: SizedBox(
          width: 360,
          child: Padding(
            padding: const EdgeInsets.all(16),
            child: Column(
              mainAxisSize: MainAxisSize.min,
              crossAxisAlignment: CrossAxisAlignment.start,
              children: [
                Text(title, style: Theme.of(context).textTheme.titleMedium),
                const SizedBox(height: 8),
                Text(message, style: Theme.of(context).textTheme.bodyMedium),
              ],
            ),
          ),
        ),
      ),
    );
  }
}

void main() {
  testWidgets('StatusBanner renders the warning state', (tester) async {
    await tester.pumpWidget(
      const MaterialApp(
        home: Center(
          child: StatusBanner(
            title: 'Payment needs review',
            message: 'Ask the account owner to approve the new card.',
          ),
        ),
      ),
    );

    expect(find.text('Payment needs review'), findsOneWidget);
    await tester.pumpAndSettle();

    await expectLater(
      find.byType(StatusBanner),
      matchesGoldenFile('goldens/status_banner_warning.png'),
    );
  });
}

Generate or refresh the image with:

flutter test test/status_banner_test.dart --update-goldens

Then compare it in normal test mode:

flutter test test/status_banner_test.dart

The update command is not a harmless convenience. The GoldenFileComparator.update API is invoked instead of comparison when autoUpdateGoldenFiles is true, and Flutter sets that when the test runner receives --update-goldens. In other words, update mode writes new truth. It should run intentionally, produce a code review diff of PNG files, and be tied to a human decision.

What To Commit And What To Ignore

Commit golden PNGs when they represent product intent. Ignore temporary failure output if your comparator or package writes diff artifacts under local build directories. A clean repository policy avoids the two worst outcomes: missing baselines in CI, or noisy generated artifacts committed after every failed run.

File typeCommit it?Notes
test/**/goldens/*.pngYesBaselines are test fixtures
test/**/failures/*.pngUsually noUseful locally, noisy in mainline
Package-specific CI goldensYes, if CI uses themAlchemist can separate goldens/ci from platform goldens
Screenshots from manual debuggingNoStore outside the repo or attach to issues
Generated HTML reportsUsually noUpload as CI artifacts instead

Teams sometimes avoid committing PNGs because they feel heavy. That usually creates a worse problem. A golden without a baseline is not a regression test. If a binary diff is hard to review in your code host, add CI artifacts or a small script to render before and after images in the pull request.

Host Drift: The Failure Mode You Will Hit First

The official matchesGoldenFile docs call out a major source of drift: custom fonts can render differently across platforms, and even the same platform can differ across Flutter versions. That warning is not theoretical. A golden generated on macOS can fail on Linux CI because text rasterization, font fallback, subpixel positioning, or antialiasing differs.

The symptom is frustrating: the UI is visually acceptable, but the diff shows tiny colored outlines around text. If the comparator is exact, even a small antialiasing difference fails the test. Before approving a mass baseline update, diagnose the host.

SymptomLikely causeFirst check
Every text edge differsFont or host renderer mismatchCompare local OS and CI OS
Only shadows differRasterization or blur behaviorDisable or standardize shadows in test theme
Layout wraps differentlyFont metrics, width, or text scale changedPrint surface size and test text scale
Many goldens fail after SDK upgradeFlutter engine or widget defaults changedReview Flutter release notes and update in one PR
One state fails intermittentlyAnimation or async load not settledPump to deterministic frame

The boring but effective policy is to generate and compare goldens on the same host class. If CI runs Linux, update goldens on Linux. If the team wants human-readable macOS goldens for local review, store them separately or use a package designed for platform-specific baselines.

Loading Fonts Deliberately

By default, Flutter widget tests use the Ahem test font in many contexts. Ahem is useful because it is deterministic, but it makes text look like blocks and can hide real typography issues. Packages such as golden_toolkit provide utilities for loading app fonts. You can also load fonts directly in test setup.

The following flutter_test_config.dart loads a font before tests run. It is a complete test configuration file, assuming the font asset path exists in the project and is declared in pubspec.yaml.

import 'dart:async';

import 'package:flutter/services.dart';
import 'package:flutter_test/flutter_test.dart';

Future<void> testExecutable(FutureOr<void> Function() testMain) async {
  TestWidgetsFlutterBinding.ensureInitialized();

  final fontLoader = FontLoader('Inter')
    ..addFont(rootBundle.load('assets/fonts/Inter-Regular.ttf'));

  await fontLoader.load();
  await testMain();
}

Use this when the text itself is part of the visual contract. Use Ahem or obscured text when text shape should not be tested and cross-host stability is more important. Alchemist formalizes that split: its docs describe platform tests with human-readable text and CI tests where text is obscured with Ahem to avoid host-specific rendering differences.

Choosing Native Goldens, Golden Toolkit, Or Alchemist

Native Flutter goldens are enough for many teams. Add a package when you need scenario grids, multiple device frames, font helpers, or a more opinionated separation between local and CI artifacts. Do not add a package just because a generated test looks shorter.

OptionCurrent status checkedBest useCaution
Native matchesGoldenFileOfficial Flutter API, activePrecise widget or screen-slice baselinesYou own surface, fonts, and comparator setup
golden_toolkit 0.15.0Pub documentation is available, API docs list 0.15.0Device scenarios, builders, font loading helpersCheck repository activity before new adoption
alchemist 0.14.0Pub package is active and changelog shows recent 0.14.0 changesSeparate local and CI goldens, scenario tables, diff thresholdsCI obscured text may be less useful for typography review
Percy or App PercyBrowserStack product, activeCross-browser or native-device screenshot reviewExternal service and screenshot-based billing

For native mobile app screenshots on real or hosted devices, a service workflow may be a better fit. See Percy visual testing complete guide when the review process needs dashboard approvals, branch baselines, and broader visual collaboration.

A Golden Toolkit Scenario Grid

golden_toolkit is useful when the same component must be captured across several states in one image. Its docs describe testGoldens, GoldenBuilder, pumpWidgetBuilder, screenMatchesGolden, and multiScreenGolden. The package also documents flutter test --update-goldens --tags=golden as a focused update workflow.

import 'package:flutter/material.dart';
import 'package:flutter_test/flutter_test.dart';
import 'package:golden_toolkit/golden_toolkit.dart';

class PlanPill extends StatelessWidget {
  const PlanPill({required this.label, required this.selected, super.key});

  final String label;
  final bool selected;

  @override
  Widget build(BuildContext context) {
    final color = selected ? Colors.green.shade700 : Colors.grey.shade300;
    return Container(
      padding: const EdgeInsets.symmetric(horizontal: 16, vertical: 10),
      decoration: BoxDecoration(
        color: color,
        borderRadius: BorderRadius.circular(20),
      ),
      child: Text(label, textDirection: TextDirection.ltr),
    );
  }
}

void main() {
  testGoldens('PlanPill states render correctly', (tester) async {
    final builder = GoldenBuilder.grid(columns: 2, widthToHeightRatio: 3)
      ..addScenario('Free plan', const PlanPill(label: 'Free', selected: false))
      ..addScenario('Pro plan selected', const PlanPill(label: 'Pro', selected: true));

    await tester.pumpWidgetBuilder(
      builder.build(),
      wrapper: materialAppWrapper(theme: ThemeData(useMaterial3: true)),
      surfaceSize: const Size(420, 220),
    );

    await screenMatchesGolden(tester, 'plan_pill_states');
  });
}

The advantage is review density: one PNG can show the matrix of states. The risk is overpacking. If a single golden contains 40 tiny variants, reviewers stop seeing details. Prefer scenario groups that fit on one screen without zooming.

An Alchemist CI-Oriented Example

Alchemist focuses on making golden tests easier to write and on separating platform goldens from CI goldens. Its documentation says CI tests obscure text and can disable shadow rendering to avoid platform-specific differences. Its 0.14.0 changelog includes a diffThreshold feature for comparison failures that depend on the image generation environment.

import 'package:alchemist/alchemist.dart';
import 'package:flutter/material.dart';

class EmptyInbox extends StatelessWidget {
  const EmptyInbox({super.key});

  @override
  Widget build(BuildContext context) {
    return const Card(
      child: Padding(
        padding: EdgeInsets.all(24),
        child: Text('No messages need your review.'),
      ),
    );
  }
}

void main() {
  goldenTest(
    'EmptyInbox renders the default state',
    fileName: 'empty_inbox',
    builder: () => GoldenTestGroup(
      scenarioConstraints: const BoxConstraints(maxWidth: 320),
      children: const [
        GoldenTestScenario(
          name: 'default',
          child: EmptyInbox(),
        ),
      ],
    ),
  );
}

Alchemist is appealing for design-system teams because it makes scenario composition concise. Still, decide up front whether CI goldens with obscured text match your risk. If your regressions are usually spacing, color, icon alignment, and clipping, obscured text may be perfect. If your regressions are font weight, baseline alignment, and copy wrapping, keep a readable platform-golden workflow too.

Custom Tolerance With goldenFileComparator

Flutter's docs show replacing goldenFileComparator with a custom comparator. This is a powerful escape hatch. Use it for known, tiny cross-host differences, not to hide real layout changes. A tolerance should be small, documented, and limited to the tests that need it.

import 'dart:typed_data';

import 'package:flutter/material.dart';
import 'package:flutter_test/flutter_test.dart';

class TolerantGoldenFileComparator extends LocalFileComparator {
  TolerantGoldenFileComparator(
    super.testFile, {
    required double precisionTolerance,
  })  : assert(precisionTolerance >= 0),
        assert(precisionTolerance <= 1),
        _precisionTolerance = precisionTolerance;

  final double _precisionTolerance;

  @override
  Future<bool> compare(Uint8List imageBytes, Uri golden) async {
    final result = await GoldenFileComparator.compareLists(
      imageBytes,
      await getGoldenBytes(golden),
    );

    final passed = result.passed || result.diffPercent <= _precisionTolerance;
    if (passed) {
      result.dispose();
      return true;
    }

    final error = await generateFailureOutput(result, golden, basedir);
    result.dispose();
    throw FlutterError(error);
  }
}

void main() {
  testWidgets('Icon badge allows tiny raster drift', (tester) async {
    final previous = goldenFileComparator;
    goldenFileComparator = TolerantGoldenFileComparator(
      Uri.parse('test/icon_badge_test.dart'),
      precisionTolerance: 0.001,
    );
    addTearDown(() {
      goldenFileComparator = previous;
    });

    await tester.pumpWidget(
      const MaterialApp(
        home: Center(child: Icon(Icons.verified, size: 48)),
      ),
    );

    await expectLater(
      find.byIcon(Icons.verified),
      matchesGoldenFile('goldens/icon_badge.png'),
    );
  });
}

The hidden danger is that tolerances become permanent folklore. If a test needs 0.001, explain why in the test name or helper. If it needs 0.05, it probably is not a trustworthy golden.

CI That Separates Compare From Update

CI should compare goldens on pull requests. Updating goldens should be a separate, explicit workflow or a local developer action. Mixing the two means a pull request can overwrite the evidence of its own regression.

name: flutter-goldens

on:
  pull_request:
  workflow_dispatch:

jobs:
  compare-goldens:
    runs-on: ubuntu-latest
    timeout-minutes: 20
    steps:
      - uses: actions/checkout@v7

      - uses: subosito/flutter-action@v2
        with:
          channel: stable
          cache: true

      - run: flutter --version

      - run: flutter pub get

      - run: flutter test --tags=golden

      - uses: actions/upload-artifact@v7
        if: failure()
        with:
          name: golden-failures-${{ github.run_id }}
          path: test/**/failures/*.png
          if-no-files-found: ignore

The subosito/flutter-action@v3 major should be verified before adopting in your repository. The GitHub-owned actions in this workflow use their current major versions. If your organization pins Flutter with FVM, Docker, mise, or a self-hosted runner, use that instead of channel: stable. Reproducibility matters more than the action brand.

An update workflow should be manual and should either commit regenerated PNGs to a branch or upload them for review. Many teams prefer local updates inside a devcontainer or Linux container that matches CI.

flutter --version
flutter pub get
flutter test --tags=golden --update-goldens
git status --short test

The final git status is part of the workflow. It reminds the developer that updating goldens changes source-controlled artifacts.

Directory-Level Tags And Test Selection

Flutter uses the Dart test runner under the hood, so tags are a practical way to separate golden tests from regular widget tests. Packages such as golden_toolkit encourage --tags=golden. Declare the tag so the runner recognizes it cleanly.

tags:
  golden:
    timeout: 2x

Then use the same selector in CI and local updates:

flutter test --tags=golden
flutter test --tags=golden --update-goldens

Use tags with restraint. A test tagged golden should actually write or compare a golden file. If everything is tagged, selection loses meaning and pull request feedback slows down.

What People Get Wrong About Golden Tests

The biggest mistake is treating golden tests as screenshots of the whole app. A full-screen golden is sometimes useful, but only if the state is deterministic and the reviewer can understand the contract. Most high-value Flutter golden tests are smaller: a form field with error text, a pricing card in selected and unselected states, a banner under text scaling, a component under right-to-left directionality, or a navigation rail at a breakpoint.

The second mistake is updating goldens during the same change that modifies many visual foundations. If a Flutter SDK upgrade, Material version change, app theme refactor, and feature redesign all happen in one pull request, the golden diff is unreadable. Split those changes. First upgrade the SDK and review mechanical baseline changes. Then land the feature visual change.

The third mistake is missing assertions before the visual comparison. If the widget never reached the intended state, matchesGoldenFile compares the wrong thing very precisely. Always assert presence of the state marker before comparing pixels.

A Practical Baseline Review Checklist

Give reviewers a checklist that turns PNG changes into engineering decisions. It should fit in the pull request description or team docs.

Review questionAccept whenReject when
Did the intended component change?The diff matches the described UI changeUnrelated states moved or recolored
Is text still readable?Contrast, weight, and wrapping remain acceptableText clips, overlaps, or shrinks unexpectedly
Are layout constraints stable?Fixed-size examples still fitContent depends on local window size
Is the baseline generated on the right host?Host matches CI policymacOS, Linux, and Windows baselines are mixed accidentally
Are animations deterministic?Capture frame is chosen intentionallyDiff shows transitional frames

Reviewing goldens is not rubber-stamping PNG churn. It is deciding whether the new reference image should become truth. That is why good naming, small surfaces, and stable fixtures matter more than screenshot count.

Frequently Asked Questions

When should I use Flutter golden tests instead of device screenshots?

Use Flutter golden tests for fast widget-level visual contracts: design-system components, empty states, error states, responsive slices, and text-scale variants. They run without installing the app on a device and can be much faster than end-to-end screenshot workflows. Use device screenshots when the contract depends on native system UI, platform views, real keyboard behavior, permission prompts, or full app navigation. Many mature teams use both, with goldens catching component drift early.

Why do golden tests pass locally but fail in CI?

The usual cause is host drift. Fonts, operating system rendering, Flutter SDK version, surface size, and antialiasing can differ between a developer laptop and CI. Start by printing flutter --version in CI, confirming the runner OS, and checking whether fonts are loaded identically. Generate baselines on the same host class that compares them. If tiny raster drift remains, consider a very small custom comparator tolerance, but do not use tolerance to hide layout changes.

Should golden files be updated automatically on pull requests?

No. Pull requests should compare goldens and expose diffs. Updating with flutter test --update-goldens writes new baselines, so it should be an explicit local or manual workflow followed by human review of the changed PNGs. Automatic updates in PR CI can erase the evidence of a regression. A good policy is compare on every relevant pull request, update only when the UI change is intentional, and isolate mass updates such as Flutter SDK upgrades.

Are Alchemist and golden_toolkit replacements for matchesGoldenFile?

They are helpers around Flutter's golden-testing model, not a different fundamental concept. golden_toolkit adds builders, wrappers, device scenarios, and focused golden tags. Alchemist adds a terse scenario API and separates platform and CI goldens, including obscured text for more stable CI output. Native matchesGoldenFile is still the underlying baseline idea. Choose a package when it reduces repeated setup or solves host-stability policy, not just to make the first test shorter.