Skip to main content
Back to Blog
Guide
2026-09-28

Firebase Test Lab: Running Android and iOS Tests on Real Devices

Use firebase test lab to run Android and iOS tests on real devices, select matrices, shard suites, manage artifacts, and wire CI with cost control.

Firebase Test Lab: Running Android and iOS Tests on Real Devices

Firebase Test Lab is Google's hosted device testing service for Android and iOS. It is active, documented in the Firebase site, and still driven by the Google Cloud CLI commands gcloud firebase test android run and gcloud firebase test ios run. For Android it supports instrumentation tests, Robo tests, and Game Loop tests. For iOS it supports XCTest, XCUITest, Game Loop, and, according to the current gcloud reference, an iOS robo test type in the command surface.

The direct payoff is simple: you can move mobile tests from "works on my emulator" to repeatable matrices across hosted physical and virtual devices, with logs, screenshots, videos, flaky-attempt status, and raw artifacts stored in Cloud Storage. The service does not remove the need for good test code. It gives QA teams a controlled place to run that code across devices they do not own. If you are still defining your device coverage model, pair this guide with Mobile Device Farm Testing Strategy 2026. If your Android suite is Espresso-based, the execution details here complement Espresso Android Testing Guide.

The official docs verified for this article include https://firebase.google.com/docs/test-lab/android/command-line, https://firebase.google.com/docs/test-lab/ios/command-line, https://firebase.google.com/docs/test-lab/usage-quotas-pricing, and the Google Cloud CLI references at https://docs.cloud.google.com/sdk/gcloud/reference/firebase/test/android/run and https://docs.cloud.google.com/sdk/gcloud/reference/firebase/test/ios/run.

What Firebase Test Lab Actually Runs

Think of Firebase Test Lab as a matrix runner, not as a test framework. You bring an Android APK and optionally a test APK, or you bring an iOS XCTest zip. Test Lab schedules those artifacts against device dimensions such as model, OS version, locale, and orientation. The result is a test matrix. If any execution in the matrix fails, the matrix is treated as failed.

PlatformTest typeYou provideBest use
AndroidInstrumentationApp APK plus test APKEspresso, UI Automator, Kaspresso, Compose UI tests
AndroidRoboApp APK, optional Robo scriptCrash discovery and broad UI exploration without writing tests
AndroidGame LoopApp APK with game loop supportGame scenarios that can run without conventional UI assertions
iOSXCTest or XCUITestZipped build products with .xctestrunDeterministic iOS UI and integration tests
iOSGame LoopIPA and scenario numbersGame validation on hosted iOS devices

The most common mistake is sending every test to every device on every pull request. That gives you slow feedback, large bills on Blaze, and failure reports nobody reads. A better model is layered: one or two fast smoke devices for pull requests, a broader physical-device set on nightly, and a release-candidate matrix that covers the oldest supported OS, latest OS, common screen sizes, and the devices where your analytics show meaningful usage.

Choose Test Types With Intent

Instrumentation tests are for known behavior. Robo tests are for exploration. They should not be evaluated with the same expectations. Instrumentation answers "does checkout still work on this Pixel device running this Android version?" Robo answers "can an automated crawler find crashes or obvious UI traps?" Game Loop answers "can a game scenario run without a manual tester driving the UI?"

DecisionInstrumentationRobo
Requires test codeYesNo, unless you add Robo scripts
Assertion qualityHigh, because your tests define expected stateLower, because crash and traversal signals dominate
Best pull-request useCritical paths and regression checksUsually too noisy unless scoped
Best scheduled useFull regression matrixExploratory crash discovery
Failure ownershipTest or app owner can inspect assertionQA needs to classify crawler path and crash evidence

For iOS, the equivalent split is between XCTest or XCUITest and Game Loop. XCTest is the path for app behavior assertions. Game Loop is specialized for games that can expose scenario numbers.

Prepare Android Artifacts

For Android instrumentation, build both the app APK and the Android test APK. The package names and output paths vary by project, but the default debug outputs are usually under app/build/outputs/apk/debug and app/build/outputs/apk/androidTest/debug.

./gradlew :app:assembleDebug :app:assembleDebugAndroidTest
test -f app/build/outputs/apk/debug/app-debug.apk
test -f app/build/outputs/apk/androidTest/debug/app-debug-androidTest.apk

A minimal instrumentation run uses --type instrumentation, --app, --test, and at least one --device flag. The Firebase docs state that --device is the preferred way to specify model, Android version, locale, and orientation, and should not be mixed with legacy dimension flags such as --device-ids, --os-version-ids, --locales, or --orientations.

gcloud firebase test android run --type instrumentation --app app/build/outputs/apk/debug/app-debug.apk --test app/build/outputs/apk/androidTest/debug/app-debug-androidTest.apk --device model=Pixel2,version=30,locale=en,orientation=portrait

To discover available Android device models and versions, use the list commands before hardcoding a matrix. Device availability changes, and the docs deliberately route users to CLI discovery rather than static lists.

gcloud firebase test android models list
gcloud firebase test android versions list
gcloud firebase test android models describe Pixel2

Keep a checked-in matrix file or script close to the app. Do not rely on a CI UI field that only one release engineer remembers. Agents can update a committed file, reviewers can diff it, and QA can tie matrix changes to product support policy.

Select Android Devices Without Overfitting

A Firebase Test Lab device matrix should represent risk, not personal preference. Include a small virtual device when fast feedback matters. Include physical devices when hardware behavior matters: camera, biometrics, manufacturer WebView behavior, graphics stack, push notification behavior, or device-specific layout issues.

Matrix layerDevice ideaTriggerTimeout posture
Pull request smokeOne stable virtual Android deviceEvery PRShort, fail fast
Merge queueOne virtual plus one physical phoneBefore merge to mainModerate
NightlySeveral OS versions and form factorsScheduledLonger, artifact-rich
Release candidateSupported oldest OS, latest OS, top usage devicesManual or release branchLong enough for full regression

The official Android get-started page notes that physical-device tests are limited to 45 minutes and virtual-device tests to 60 minutes. Treat those limits as upper bounds, not goals. A 40-minute pull-request matrix becomes a tax on every engineer. If you need full coverage, schedule it.

Use Test Targets And Sharding Deliberately

The Android gcloud reference documents --test-targets for package, class, method, and annotation filters supported by AndroidJUnitRunner. Use it to keep pull-request runs focused. For example, run only smoke tests on every PR and reserve large annotations for nightly runs.

gcloud firebase test android run --type instrumentation --app app/build/outputs/apk/debug/app-debug.apk --test app/build/outputs/apk/androidTest/debug/app-debug-androidTest.apk --device model=Pixel2,version=30,locale=en,orientation=portrait --test-targets "annotation com.example.test.Smoke"

Firebase docs also describe uniform and manual sharding flags in beta, and the Android command reference exposes --num-uniform-shards and --test-targets-for-shard in the current command surface. Sharding is not a free speed button. It multiplies device executions, which can multiply cost and queueing. Use it when a long suite has independent tests and deterministic fixtures.

gcloud beta firebase test android run --type instrumentation --app app/build/outputs/apk/debug/app-debug.apk --test app/build/outputs/apk/androidTest/debug/app-debug-androidTest.apk --device model=Pixel2,version=30,locale=en,orientation=portrait --num-uniform-shards=4

If you use Android Test Orchestrator, the Firebase Android CLI docs show --use-orchestrator and --no-use-orchestrator. Orchestrator runs each test in its own instrumentation instance, which can reduce shared-state leakage. It also changes coverage collection, so the docs show different coverageFile and coverageFilePath environment variables depending on orchestrator usage.

gcloud firebase test android run --type instrumentation --app app/build/outputs/apk/debug/app-debug.apk --test app/build/outputs/apk/androidTest/debug/app-debug-androidTest.apk --device model=Pixel2,version=30,locale=en,orientation=portrait --use-orchestrator --environment-variables clearPackageData=true,coverage=true,coverageFilePath=/sdcard/Download/ --directories-to-pull /sdcard/Download

What people get wrong: they turn on sharding and flaky retries at the same time, then celebrate the green matrix without checking whether the underlying suite got less deterministic. --num-flaky-test-attempts is useful because Test Lab can mark an execution flaky when it fails first and passes on retry. It is not a substitute for owning the failure.

Run Android Robo Tests For Discovery

Robo tests require less setup. You upload the app APK and let Test Lab crawl the UI. This is useful early in a project, after navigation refactors, and before release candidates. It is less useful as a strict PR gate because failures may require classification.

gcloud firebase test android run --type robo --app app/build/outputs/apk/debug/app-debug.apk --device model=Pixel2,version=30,locale=en,orientation=portrait --timeout=5m

Use Robo results to file product bugs, missing accessibility labels, unexpected crash paths, and login-wall problems. Do not compare Robo coverage to a hand-written Espresso suite. They answer different questions.

Prepare iOS XCTest Packages

The iOS docs describe building for a Generic iOS Device, then zipping the Debug-iphoneos directory and the generated .xctestrun file. They also recommend verifying app and runner signatures with codesign --verify --deep --verbose before upload.

xcodebuild build-for-testing -workspace Shop.xcworkspace -scheme Shop -destination "generic/platform=iOS" -derivedDataPath build/ios_testlab
cd build/ios_testlab/Build/Products
zip -r MyTests.zip Debug-iphoneos Shop_iphoneos18.0-arm64.xctestrun

Run the package with gcloud firebase test ios run --test and one or more --device flags. The iOS docs also warn that tests can fail because the Xcode version used to build the test is incompatible with the default Xcode version used by Test Lab. Use --xcode-version when you need to pin a supported version.

gcloud firebase test ios run --test build/ios_testlab/Build/Products/MyTests.zip --device model=iphone14,version=16.6,locale=en_US,orientation=portrait --xcode-version=15

To discover iOS model and version IDs, use the corresponding list commands.

gcloud firebase test ios models list
gcloud firebase test ios versions list
gcloud firebase test ios models describe iphone14

The gcloud iOS reference documents additional flags including --xctestrun-file, --timeout, --num-flaky-test-attempts, --record-video, --results-bucket, --results-dir, --results-history-name, and --test-special-entitlements. Use --xctestrun-file when you need to customize or shard methods by changing the test plan metadata outside the zip.

Manage Results And Artifacts

Firebase Test Lab prints a results link when a run starts and stores summaries in the Firebase console. The Android get-started page says raw test results contain logs and app failure details and are automatically stored in a Google Cloud bucket. If you do not specify a bucket, Test Lab creates one. If you specify your own bucket, you are responsible for its storage cost.

ArtifactWhere to lookWhy QA cares
Console summaryFirebase console Test Lab historyFast pass, fail, flaky overview
LogsRaw Cloud Storage artifactsRoot-cause exceptions and runner output
ScreenshotsTest execution details and storageVisual proof of UI state
VideoExecution artifacts when recording is enabledTiming and transition diagnosis
Coverage filesPulled directory in storageRegression scope analysis when coverage is configured

Use --results-bucket and --results-dir when you need predictable artifact paths. The Android command reference warns that --results-dir must be unique for each test matrix, otherwise results can be overwritten or intermingled.

gcloud firebase test android run --type instrumentation --app app/build/outputs/apk/debug/app-debug.apk --test app/build/outputs/apk/androidTest/debug/app-debug-androidTest.apk --device model=Pixel2,version=30,locale=en,orientation=portrait --results-bucket=my-mobile-test-results --results-dir=android/pr-123/run-456 --results-history-name=android-pr-smoke

For CI, make results-dir include a unique run identifier from the CI system. In GitHub Actions, that usually means ${{ github.run_id }} or ${{ github.run_attempt }} in the workflow source. In shell scripts, pass the value as an argument rather than building paths through string concatenation in several places.

Understand Quotas, Pricing, And Cost Control

The Firebase usage, quotas, and pricing page is current and important. It says Test Lab and Android Device Streaming have Cloud API quota and testing quota included in the standard Spark and Blaze plans. Spark currently lists up to 15 total test runs per day: 10 on virtual devices and 5 on physical devices. Blaze starts with no-cost daily test time similar to Spark: 30 minutes per day on physical devices and 60 minutes per day on virtual devices. Usage above those no-cost limits is charged by device minute, rounded up to the nearest minute, with official hourly rates listed as $5 per physical-device hour and $1 per virtual-device hour.

Cost leverWhy it moves spendControl
Device countEach matrix dimension creates executionsUse narrow PR matrix and broad scheduled matrix
Physical devicesHigher billed rate than virtual devicesReserve for hardware-risk paths
TimeoutLong timeouts can consume minutes even on stuck testsSet realistic --timeout per suite
ShardsParallelism can multiply executionsShard only long, deterministic suites
ArtifactsCustom buckets can add storage chargesRetention policy and failure-focused uploads

Budget alerts are alerts, not caps. The official page explicitly warns that budget alerts do not cap usage or charges. If cost control matters, enforce it in CI: separate workflows, manual approval for release matrices, scheduled nightly limits, and a script that refuses to launch an oversized matrix on pull requests.

Build A CI Workflow That Separates Build From Device Execution

A robust CI flow has four phases: build artifacts, authenticate to Google Cloud, run the matrix, and publish a pointer to results. Do not mix build logic and matrix logic in one unreadable command. The workflow below installs the Google Cloud CLI through apt, authenticates with a service account JSON secret, builds Android APKs, and runs a PR smoke matrix. It uses only the current major versions specified for first-party GitHub Actions in this workspace.

name: firebase-test-lab-android

on:
  pull_request:
  workflow_dispatch:

jobs:
  android-smoke:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7

      - uses: actions/setup-java@v6
        with:
          distribution: temurin
          java-version: "21"

      - name: Install Google Cloud CLI
        run: |
          sudo apt-get update
          sudo apt-get install -y apt-transport-https ca-certificates gnupg curl
          curl https://packages.cloud.google.com/apt/doc/apt-key.gpg | sudo gpg --dearmor -o /usr/share/keyrings/cloud.google.gpg
          echo "deb [signed-by=/usr/share/keyrings/cloud.google.gpg] https://packages.cloud.google.com/apt cloud-sdk main" | sudo tee /etc/apt/sources.list.d/google-cloud-sdk.list
          sudo apt-get update
          sudo apt-get install -y google-cloud-cli

      - name: Authenticate Google Cloud
        run: |
          printf '%s' '${{ secrets.GCP_SA_KEY }}' > key.json
          gcloud auth activate-service-account --key-file=key.json
          gcloud config set project "${{ vars.FIREBASE_PROJECT_ID }}"

      - name: Build Android APKs
        run: ./gradlew :app:assembleDebug :app:assembleDebugAndroidTest

      - name: Run Firebase Test Lab smoke matrix
        run: |
          gcloud firebase test android run --type instrumentation --app app/build/outputs/apk/debug/app-debug.apk --test app/build/outputs/apk/androidTest/debug/app-debug-androidTest.apk --device model=Pixel2,version=30,locale=en,orientation=portrait --test-targets "annotation com.example.test.Smoke" --timeout=10m --num-flaky-test-attempts=1 --results-bucket="${{ vars.FTL_RESULTS_BUCKET }}" --results-dir="android/pr-${{ github.event.pull_request.number }}/run-${{ github.run_id }}-${{ github.run_attempt }}"

For iOS, the same idea applies, but use macOS runners for Xcode packaging and then call gcloud firebase test ios run. Keep Xcode version choices visible because Test Lab compatibility can be the difference between a real app failure and an infrastructure failure.

name: firebase-test-lab-ios

on:
  workflow_dispatch:

jobs:
  ios-xctest:
    runs-on: macos-latest
    steps:
      - uses: actions/checkout@v7

      - name: Build XCTest package
        run: |
          xcodebuild build-for-testing -workspace Shop.xcworkspace -scheme Shop -destination "generic/platform=iOS" -derivedDataPath build/ios_testlab
          cd build/ios_testlab/Build/Products
          zip -r MyTests.zip Debug-iphoneos Shop_iphoneos18.0-arm64.xctestrun

      - name: Run iOS tests in Firebase Test Lab
        run: |
          gcloud firebase test ios run --test build/ios_testlab/Build/Products/MyTests.zip --device model=iphone14,version=16.6,locale=en_US,orientation=portrait --xcode-version=15 --timeout=20m --results-bucket="${{ vars.FTL_RESULTS_BUCKET }}" --results-dir="ios/manual/run-${{ github.run_id }}-${{ github.run_attempt }}"

The iOS workflow assumes gcloud is already available on the runner or installed in an earlier organization-standard setup step. If it is not, use the same installation approach as the Android workflow or a verified organization action with a pinned major.

Failure Mode: Matrix Fails But Local Tests Pass

Scenario: Android checkout instrumentation passes on a developer emulator and fails on Firebase Test Lab only on one physical device. The matrix reports exit code 10, meaning one or more test cases did not pass. The screenshot shows the checkout screen, but the assertion for "Place order" fails.

Diagnosis:

SignalInterpretationNext check
One device fails, others passDevice or OS-specific behavior is likelyCompare screen size, API level, navigation mode, and locale
Screenshot shows the screenNavigation workedInspect hierarchy and view visibility
Button text assertion failsCopy, localization, or layout state changedPrefer resource ID plus enabled state when text is not the contract
Video shows spinner over buttonApp is still loadingAdd real idling or wait for loading view to disappear
Retry passesFlaky app state or backend data raceFix fixture setup before increasing retries

The bad fix is to add a longer sleep before the assertion. The better fix is to assert the loaded state that makes the button actionable: loading indicator gone, cart summary present, and order button enabled. If this is a Kaspresso or Espresso test, the test should assert a real side effect after tapping, such as an order confirmation id, not only a status code from an API mock.

Operational Rules For QA Teams

Write down who owns each matrix. A PR smoke matrix is owned by feature teams because it blocks their work. Nightly device coverage is owned by QA or release engineering because it tracks product risk. Release-candidate matrices are owned by the release captain because they trade time for confidence.

Keep test lab commands reproducible outside CI. A developer should be able to copy the command, swap paths to local APKs or XCTest zips, and reproduce a failure. Avoid hidden UI-only configuration. Also keep a small document with accepted failure categories: app bug, test bug, infrastructure error, device capacity issue, unsupported OS combination, and flaky retry. Without categories, every red matrix becomes a debate.

For AI coding agents, attach three things to a failure ticket: the exact gcloud command, the Firebase result link, and the artifact path. The agent can then reason from concrete evidence instead of guessing which device, test target, or timeout was used.

Firebase Test Lab Compared With A Private Device Farm

Firebase Test Lab is excellent when you want quick access to hosted devices without maintaining hardware. A private device farm is better when you need special peripherals, custom OS images, lab-only network conditions, or extremely high run volume where owned hardware is cheaper. Many mature teams use both: Firebase for broad hosted coverage and an internal bench for specialized hardware.

NeedFirebase Test Lab fitPrivate farm fit
Fast setupStrongWeak unless farm already exists
Special hardwareLimitedStrong
Device maintenanceGoogle-ownedTeam-owned
Cost predictabilityGood with matrix disciplineGood after upfront investment
CI integrationStrong through gcloudDepends on farm tooling
Debugging physical device stateArtifact-basedDirect access possible

Do not build a private lab just because a cloud run failed twice. Build it when your product risk cannot be represented by hosted devices, or when the economics of repeated testing clearly justify maintenance.

Frequently Asked Questions

Is Firebase Test Lab free?

It has no-cost quota through the Spark plan and no-cost daily test time on Blaze, but it is not unlimited. The official pricing page currently lists Spark at up to 15 test runs per day total, split between virtual and physical devices. On Blaze, usage above the no-cost daily time is billed by device minute, rounded up, with separate physical and virtual rates. Always check the official pricing page before expanding a matrix.

Can Firebase Test Lab run iOS UI tests?

Yes. Firebase Test Lab can run XCTest and XCUITest packages on hosted iOS devices. You build for testing with Xcode, zip the build products and .xctestrun file, then call gcloud firebase test ios run --test with one or more --device flags. Pay attention to Xcode compatibility. If the default Test Lab Xcode version is not compatible with your package, set --xcode-version explicitly.

Should I use Robo tests or instrumentation tests for Android?

Use instrumentation tests when you know the behavior you want to verify and need meaningful assertions. Use Robo tests when you want automated exploration, crash discovery, or quick coverage before a suite exists. Robo is useful, but it does not replace Espresso, UI Automator, Kaspresso, or Compose UI tests. Most teams use instrumentation for gates and Robo for scheduled discovery or release-candidate exploration.

How should I handle flaky results in Firebase Test Lab?

Use --num-flaky-test-attempts to identify executions that fail first and pass on retry, then treat the flaky label as evidence, not forgiveness. Inspect videos, logs, and screenshots. If the app is still loading, fix synchronization. If data varies, fix fixtures. If only one device fails, compare OS, model, screen, and locale. Keep retry counts low on PRs so real instability remains visible.