Skip to main content
Back to Blog
AI Testing
2026-09-28

Browserbase for AI Agents and Test Automation: Sessions, Stagehand, and Debugging

Browserbase guide for QA engineers: create cloud browser sessions, connect Playwright, use Stagehand, persist contexts, debug replays, and control cost.

Browserbase for AI Agents and Test Automation: Sessions, Stagehand, and Debugging

Browserbase is a hosted browser infrastructure platform for AI agents, browser automation, web data retrieval, and end-to-end testing. The current docs describe a broader platform than just remote Chrome: Browser Sessions, Fetch, Search, Agents, Functions, Agent Identity, Model Gateway, Contexts, Proxies, Session Inspector, Live View, and Session Replay all sit behind the same Browserbase account model.

For QA engineers, the useful mental model is simple: Browserbase gives you cloud Chromium sessions that your existing tools can drive over CDP or WebDriver, plus observability and persistence features that are painful to build yourself. Playwright still writes the deterministic test. Stagehand adds AI actions, extraction, and observation when selectors are unstable or the workflow benefits from natural-language steps. Browserbase supplies the remote browser fleet, session URLs, recordings, logs, live debugging, proxies, and contexts.

Stagehand 4.0.0 shipped on August 10, 2026 and @browserbasehq/stagehand@4.1.0 (September 9, 2026) is the npm latest tag, while 3.7.3 remains available under the v3-latest tag for teams that have not migrated. The v4 docs emphasize TypeScript, Python, and Go APIs, and the samples below use the v4 browserbase.launch() entry point. Browserbase and Stagehand are active, but their docs and packages move quickly. Pin package versions, read the migration notes before upgrading, and keep plain Playwright coverage for critical assertions. For related agent-browser patterns, compare Stagehand AI Browser Automation Guide 2026 and Skyvern AI Browser Automation Guide.

Product Surface QA Teams Actually Use

The Browserbase documentation currently positions Browserbase as a complete platform for browser agents: cloud browsers, web search, page fetching, sandbox runtime, and model access. QA teams do not need every product on day one. Most testing programs start with sessions, then add contexts, proxies, replay, and Stagehand only when those solve a real bottleneck.

Browserbase surfaceWhat it doesQA use case
Browser SessionsIsolated cloud browser instances with connection URLsRun Playwright, Puppeteer, Selenium, or Stagehand remotely
Session InspectorLive debugging view with browser state, network, console, metrics, and replayTriage CI failures without guessing from logs only
Session ReplayHLS replay metadata and page playlistsAttach recordings to failed builds or internal dashboards
ContextsPersist Chromium user data across sessionsAvoid repeated login and keep authenticated state
ProxiesBuilt-in or external proxy routing with geolocation optionsValidate regional behavior and identity-sensitive flows
StagehandAI primitives plus Playwright-style page APIsHandle changing UI where deterministic selectors are too brittle
FunctionsRun TypeScript next to Browserbase browsersMove automation close to the browser for latency and deployment

The trap is using all of it because it exists. Start with a remote Playwright smoke suite, capture replay links for failures, and add context persistence for logins. Once that is stable, introduce Stagehand for the small set of actions where AI actually reduces maintenance.

Create Sessions And Connect Playwright

The core Browserbase loop is create, connect, use, close. The Sessions API returns an id, connectUrl, region, status, expiration data, and related metadata. The docs show @browserbasehq/sdk with playwright-core for Node.js, and the project can be inferred from the API key.

npm install --save-dev playwright-core @browserbasehq/sdk

This Playwright sample checks required environment, creates a session, connects over CDP, uses the default context and page, asserts useful content, and closes the browser. It avoids relying on a local browser binary.

import assert from "node:assert/strict";
import { chromium } from "playwright-core";
import { Browserbase } from "@browserbasehq/sdk";

const apiKey = process.env.BROWSERBASE_API_KEY;
assert.ok(apiKey, "BROWSERBASE_API_KEY is required");

const bb = new Browserbase({ apiKey });
const session = await bb.sessions.create({
  region: "us-west-2",
  userMetadata: {
    suite: "smoke",
    owner: "qa",
  },
});

const browser = await chromium.connectOverCDP(session.connectUrl);

try {
  const context = browser.contexts()[0];
  const page = context.pages()[0];
  await page.goto("https://www.browserbase.com/", {
    waitUntil: "domcontentloaded",
  });

  const title = await page.title();
  assert.match(title, /Browserbase/i);

  const bodyText = await page.locator("body").textContent();
  assert.match(bodyText ?? "", /browser/i);
} finally {
  await browser.close();
}

Browserbase docs call out one easy-to-miss rule: after creating a session, you have five minutes to connect before it terminates. That affects queue-based CI. If your test runner creates all sessions up front and then waits behind a long build step, sessions can expire before Playwright attaches. Create the session as close as possible to the point of use, or use keep-alive when the workflow requires reconnection.

Session Configuration That Changes Test Outcomes

The Create Session API exposes practical settings: region, timeout, keepAlive, proxies, proxySettings, browserSettings, and userMetadata. The API reference lists available regions as us-west-2, us-east-1, eu-central-1, and ap-southeast-1. Timeout ranges from 60 seconds to 21600 seconds, which is 6 hours.

SettingWhy it mattersQA recommendation
regionLatency and geography can affect app behaviorPick the closest region unless testing location
timeoutLong flows need more than the project defaultSet explicit timeouts for suites over a few minutes
keepAliveLets sessions survive disconnects on eligible plansUse for human-in-the-loop and reconnect debugging
browserSettings.recordSessionReplay is enabled by default in docsKeep on for CI triage unless retention policy forbids it
browserSettings.logSessionLogs help explain browser failuresKeep on for non-sensitive test environments
browserSettings.viewportRemote browser size affects responsive UIPin width and height per suite
userMetadataLets you tag sessionsInclude suite, commit, shard, and test owner

Here is a session shape for a CI smoke shard that needs deterministic viewport and useful metadata:

import assert from "node:assert/strict";
import { Browserbase } from "@browserbasehq/sdk";

const apiKey = process.env.BROWSERBASE_API_KEY;
assert.ok(apiKey, "BROWSERBASE_API_KEY is required");

const bb = new Browserbase({ apiKey });

const session = await bb.sessions.create({
  region: "us-east-1",
  timeout: 900,
  browserSettings: {
    viewport: {
      width: 1440,
      height: 1000,
    },
    recordSession: true,
    logSession: true,
  },
  userMetadata: {
    suite: "checkout-smoke",
    shard: process.env.CI_NODE_INDEX ?? "0",
    commit: process.env.GITHUB_SHA ?? "local",
  },
});

assert.match(session.id, /^[a-zA-Z0-9_-]+$/);
console.log("Session URL: https://browserbase.com/sessions/" + session.id);

The userMetadata object is underrated. When a failure report contains only a Browserbase session ID, someone still has to map it back to a test. Metadata lets you filter and correlate sessions by suite, commit, shard, customer sandbox, or agent run ID.

Persist Logins With Contexts

Browserbase Contexts persist Chromium user data across sessions. The docs say they store cookies, localStorage, IndexedDB, session storage, service workers, web data, browser preferences, and site settings, while not including the browser HTTP cache. Context data is encrypted at rest. By default, every Browserbase session starts fresh, so Contexts are the way to reuse login state.

The workflow is create a Context, start a session with that Context, log in, close the session with persist: true, wait a few seconds for synchronization, then reuse the Context in later sessions.

import assert from "node:assert/strict";
import { chromium } from "playwright-core";
import { Browserbase } from "@browserbasehq/sdk";

const apiKey = process.env.BROWSERBASE_API_KEY;
assert.ok(apiKey, "BROWSERBASE_API_KEY is required");
assert.ok(process.env.APP_EMAIL, "APP_EMAIL is required");
assert.ok(process.env.APP_PASSWORD, "APP_PASSWORD is required");

const bb = new Browserbase({ apiKey });
const context = await bb.contexts.create({
  name: "qa-demo-login",
});

const session = await bb.sessions.create({
  browserSettings: {
    context: {
      id: context.id,
      persist: true,
    },
  },
});

const browser = await chromium.connectOverCDP(session.connectUrl);

try {
  const page = browser.contexts()[0].pages()[0];
  await page.goto("https://app.example.test/login");
  await page.getByLabel("Email").fill(process.env.APP_EMAIL);
  await page.getByLabel("Password").fill(process.env.APP_PASSWORD);
  await page.getByRole("button", { name: "Sign in" }).click();
  await page.getByRole("heading", { name: "Dashboard" }).waitFor();

  const heading = await page.getByRole("heading", { name: "Dashboard" }).textContent();
  assert.equal(heading, "Dashboard");
} finally {
  await browser.close();
}

console.log("Context ID: " + context.id);

What people get wrong is sharing one Context across parallel tests. Browserbase docs warn that simultaneous logins with the same Context can cause sites to force logout, and they recommend one Context per site and per login. For QA, that usually means one Context per test account class, not one global super-context. If you need ten parallel authenticated shards, create ten accounts and ten Contexts.

Stagehand: Use AI Where Selectors Are The Bottleneck

Stagehand is built by Browserbase for browser agents. The docs describe three AI primitives: act, extract, and observe, alongside Playwright-style page APIs such as goto, click, type, locator, and screenshot. Stagehand v3 can connect to Browserbase sessions, and the current docs emphasize that you can mix AI-powered actions with deterministic browser control.

That mix is the whole point for QA. Do not replace every stable locator with natural language. Use Stagehand when the page is third-party, changes often, or needs extraction from inconsistent structures. Use Playwright locators when you own the app and can add roles, labels, and test IDs.

Step typePrefer PlaywrightPrefer Stagehand
Owned app button with accessible nameYesRarely
Third-party portal with changing labelsSometimesYes
Extract product data from varied cardsMaybeYes with schema
Assert checkout success in your appYesNo
Explore available actions before scriptingNoYes, with observe
Enter credentialsYes, never send secrets to the modelUse observe only to find selectors

This pattern uses Stagehand to discover selectors, then fills credentials through deterministic page APIs so secrets do not enter the model prompt.

import assert from "node:assert/strict";
import { browserbase, Stagehand } from "@browserbasehq/stagehand";

const apiKey = process.env.BROWSERBASE_API_KEY;
assert.ok(apiKey, "BROWSERBASE_API_KEY is required");
assert.ok(process.env.APP_EMAIL, "APP_EMAIL is required");
assert.ok(process.env.APP_PASSWORD, "APP_PASSWORD is required");

// Stagehand v4 launches the Browserbase session itself, with its extension attached.
const browser = await browserbase.launch({ apiKey });

try {
  const stagehand = await Stagehand.create({ browser });
  const [page] = await browser.context.pages();

  await page.goto("https://app.example.test/login");
  const emailFields = await stagehand.observe("find the email input");
  const passwordFields = await stagehand.observe("find the password input");

  assert.ok(emailFields.data.length > 0, "email selector not found");
  assert.ok(passwordFields.data.length > 0, "password selector not found");

  await page.locator(emailFields.data[0].selector).fill(process.env.APP_EMAIL);
  await page.locator(passwordFields.data[0].selector).fill(process.env.APP_PASSWORD);
  await page.getByRole("button", { name: "Sign in" }).click();
  await page.getByRole("heading", { name: "Dashboard" }).waitFor();
} finally {
  await browser.close();
}

The failure mode here is easy to diagnose. If observe returns no candidates, the problem is element discovery. If candidates exist but login fails, the issue is credentials, app state, or post-submit behavior. Splitting AI discovery from deterministic action makes the run explainable.

Debugging With Live View, Inspector, And Replay

Browserbase sessions are not black boxes. The docs describe Session Inspector with live browser state, real-time network requests and responses, console output, performance metrics, resource usage, and replay. Live View can also be embedded into your application and supports human-in-the-loop control for authentication, CAPTCHAs, or unexpected errors.

The Playwright quickstart shows bb.sessions.debug(session.id) returning debug URLs. Use that in local repro scripts, but do not print sensitive URLs into public logs.

import assert from "node:assert/strict";
import { Browserbase } from "@browserbasehq/sdk";

const apiKey = process.env.BROWSERBASE_API_KEY;
assert.ok(apiKey, "BROWSERBASE_API_KEY is required");

const bb = new Browserbase({ apiKey });
const session = await bb.sessions.create({
  userMetadata: {
    purpose: "debug-demo",
  },
});

const debugUrls = await bb.sessions.debug(session.id);
assert.ok(debugUrls.debuggerUrl.startsWith("https://"));

console.log("Session: https://browserbase.com/sessions/" + session.id);
console.log("Debugger URL: " + debugUrls.debuggerUrl);

For automated failure reporting, prefer storing the Browserbase session page URL and replay metadata behind your own access controls. The replay docs say the metadata response lists each tab recording and the page response is an HLS .m3u8 playlist whose segment URLs are signed and valid for a limited time.

import assert from "node:assert/strict";
import { Browserbase } from "@browserbasehq/sdk";

const apiKey = process.env.BROWSERBASE_API_KEY;
const sessionId = process.env.BROWSERBASE_SESSION_ID;
assert.ok(apiKey, "BROWSERBASE_API_KEY is required");
assert.ok(sessionId, "BROWSERBASE_SESSION_ID is required");

const bb = new Browserbase({ apiKey });
const meta = await bb.sessions.replays.retrieve(sessionId);
assert.ok(meta.pages.length > 0, "session replay has no pages");

const firstPage = meta.pages[0];
const playlist = await bb.sessions.replays.retrievePage(sessionId, firstPage.pageId);
const m3u8 = await playlist.text();
const hasSegment = m3u8.split("\n").some((line) => line.startsWith("https://"));

assert.equal(hasSegment, true);

Do not call the replay API directly from a browser client with your API key. The docs explicitly recommend proxying through your own backend if embedding an HLS player. For QA dashboards, make the backend enforce the same access policy as your CI logs.

Proxies, Identity, And Regional Test Design

Browserbase supports built-in residential proxies, custom external proxies, geolocation settings, and ordered routing rules. The proxy docs say proxies: true makes a best-effort attempt to use a US-based proxy by default, while array configuration can specify Browserbase proxies, external proxies, geolocation, domain patterns, and exclusions.

Use proxies to test location-sensitive behavior, not as a magic fix for every block. The docs warn that provider restrictions can affect categories such as financial services, government domains, streaming, ticketing, webmail, and gambling, and that ERR_TUNNEL_CONNECTION_FAILED can mean an unsupported site, unsupported city, or temporary proxy problem.

Proxy patternExample useQA caution
proxies: trueQuick US-ish proxy coverageBest effort geography, not a location assertion
Browserbase geolocationCountry-level regional behaviorPrefer broad country before city
External proxyCorporate egress or approved providerValidate credentials at session creation
Domain routingDifferent proxy for selected domainsOrder matters, first matching rule wins
type: "none" exclusionKeep internal domains directPut exclusions before fallback proxies

Here is a focused regional session:

import assert from "node:assert/strict";
import { Browserbase } from "@browserbasehq/sdk";

const apiKey = process.env.BROWSERBASE_API_KEY;
assert.ok(apiKey, "BROWSERBASE_API_KEY is required");

const bb = new Browserbase({ apiKey });
const session = await bb.sessions.create({
  proxies: [
    {
      type: "browserbase",
      geolocation: {
        country: "GB",
      },
    },
  ],
  userMetadata: {
    suite: "regional-pricing",
    country: "GB",
  },
});

assert.ok(session.connectUrl.startsWith("wss://"));

For custom proxies with private certificate authorities, Browserbase supports uploading PEM certificates and referencing them through proxySettings.caCertificates. That is cleaner than disabling certificate validation globally when testing corporate egress flows.

CI Architecture For Browserbase Tests

Browserbase moves the browser out of your CI runner, but it does not remove test architecture. You still need sharding, timeouts, artifact capture, and secret hygiene. The difference is that the expensive browser process lives remotely and the failure evidence lives in Browserbase.

name: browserbase-smoke

on:
  pull_request:
  workflow_dispatch:

jobs:
  smoke:
    runs-on: ubuntu-24.04
    steps:
      - uses: actions/checkout@v7
      - uses: actions/setup-node@v7
        with:
          node-version: "22"
          cache: "npm"
      - run: npm ci
      - name: Run Browserbase Playwright smoke
        env:
          BROWSERBASE_API_KEY: ${{ secrets.BROWSERBASE_API_KEY }}
          GITHUB_SHA: ${{ github.sha }}
        run: node tests/browserbase-smoke.mjs
      - uses: actions/upload-artifact@v7
        if: always()
        with:
          name: browserbase-results-${{ github.run_id }}
          path: test-results

A production workflow should write a small JSON file per test with the session ID, app URL, commit SHA, shard, and failure reason. Upload that as an artifact and, if your organization allows it, add the Browserbase session URL to the test report.

import { mkdir, writeFile } from "node:fs/promises";
import assert from "node:assert/strict";

const sessionId = process.env.BROWSERBASE_SESSION_ID;
assert.ok(sessionId, "BROWSERBASE_SESSION_ID is required");

await mkdir("test-results", { recursive: true });
await writeFile(
  "test-results/browserbase-session.json",
  JSON.stringify(
    {
      sessionId,
      sessionUrl: "https://browserbase.com/sessions/" + sessionId,
      commit: process.env.GITHUB_SHA ?? "local",
      suite: "checkout-smoke",
    },
    null,
    2,
  ),
);

If you use Playwright Test, keep filtering normal and explicit. For example, npx playwright test --grep @browserbase selects tagged tests, while npx playwright test tests/checkout.spec.ts -g "guest checkout" narrows by title. Do not ask agents to invent runner flags. Playwright uses --grep or -g.

Cost, Rate Limits, And Plan Design

The Browserbase plans page currently lists Free, Developer, Startup, and Scale. Free includes 3 concurrent sessions, 1 browser hour, and 15 minute session duration. Developer is $20 with 25 concurrent browsers, 100 browser hours, 6 hour session duration, and proxy allocation. Startup is $99 with 100 concurrent browsers and 500 browser hours. Scale is custom with higher concurrency, enterprise controls, and compliance options.

The same docs say browser time is billed by the minute with a one-minute minimum per session, and session creation has per-minute limits. That has direct test-design consequences.

Cost leverBad patternBetter pattern
Session startupCreate a new session for every tiny assertionGroup related checks into one browser flow
ParallelismFan out beyond creation rateCap shards to plan limits
TimeoutsLeave defaults implicitSet suite-specific timeout values
ReplaysRecord all sensitive exploratory sessionsAlign recording with data policy
ContextsRe-login on every testPersist stable login state per account
ProxiesRoute all traffic through proxiesUse proxies only for tests that need them

Do not optimize cost by deleting assertions. Optimize by choosing the right evidence layer: Search or Fetch for content discovery, Browser Sessions for real interactions, Stagehand for resilient exploration, and Playwright for deterministic regression checks.

A Practical Failure Mode: Session Expired Before Connect

Here is a realistic Browserbase failure: CI creates 40 sessions at the start of a job, then the runner installs dependencies, builds the app, starts the server, and finally tries to connect. Some sessions fail with CDP connection errors. The root cause is not Playwright flake. Browserbase docs say a newly created session must be connected within five minutes before it terminates.

The diagnosis is to log creation time and connect time, then move session creation into the test fixture after the app is ready. A useful guard looks like this:

import assert from "node:assert/strict";
import { chromium } from "playwright-core";
import { Browserbase } from "@browserbasehq/sdk";

const apiKey = process.env.BROWSERBASE_API_KEY;
assert.ok(apiKey, "BROWSERBASE_API_KEY is required");

export async function withBrowserbasePage(run) {
  const bb = new Browserbase({ apiKey });
  const session = await bb.sessions.create({
    timeout: 1200,
    userMetadata: {
      fixture: "withBrowserbasePage",
    },
  });

  const browser = await chromium.connectOverCDP(session.connectUrl);
  try {
    const context = browser.contexts()[0];
    const page = context.pages()[0];
    await run({ page, session });
  } finally {
    await browser.close();
  }
}

Create late, connect immediately, and close deliberately. That one habit removes a surprising amount of remote-browser flake.

Frequently Asked Questions

Is Browserbase only for AI agents?

No. Browserbase is useful for AI agents, but QA teams can use it as remote browser infrastructure for Playwright, Puppeteer, and Selenium without adding AI to the test. The agent-specific value appears when you add Stagehand, Agents, or natural-language workflows. A conservative rollout is to move a small Playwright smoke suite to Browserbase first, prove replay and debugging value, then add Stagehand for brittle third-party workflows.

Should Stagehand replace Playwright locators?

Not across your owned application. Stable Playwright locators are still better for deterministic regression tests, especially roles, labels, and well-named test IDs. Stagehand is strongest when selectors are unknown, changing, or outside your control. A good hybrid pattern is using Stagehand observe to find candidates, then using Playwright-style APIs for secrets, clicks, assertions, and side-effect checks. That keeps AI helpful without making every test probabilistic.

How do Browserbase Contexts differ from Playwright storage state?

Playwright storage state usually serializes cookies and localStorage for a browser context. Browserbase Contexts persist the underlying Chromium user data directory across Browserbase sessions, including cookies, localStorage, IndexedDB, session storage, service workers, web data, preferences, and security state. They are useful for remote authenticated sessions, but they should not be shared recklessly across parallel tests. Use one Context per site and login identity when reliability matters.

What should I upload from CI failures?

Upload your normal test report plus a small Browserbase metadata file containing the session ID, Browserbase session URL, suite name, shard, commit, and failure message. Keep API keys out of artifacts. If replay is allowed by your data policy, link to the Browserbase session or route replay metadata through an internal dashboard. For sensitive flows, consider disabling recording or using enterprise retention controls instead of spreading recordings through generic CI artifacts.