End-to-end testing

Playwright: cross-platform end-to-end testing that holds up

How to run one Playwright suite on Windows, macOS and Linux across Chromium, Firefox and WebKit — setup, flaky-test traps, traces and CI sharding.

At a glance

ToolPlatformsLicenceBest for
Playwright Test
  • Windows
  • macOS
  • Linux
Apache-2.0 Web and Electron end-to-end tests in TypeScript, Python, .NET or Java, with one API for three browser engines.
Trace Viewer
  • Windows
  • macOS
  • Linux
Apache-2.0 Replays a failed test step by step: DOM snapshots, network, console and the exact action that broke.

A web application is not tested until it has been driven by a real browser, on the operating systems your users actually run. Playwright does exactly that from a single test suite: Chromium, Firefox and WebKit, on Windows, macOS and Linux, with the same API in TypeScript, Python, .NET and Java. This guide shows how to set it up so that the suite stays trustworthy as it grows.

Why Playwright for cross-platform work

  • Three engines, one API. WebKit is the engine behind Safari. Running it on Linux or Windows in CI catches most Safari-only layout and API bugs without a Mac in the loop — though it is not a substitute for a final check in real Safari.
  • Auto-waiting. Every action waits for the element to be attached, visible, stable and enabled. Most sleep() calls simply disappear.
  • Web-first assertions. expect(locator).toHaveText() retries until the condition holds or the timeout expires, instead of reading the page once and hoping.
  • Isolation by default. Each test gets a fresh browser context — cookies, storage and permissions included — in milliseconds.

Installing on each operating system

The Node.js flavour is the reference implementation and receives features first. Initialise a project, then download the browser builds Playwright is tested against:

npm init playwright@latest
npx playwright install --with-deps

--with-deps matters on Linux: it installs the system libraries the browsers need (fonts, codecs, GTK pieces) through the distribution's package manager. On Windows and macOS it is harmless. Browsers are cached per user — %USERPROFILE%\AppData\Local\ms-playwright on Windows, ~/Library/Caches/ms-playwright on macOS and ~/.cache/ms-playwright on Linux — which is the directory to cache in CI.

One configuration, three engines

Projects let the same tests run once per browser. Keep the operating-system dimension out of the config: that is the CI matrix's job, as described in our GitHub Actions matrix guide.

playwright.config.ts
import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  fullyParallel: true,
  forbidOnly: !!process.env.CI,
  retries: process.env.CI ? 2 : 0,
  reporter: [['html', { open: 'never' }], ['list']],
  use: {
    baseURL: process.env.BASE_URL,
    trace: 'on-first-retry',
    screenshot: 'only-on-failure',
  },
  projects: [
    { name: 'chromium', use: { ...devices['Desktop Chrome'] } },
    { name: 'firefox', use: { ...devices['Desktop Firefox'] } },
    { name: 'webkit', use: { ...devices['Desktop Safari'] } },
  ],
});

Note that baseURL comes from the environment: the same suite then runs against a local server, a preview deployment or staging without editing a file.

Locators that survive redesigns

Most flaky suites are not flaky — they are coupled to markup. Prefer locators that describe what a user perceives:

await page.getByRole('button', { name: 'Sign in' }).click();
await page.getByLabel('Email').fill('qa@example.test');
await expect(page.getByRole('heading', { level: 1 })).toHaveText('Dashboard');

Role and label locators double as a cheap accessibility check: if a button cannot be found by its accessible name, a screen-reader user cannot find it either. Keep data-testid for the rare element that has no meaningful role.

The cross-platform traps

SymptomUsual causeFix
Screenshot diffs only on one OSFont rendering and anti-aliasing differ per systemKeep one set of baselines per OS (the default naming does this) or generate them in Docker
Keyboard shortcut works on Windows, fails on macOSControl versus MetaUse ControlOrMeta, e.g. page.keyboard.press('ControlOrMeta+A')
File upload test fails on WindowsHard-coded / path separatorsBuild paths with path.join()
Test passes locally, times out in CIWaiting for a navigation armed after the clickAssert on the resulting state with await expect(page).toHaveURL(...) instead of waiting for an event
Dates differ between machinesTime zone and locale of the hostPin timezoneId and locale in use

Debugging failures you cannot reproduce

With trace: 'on-first-retry', every test that fails once records a trace on its retry. Download the HTML report artefact from CI and open the trace:

npx playwright show-trace test-results/checkout-webkit-retry1/trace.zip

The trace viewer shows a DOM snapshot before and after each action, the network log, console output and the source line. It answers in a minute the question "what was on screen when it failed on the macOS runner?" that would otherwise take an afternoon of guesswork.

Testing Electron applications

Playwright can launch an Electron application and drive its windows like pages. The support is labelled experimental, but it is the most practical way to test an Electron app on all three systems with the same tooling as the web front end:

import { _electron as electron, expect, test } from '@playwright/test';

test('main window opens', async () => {
  const app = await electron.launch({ args: ['.'] });
  const window = await app.firstWindow();
  await expect(window).toHaveTitle(/My App/);
  await app.close();
});

For native toolkits — Win32, WPF, Cocoa, GTK — Playwright cannot help; see desktop UI automation on all three systems instead.

Scaling the suite in CI

  • Shard large suites across machines with npx playwright test --shard=1/4, then merge the blob reports.
  • Cache the browsers, keyed on the Playwright version, to save a minute or more per job.
  • Upload the report as an artefact on failure, always — a red job without its trace is a rerun waiting to happen.
  • Retries are a symptom. Track tests marked flaky in the report and fix them; a retry that becomes routine hides real regressions.

Verdict

For web and Electron applications, Playwright is the strongest cross-platform choice available today: one suite, three engines, three operating systems, and debugging tools that make CI failures explainable. Pair it with an OS matrix in CI and you will hear about the Safari-only bug before your users do.

  • playwright
  • e2e
  • browser testing
  • webkit
  • electron
  • ci