At a glance
| Tool | Platforms | Licence | Best for |
|---|---|---|---|
| Playwright Test |
|
Apache-2.0 | Web and Electron end-to-end tests in TypeScript, Python, .NET or Java, with one API for three browser engines. |
| Trace Viewer |
|
Apache-2.0 | Replays a failed test step by step: DOM snapshots, network, console and the exact action that broke. |
A web application is not tested until it has been driven by a real browser, on the operating systems your users actually run. Playwright does exactly that from a single test suite: Chromium, Firefox and WebKit, on Windows, macOS and Linux, with the same API in TypeScript, Python, .NET and Java. This guide shows how to set it up so that the suite stays trustworthy as it grows.
Why Playwright for cross-platform work
- Three engines, one API. WebKit is the engine behind Safari. Running it on Linux or Windows in CI catches most Safari-only layout and API bugs without a Mac in the loop — though it is not a substitute for a final check in real Safari.
- Auto-waiting. Every action waits for the element to be attached, visible, stable and enabled. Most
sleep()calls simply disappear. - Web-first assertions.
expect(locator).toHaveText()retries until the condition holds or the timeout expires, instead of reading the page once and hoping. - Isolation by default. Each test gets a fresh browser context — cookies, storage and permissions included — in milliseconds.
Installing on each operating system
The Node.js flavour is the reference implementation and receives features first. Initialise a project, then download the browser builds Playwright is tested against:
npm init playwright@latest
npx playwright install --with-deps--with-deps matters on Linux: it installs the system libraries the browsers need (fonts, codecs, GTK pieces) through the distribution's package manager. On Windows and macOS it is harmless. Browsers are cached per user — %USERPROFILE%\AppData\Local\ms-playwright on Windows, ~/Library/Caches/ms-playwright on macOS and ~/.cache/ms-playwright on Linux — which is the directory to cache in CI.
One configuration, three engines
Projects let the same tests run once per browser. Keep the operating-system dimension out of the config: that is the CI matrix's job, as described in our GitHub Actions matrix guide.
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: './tests',
fullyParallel: true,
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 2 : 0,
reporter: [['html', { open: 'never' }], ['list']],
use: {
baseURL: process.env.BASE_URL,
trace: 'on-first-retry',
screenshot: 'only-on-failure',
},
projects: [
{ name: 'chromium', use: { ...devices['Desktop Chrome'] } },
{ name: 'firefox', use: { ...devices['Desktop Firefox'] } },
{ name: 'webkit', use: { ...devices['Desktop Safari'] } },
],
});Note that baseURL comes from the environment: the same suite then runs against a local server, a preview deployment or staging without editing a file.
Locators that survive redesigns
Most flaky suites are not flaky — they are coupled to markup. Prefer locators that describe what a user perceives:
await page.getByRole('button', { name: 'Sign in' }).click();
await page.getByLabel('Email').fill('qa@example.test');
await expect(page.getByRole('heading', { level: 1 })).toHaveText('Dashboard');Role and label locators double as a cheap accessibility check: if a button cannot be found by its accessible name, a screen-reader user cannot find it either. Keep data-testid for the rare element that has no meaningful role.
The cross-platform traps
| Symptom | Usual cause | Fix |
|---|---|---|
| Screenshot diffs only on one OS | Font rendering and anti-aliasing differ per system | Keep one set of baselines per OS (the default naming does this) or generate them in Docker |
| Keyboard shortcut works on Windows, fails on macOS | Control versus Meta | Use ControlOrMeta, e.g. page.keyboard.press('ControlOrMeta+A') |
| File upload test fails on Windows | Hard-coded / path separators | Build paths with path.join() |
| Test passes locally, times out in CI | Waiting for a navigation armed after the click | Assert on the resulting state with await expect(page).toHaveURL(...) instead of waiting for an event |
| Dates differ between machines | Time zone and locale of the host | Pin timezoneId and locale in use |
Debugging failures you cannot reproduce
With trace: 'on-first-retry', every test that fails once records a trace on its retry. Download the HTML report artefact from CI and open the trace:
npx playwright show-trace test-results/checkout-webkit-retry1/trace.zipThe trace viewer shows a DOM snapshot before and after each action, the network log, console output and the source line. It answers in a minute the question "what was on screen when it failed on the macOS runner?" that would otherwise take an afternoon of guesswork.
Testing Electron applications
Playwright can launch an Electron application and drive its windows like pages. The support is labelled experimental, but it is the most practical way to test an Electron app on all three systems with the same tooling as the web front end:
import { _electron as electron, expect, test } from '@playwright/test';
test('main window opens', async () => {
const app = await electron.launch({ args: ['.'] });
const window = await app.firstWindow();
await expect(window).toHaveTitle(/My App/);
await app.close();
});For native toolkits — Win32, WPF, Cocoa, GTK — Playwright cannot help; see desktop UI automation on all three systems instead.
Scaling the suite in CI
- Shard large suites across machines with
npx playwright test --shard=1/4, then merge the blob reports. - Cache the browsers, keyed on the Playwright version, to save a minute or more per job.
- Upload the report as an artefact on failure, always — a red job without its trace is a rerun waiting to happen.
- Retries are a symptom. Track tests marked flaky in the report and fix them; a retry that becomes routine hides real regressions.
Verdict
For web and Electron applications, Playwright is the strongest cross-platform choice available today: one suite, three engines, three operating systems, and debugging tools that make CI failures explainable. Pair it with an OS matrix in CI and you will hear about the Safari-only bug before your users do.