Playwright Sharding in GitHub Actions: One Report

To shard Playwright tests in GitHub Actions, run npx playwright test --shard=${{ matrix.shardIndex }}/${{ matrix.shardTotal }} across a matrix of jobs, switch the reporter to blob on CI, upload each shard’s blob-report folder as an artifact, then run npx playwright merge-reports in a final job to get one HTML report. Set fullyParallel: true or the shards will be lopsided.

The docs give you a workflow that runs first time. Where it falls short is balance, setup cost and shared test data, which is what most of this post covers. Everything here was checked against Playwright 1.63, the current release as of October 2026.

What does —shard actually split?

--shard=2/4 tells Playwright to run the second quarter of the suite and ignore the rest. Each shard is a separate playwright test process on a separate machine. They do not talk to each other.

What gets divided depends on one setting. Without fullyParallel, Playwright shards by file: a whole spec file lands on one shard. The sharding guide is blunt about the consequence: if some files hold many more tests than others, “certain shards may end up running significantly more tests, while others may run fewer or even none.”

With fullyParallel: true, the split happens per test, and each shard gets an even count. That is the mode you want, with two caveats from the docs:

  • Tests skipped statically with test.skip() or test.fixme() are not counted when balancing, because they never run.
  • Without fullyParallel, the same file can land on different shards for different projects, so a Chromium and a WebKit project do not necessarily pair up.

Note that “even” means an even number of tests, not even wall-clock time. One shard can still draw the five slowest checkout tests. More on reading that below.

If a file uses test.describe.configure({ mode: 'serial' }), its tests stay together as a group. A suite full of serial files shards about as well as one with fullyParallel switched off.

The config change

Two settings matter in playwright.config.ts, fullyParallel and reporter:

import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  fullyParallel: true,
  reporter: process.env.CI ? 'blob' : 'html',
  retries: process.env.CI ? 2 : 0,
  use: {
    trace: 'on-first-retry',
  },
  projects: [
    { name: 'chromium', use: { ...devices['Desktop Chrome'] } },
  ],
});

The blob reporter writes everything about the run (results, traces, screenshots, attachments) into a zip in blob-report. According to the reporter docs, the file is named report-<hash>-<shard_number>.zip when sharding, so files from different shards never collide when you pour them into one folder. You can override the output path with the PLAYWRIGHT_BLOB_OUTPUT_FILE environment variable if your CI layout needs it.

Keep html locally. Nobody wants to unzip a blob to read a failure on their own laptop.

Playwright sharding in GitHub Actions: the workflow

This is the full workflow, adapted from the Playwright sharding guide with the two jobs in one file:

name: Playwright Tests
on:
  push:
    branches: [main]
  pull_request:
    branches: [main]
jobs:
  playwright-tests:
    runs-on: ubuntu-latest
    strategy:
      fail-fast: false
      matrix:
        shardIndex: [1, 2, 3, 4]
        shardTotal: [4]
    steps:
      - uses: actions/checkout@v6
      - uses: actions/setup-node@v6
        with:
          node-version: lts/*
      - run: npm ci
      - run: npx playwright install --with-deps chromium
      - run: npx playwright test --shard=${{ matrix.shardIndex }}/${{ matrix.shardTotal }}
      - name: Upload blob report
        if: ${{ !cancelled() }}
        uses: actions/upload-artifact@v4
        with:
          name: blob-report-${{ matrix.shardIndex }}
          path: blob-report
          retention-days: 1

  merge-reports:
    if: ${{ !cancelled() }}
    needs: [playwright-tests]
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v6
      - uses: actions/setup-node@v6
        with:
          node-version: lts/*
      - run: npm ci
      - name: Download blob reports
        uses: actions/download-artifact@v5
        with:
          path: all-blob-reports
          pattern: blob-report-*
          merge-multiple: true
      - name: Merge into HTML report
        run: npx playwright merge-reports --reporter html ./all-blob-reports
      - name: Upload HTML report
        uses: actions/upload-artifact@v4
        with:
          name: html-report--attempt-${{ github.run_attempt }}
          path: playwright-report
          retention-days: 14

Four details in that file decide whether you get a usable report on a failing run.

fail-fast: false

A GitHub matrix cancels its sibling jobs when one fails, by default. For a test suite that is the wrong behaviour: shard 2 fails, shards 1, 3 and 4 get cancelled, and you only see a quarter of your failures. Turn it off and let every shard finish.

if: ${{ !cancelled() }}

The test step exits non-zero when a test fails, and GitHub skips later steps after a failure. Without this condition the blob report is never uploaded on exactly the runs where you need it. The same condition on the merge job makes it run after failed shards, while still respecting a manual cancel.

One artifact name per shard

Artifacts made with actions/upload-artifact@v4 are immutable, as the upload-artifact README explains, so four jobs cannot write into one shared artifact. Each shard uploads blob-report-N, and the download step gathers them back with pattern: blob-report-* and merge-multiple: true. The versions above match Playwright’s own example; newer majors of both actions exist, so bump them when you next touch the file.

The merge job does not install browsers

merge-reports only needs @playwright/test from npm ci. Skip playwright install there; it costs time and does nothing.

On the test jobs, install only the browsers your projects use. --with-deps chromium is much lighter than the full set.

How many shards should you use?

Fewer than you think. Every shard repeats checkout, npm ci and the browser install before it runs a single test. Playwright’s CI guide also advises against caching browser binaries, because restoring the cache takes about as long as downloading them and the Linux system dependencies cannot be cached anyway. So that setup cost is fixed per shard, and past a certain count you are paying for setup, not tests.

Look at the merged report to find that point. The HTML report has a Speedboard tab (added in 1.57) that sorts tests by duration, and since 1.58 it shows a Timeline when you open a merged report. If one shard runs far longer than the rest, the fix is usually a few slow tests to split up, not more shards.

The CI guide also recommends workers: 1 on CI “to prioritize stability and reproducibility”, and suggests sharding as the way to get wider parallelism instead. That is a reasonable default. If your tests are well isolated, try two workers per shard before adding a fifth or sixth shard, since extra workers cost no extra setup.

Shared state breaks when you shard

Sharding exposes tests that secretly depend on each other. Two tests that edit the same account settings were fine when one worker ran them in order. Spread across four machines at once, they race.

Playwright 1.63 added test locks for this:

import { test } from '@playwright/test';

test('update user settings', { lock: 'user-settings' }, async ({ page }) => {
  // never runs at the same time as other tests holding 'user-settings'
});

The docs say locks work “across files, worker processes and projects”. They do not mention shards, and each shard is an independent process on its own runner with nothing coordinating it, so do not rely on a lock to stop shard 1 and shard 3 colliding. For anything that touches shared external state, give each shard its own data. A test account per shardIndex, passed in through an environment variable, is the simplest version.

What I would do

Turn on fullyParallel, switch to the blob reporter on CI, and start with the two-job workflow above at three or four shards. Read the merged report after a week of runs and adjust the count from what the Speedboard tells you. If the suite takes under five minutes on one runner, do not shard at all; a single job with a couple of workers is simpler and easier to debug. And check what the end-to-end suite is testing in the first place: component checks often belong in Vitest browser mode and pure logic in Node’s built-in test runner, where they run in seconds without a matrix.

Wiring test suites, preview deploys and release gates into a pipeline that stays fast is a regular part of our cloud hosting and deployment work. If your end-to-end suite has become the slowest step in your pipeline, that is a good place to start.

Need this built properly?

Whoooop Ltd has spent 15+ years building and maintaining web applications in TypeScript, React, Node.js and serverless — the same ground this post covers.

Get in touch