Every growing test suite eventually meets a wall. On one program I led, the wall was literal: CI jobs had a hard four-hour cap, and anything still running at the limit was stopped. Full regression across a portfolio of applications was landing between three and a half hours and just under four, depending on what was in scope. Several applications shared that single budget. One more feature, and releases would start shipping with an incomplete regression.
Here’s the approach that works, in the order I’d apply it.
1. Shard across containers
Playwright can split a suite into shards natively. Run each shard in its own container, in parallel, and a four-hour serial run becomes a fraction of that on wall-clock time.
# .github/workflows/regression.yml strategy: fail-fast: false matrix: shard: [1, 2, 3, 4, 5, 6, 7, 8] steps: - run: npx playwright test --shard=${{ matrix.shard }}/8 --reporter=blob - uses: actions/upload-artifact@v4 with: { name: blob-${{ matrix.shard }}, path: blob-report } # a final job runs: npx playwright merge-reports ./all-blobs
The merged report matters more than it looks. Eight separate reports mean nobody reads any of them. One report, with traces for every failure, is what a release manager actually uses.
2. Make the tests safe to run in parallel
Sharding exposes every hidden dependency between tests. Two tests that edit the same record, or a test that assumes another one ran first, will fail randomly under parallel load. The fixes are unglamorous and essential: each worker creates its own test data, tests never depend on execution order, and shared reference data is read-only.
3. Remove the waiting
Fixed sleeps are the silent budget killer. A single waitForTimeout(5000) in a shared helper, called several hundred times across a suite, can add up to an hour of machine time per run. We block fixed waits at commit time and replace them with web-first assertions that wait exactly as long as the page needs.
4. Balance shards by duration, not file count
An even split of files is rarely an even split of time. One shard gets the three slowest end-to-end journeys and finishes forty minutes after the rest. Use historical run times to rebalance, and split the slowest spec files so no single shard sets the pace.
5. Stop running everything, every time
The biggest win isn’t speed. It’s selection. We tier the suite:
- Smoke: minutes long, on every commit.
- Targeted: the tests mapped to what changed, on every pull request.
- Full regression: on release candidates and nightly.
Mapping tests to features and applications, through tags and traceability IDs, is what makes targeted runs possible. It’s also why we enforce tags and IDs at commit time rather than hoping people add them.
6. Quarantine flaky tests out loud
Retries hide flakiness and burn budget. A test that needs a retry to pass is flagged, tracked, and moved to a quarantine lane that doesn’t block the release, with an owner and a deadline. A quiet retry policy is how a suite loses its credibility one run at a time.
What you get
With sharding, parallel-safe data, and tiered selection in place, full regression fits comfortably inside the window again, with room to grow, and most changes get meaningful feedback in minutes rather than hours. The suite stops being the thing that decides when you can ship.
Takeaways
- Shard across containers and merge into one report with traces.
- Make every test independent before you parallelize.
- Ban fixed waits at commit time.
- Balance shards by historical duration.
- Tier the suite so full regression isn’t the only option.
- Quarantine flaky tests visibly, with owners.