Est.
FeaturesLong read

Running Headless Browser Tests in a CI Pipeline

Headless browsers eliminate graphics overhead to run tests faster and in parallel on CI servers.

Editor-at-Large · · 10 min read
Cover illustration for “Running Headless Browser Tests in a CI Pipeline”
Features · October 5, 2026 · 10 min read · 2,272 words

A continuous integration server has no monitor plugged into it, no graphics card doing any real work, and no window manager waiting to draw a browser tab. Yet the tests still need to run. The browser has to work without ever producing a picture for anyone to look at. That's the whole premise of headless mode, because it isn't a setting teams flip on to save a little time; it's the only way browser tests execute at all on a build server. A headless browser is not a stripped-down or simulated version of the real thing. It runs the exact same rendering engine as the browser on a developer's laptop: full HTML parsing, CSS layout, JavaScript execution, cookies, local storage, session handling, all of it happens the same way. Most teams split the work: headless mode runs in CI, where speed and scale take priority over letting a human watch the page load, and headed mode gets reserved for a developer's own machine when a test is failing and someone needs to see what the browser is actually doing. Because of this split, a test suite can run on every single pull request instead of being saved for an overnight batch job, so a team can trust its own code much faster.

The resource and speed case for headless execution in CI

The case for headless mode in CI isn't about shaving a few seconds off a build. It changes what's actually possible to run on the hardware most teams already have. A headless browser instance uses far less memory than a headed one, because it never has to allocate the buffers, compositing layers, and window-management overhead that come with displaying anything on a screen. That difference matters most at scale: a CI runner that could only support a handful of headed browser instances at once can typically support far more headless ones, and that gap is what makes real parallelization possible. The real payoff is parallelism. When you run many headless browser instances side by side on the same hardware, you shorten the feedback loop on every pull request, and a test suite that might take an hour finishes in a few minutes. For a team that pushes code multiple times a day, tests either run constantly as a safety net or get treated as a chore nobody wants to wait on.

Choosing between Playwright, Puppeteer, Selenium, and Cypress for headless CI work

Once a team accepts that headless is the only sane way to run browser tests in CI, the next decision is which tool actually drives the browser. There's no universal right answer here. The choice depends on what language the team already writes in, what kind of application is being tested, whether cross-browser coverage matters, and how much tolerance there is for managing drivers and version mismatches by hand.

Playwright, built by Microsoft and released under the Apache 2.0 license, supports headless execution natively across Chromium, Firefox, and WebKit through a single API, so one test file can run against three different rendering engines without rewriting anything. One of its more useful structural features is auto-waiting: Playwright waits for an element to actually be ready for interaction before acting on it, which is built into the tool's behavior rather than something a developer configures test by test. The tool leans JavaScript and TypeScript first. Bindings exist for other languages, but the center of gravity of the ecosystem, the documentation, the community answers, the GitHub issues, is in JS. A 2026 review from the DEV Community found Playwright to be one of the more dependable choices for modern web applications, citing headless speed gains as high as 15x over headed test runs, crediting the tool's balance of speed, stability, and cross-browser reach, and naming its JavaScript-first design as the main constraint for teams working in other languages.

Puppeteer, also from Google and also Apache 2.0, is a Node.js library you build specifically around Chromium. It talks to the browser through the Chrome DevTools Protocol directly, so you get a tight, efficient connection for Chromium-based work, including headless automation and web scraping. A 2026 comparison of automation tools from Firecrawl frames Puppeteer as a tool optimized for headless efficiency, with direct CDP access making it especially well suited to Chromium-primary workflows and Firefox as a secondary option.

Selenium is the oldest name in the group and still has the broadest language support, with bindings for Java, Python, JavaScript, C#, and several others, backed by the longest community history of any tool in this list. Setting up headless Chrome in Selenium requires passing --headless=new as an argument to ChromeOptions. The older convenience method, setHeadless(true), was removed in Selenium 4.10.0, so any guide or snippet referencing it is out of date. Selenium Manager, bundled into modern versions of the framework, automates the download and version-matching of ChromeDriver and GeckoDriver, which cuts down on manual driver maintenance without eliminating it outright, particularly for teams running their own fleet of self-managed CI runners. Selenium tends to fit best in multi-language or legacy enterprise environments, where you can write tests in Java or C# alongside existing codebases, and that outweighs the per-action speed advantages of newer tools.

Cypress, available under an MIT license with a paid tier for additional features, is built as an all-in-one framework with a strong developer experience at its center, and it's primarily a JavaScript tool. It runs headlessly by default through its cypress run command, supporting Chrome-family browsers including Edge and Chrome for Testing, along with Firefox, and experimental WebKit support. Its engine coverage is narrower than Playwright's, but Firecrawl's comparison positions Cypress as a tool optimized specifically for developer experience and component-level testing, with an MIT-licensed core at the heart of the project.

No one tool wins outright here. If a team is already standardized on Java, it might reasonably choose Selenium despite its lower per-action speed, while a team running a Chromium-only internal tool might find Puppeteer's direct CDP access is exactly enough. Matching the tool to the situation is the actual job.

Choosing a headless browser engine: Chrome, Firefox, WebKit, and Lightpanda

The framework a team picks and the browser engine that framework drives are two separate decisions, and conflating them is an easy mistake for anyone new to browser automation. Playwright can drive all three major engines, Chromium, Firefox, and WebKit, from one API. Puppeteer is tied to Chromium. Selenium can reach any of the major engines, but each one needs its own driver configured separately. Picking a framework doesn't lock in a browser, and picking a browser doesn't lock in a framework; the two choices sit side by side.

Headless Chrome has shipped with --headless=new as its default mode since Chrome 112, and it uses the same rendering pipeline as the regular, visible version of Chrome, so results from a headless run reflect what a real user's browser would actually show. It's the most common engine choice in CI today, and it carries the largest body of documentation and community troubleshooting of any option on this list, which matters quite a bit when something breaks at two in the morning before a release.

For any team that cares about how a page actually renders in Safari, and macOS runners remain expensive and limited in availability on most CI platforms, this closes a real gap.

Lightpanda represents a different idea. Some Playwright features carry caveats when running over a CDP connection like this one, so it isn't a drop-in replacement in every case. The trade-offs: no screenshots, no PDF generation, partial support for complex single-page applications, no stealth or anti-bot capabilities, and the project has only recently reached version 1.0 after moving out of beta. Lightpanda suits CI jobs that need raw throughput on purely functional tests and have no use for screenshots or visual checks. It isn't a general substitute for Chromium, and treating it as one would mean losing capabilities most test suites still rely on somewhere.

Configuring headless Chrome with Selenium in a CI environment

Getting headless Chrome running correctly under Selenium in a CI environment comes down to a short list of flags that aren't obvious from the API surface alone, and getting them wrong produces the most familiar failure pattern in browser automation: a test suite that runs fine on a developer's laptop and crashes the moment it hits Docker.

Two flags are close to mandatory in any containerized or locked-down CI environment. The first, --no-sandbox, exists because Chrome's sandboxing depends on kernel features that many container setups restrict by default, and leaving this flag off tends to produce crashes with no clear error message pointing at the cause. So the fix makes Chrome use /tmp instead of shared memory, and that sidesteps the problem.

A third flag, --window-size, deals with layout. If you set a realistic window size as part of the CI configuration, those tests stay honest.

A minimal setup in Python might look like this:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
options.add_argument("--headless=new")
options.add_argument("--no-sandbox")
options.add_argument("--disable-dev-shm-usage")
options.add_argument("--window-size=1920,1080")

driver = webdriver.Chrome(options=options)

Self-managed runner fleets are the exception: anyone maintaining their own CI infrastructure still needs to check driver compatibility after a Chrome update ships, since automatic matching only works as well as the environment it's running in.

Firefox follows the same basic structure through GeckoDriver and FirefoxOptions. The specific flag names differ, but if you set headless mode, disable sandboxing where needed, and fix a sane window size, the same logic carries over directly.

Setting up Playwright for headless CI, including GitHub Actions and Docker

Playwright takes a different approach to setup than Selenium does, mostly because headless is already the default behavior, so you don't need a special flag to turn it on. The configuration decisions that actually matter in CI are about how browsers get installed, which container image the job runs in, and how much parallelism the pipeline is configured to use.

Playwright manages its own browser binaries rather than relying on the system's installed browsers. Running npx playwright install downloads the specific browser versions Playwright has tested against, and in CI, you either run this step as part of the pipeline or bake it directly into a Docker image ahead of time. Either way, you have no separate driver to install and no ChromeDriver version mismatch to chase down, because Playwright controls both sides of that relationship.

For teams running in Docker, the official image hosted at mcr.microsoft.com/playwright comes with the correct browser versions and all required system dependencies already bundled in. Using it as a base image removes an entire category of failure: the "missing shared library" error that appears when a minimal container image lacks some font package or rendering dependency Chrome assumes is present.

A typical GitHub Actions job follows a simple sequence: install Node dependencies, run npx playwright install --with-deps to pull down both the browsers and their system-level dependencies in one step, then run the test command itself. That --with-deps flag matters specifically in CI, because it handles the operating-system packages a developer's own machine already has installed, and that would otherwise go unnoticed until a build fails.

Playwright's built-in auto-waiting removes a large share of the timing-related flakiness that occurs in test suites written without explicit waits.

Playwright is free and open-source under Apache 2.0. You pay no per-seat fee, no per-test charge, and no subscription to run it in a private CI pipeline, and that matters if you're scaling a test suite across many parallel jobs where a per-test pricing model would get expensive fast.

The most common CI failures

A test suite that passes locally and breaks in CI almost always traces back to one of a small number of causes, and each one has a fix that doesn't require guesswork.

Browser version mismatch happens when the browser binary running in CI differs from the version installed on a developer's machine, and even small version differences can make rendering behave differently or expose API differences, so tests fail with no code change behind them. The fix is to pin browser versions explicitly, whether through Playwright's own install mechanism, a locked Docker image tag, or a lockfile, and to update that pinned version on purpose rather than letting the CI runner quietly inherit whatever version happens to be installed on it.

Missing system dependencies cause a different kind of failure. Headless Chrome depends on a set of shared libraries, fonts, NSS, libgbm among them, that a minimal Docker image often doesn't include by default. The fix is to build from the framework's official Docker image or to run its dependency-install script, playwright install-deps in Playwright's case, as a dedicated step in CI setup rather than assuming the base image has everything Chrome needs.

Viewport and layout failures trace back to headless Chrome's small default resolution, and that breaks any test checking for elements that only render above a certain screen width. Setting --window-size in Selenium or the viewport value in Playwright's configuration to match the application's actual target resolution resolves this directly.

Flaky selectors round out the list, and they come from a different source: tests written against CSS class names generated by a build tool, which change every time the build output changes, breaking tests that had nothing wrong with them. The fix is to target stable attributes instead, data-testid values or ARIA roles, which stay constant across builds regardless of how the underlying styling or markup shifts.

None of these failure modes require exotic debugging once they're recognized for what they are. Each one points back to a specific, fixable mismatch between what the test assumes and what the CI environment actually provides.

Sources

  1. Chrome Headless mode
  2. Continuous Integration

More in Features