Flaky Test Root Causes in Browser Automation Suites
Flaky tests hide deterministic bugs, not random noise.

Flaky Test Root Causes in Browser Automation Suites
Flaky tests: a structural problem with known causes
A flaky test passes and fails on the exact same code revision, with nothing changed in between. A flaky test that passes and fails on the exact same code revision, with nothing changed in between, means the failure has nothing to do with the code under test, and everything to do with something the test depends on that it shouldn't. Timing, environment state, external services, execution order, these are the actual variables at play, not the logic being tested. Once that framing lands, the entire premise of "the test is just flaky" starts to look like an evasion rather than an explanation.
Intermittency is the symptom, not the diagnosis. Every flaky test has a deterministic triggering condition, one that happens to not fire on every single run. That's not noise. That's a bug in the test's assumptions, waiting for its trigger.
This distinction changes what a team does next. Teams that treat flakiness as random respond with retries: rerun the failing test, if it passes the second time, ship it. Teams that treat it as deterministic respond with diagnosis, tracing the actual triggering condition back to one of a handful of known categories, and the structural issue that produces it is what they fix. One of these approaches makes the problem worse over time. The other makes it disappear.
And the scale of the problem argues for the second approach. Worse, this isn't a static background hum. The share of teams experiencing flakiness climbed from 10% in 2022 to 26% in 2025, a rise of roughly 160% in three years scrolltest.com. Something about the direction modern test suites are moving, more async operations, more parallel execution, more distributed services, is making the underlying conditions more common, not less. 59% of developers encounter flaky tests at least monthly, according to scrolltest.com's 2026 guide, which aggregates data from the StackOverflow Developer Survey StackOverflow Developer Survey, aggregated.
The real cost: how flakiness erodes CI trust faster than it burns compute
Compute is a recoverable cost. A team can throw more runners at reruns, add capacity, absorb the bill. What flakiness actually destroys is harder to buy back: the assumption that a red build means something is broken.
Microsoft's research on this found that developers who run into flaky tests become significantly less likely to investigate the next failure they see. That mechanism isn't a one-time event, it's a habit that compounds. Once an engineer has been burned two or three times by a failure that turned out to be nothing, they stop reading stack traces closely https://contextqa.com/blog/what-is-flaky-test-automation-fix/. They stop requiring a green build before merging. They start clicking rerun as a reflex, and the reflex outlives the specific test that taught it to them.
That reflex is re-run and hope. It feels harmless because most of the time it is. When a failure gets assumed to be noise, whatever real regression is hiding behind it goes uninvestigated, silently, indefinitely. One documented case: a payment bug sat behind a test that had been retried and auto-passed fourteen separate times before anyone looked closely enough to catch it. Fourteen chances to catch a payment defect, and fourteen times the team's tooling told them everything was fine.
The dollar figures this produces are not small. Atlassian's own numbers put flakiness-related investigation at over 150,000 developer hours a year across their engineering org, and in the Jira Frontend repository specifically, flaky tests accounted for up to 21% of master build failures testbooster.ai smartkeys.org. Slack's experience before it built automated flaky-test detection was worse still: the main branch held a 20% pass rate, and 57% of failing builds traced back to test job failures rather than actual developer mistakes, CI issues, or infrastructure problems testbooster.ai TestDino Flaky Test Benchmark Report 2026 Gartner. After Slack put automated detection in place, false failures dropped to under 4%, which says as much about how bad the baseline was as it does about the fix testbooster.ai TestDino Flaky Test Benchmark Report 2026 Gartner.
As teams zoom out further, the time sink becomes a structural feature of how QA teams spend their days Autonoma State of QA 2025 Diffie customer data, 2025–2026 Gartner aithinkerlab.com. Autonoma's State of QA 2025 research puts test maintenance, flakiness-fighting included, at roughly 40% of QA team time Autonoma State of QA 2025 Diffie customer data, 2025–2026 Gartner aithinkerlab.com. That's 40% not spent finding actual bugs Autonoma State of QA 2025 Diffie customer data, 2025–2026 Gartner aithinkerlab.com. That's 40% spent fighting the test infrastructure itself Autonoma State of QA 2025 Diffie customer data, 2025–2026 Gartner aithinkerlab.com.
So what, precisely, is generating all this? That's the question the rest of this piece works through, moving from the largest single cause down through the ones that hide inside it.
Async timing issues: the single largest root cause and why sleep statements make it worse
Start with the biggest number on the board. Academic research examining 201 bug fixes across 51 Apache open-source projects (Luo et al., FSE 2014) found that async timing issues accounted for 45% of flaky test fixes, making it by a wide margin the single largest category of root cause testbooster.ai TestDino Flaky Test Benchmark Report 2026 Gartner. A separate, more recent benchmark (TestDino, 2026) is at the same 45% figure for async waits, and adds that concurrency and race conditions account for another 20% on top of that testbooster.ai TestDino Flaky Test Benchmark Report 2026 Gartner. Combined, these two related categories explain something close to two-thirds of all flakiness industry-wide testbooster.ai TestDino Flaky Test Benchmark Report 2026 Gartner.
The mechanics are almost embarrassingly simple once you see them written out. A test clicks Submit, the API call fires, the UI is supposed to render a success banner, and the test immediately asserts that the banner is visible. On a fast machine, the API responds in milliseconds, the banner renders, the assertion passes. On a slower CI worker, though, the API takes a beat longer, and the assertion fires before the banner element even exists in the DOM. Same code. Same test. Different outcome, purely because of how long the machine happened to take that particular run.
The code isn't wrong in either case. The application worked correctly both times, the banner did eventually render. What was wrong was the test's assumption about how long "eventually" takes.
The instinctive fix, adding a sleep statement or a hardcoded wait, makes intuitive sense and is almost always the wrong move. It doesn't solve the synchronization problem, it just makes the timing window wider and the test slower. A two-second sleep might paper over the problem on a normal day and still fail on a day when the CI runner is under heavier load than usual https://contextqa.com/blog/what-is-flaky-test-automation-fix/. The underlying race between "did the assertion fire" and "did the banner render" is never actually resolved, just pushed back a couple seconds.
What actually solves this is waiting for a condition, not a duration. Playwright's auto-wait mechanism is the clearest example in modern tooling: instead of waiting a fixed amount of time, it waits until an element is genuinely present, visible, and interactive before letting a test act on it pie.inc Deloitte. Teams that migrate from Selenium to Playwright report up to 50% fewer flaky tests, largely credited to this one architectural difference pie.inc Deloitte. The pattern generalizes past that specific tool: wait for the network to go idle, wait for the data to load, wait for the element to be actionable, never wait for a clock.
There's a further wrinkle here that complicates the picture in a useful way. A 2025 FSE study found that 46.5% of flaky tests are what it calls Resource-Affected Flaky Tests, meaning their failure rate is directly shaped by how much CPU or memory was available at the moment they ran testbooster.ai FSE 2025 Gartner. This overlaps heavily with the timing category, because a resource-starved CI runner slows down async operations enough to push them past whatever threshold the test silently assumed testbooster.ai FSE 2025 Gartner. Which raises a practical question for anyone debugging this: if a test passes reliably on a developer's fast local machine but flakes specifically on a shared CI runner with variable CPU allocation, is that a code defect at all? Often, it isn't. It's a timing assumption meeting a resource constraint it was never designed to survive.
Brittle selectors and DOM drift: real but overstated, and why its share is often misread
Most engineers, asked to guess the top cause of flaky tests, will say selectors. That instinct is understandable, and it's also measurably wrong in scale. Analysis of real production test-suite failures from QA Wolf (January 2026) found DOM changes and brittle selectors account for only around 28% of failures, and more than 70% trace back to timing, test data problems, runtime errors, and rendering issues instead QA Wolf, January 2026 diffie.ai.
Why does perception run so far ahead of the actual numbers here? Visibility is the likely answer. A broken locator throws an obvious, legible error: "element not found," with a clean stack trace pointing right at the missing selector. An async timing failure, by contrast, often shows up as "element not visible," a message that looks nearly identical on the surface but points to a completely different root cause. Engineers see the selector-shaped error message far more often than they correctly diagnose the timing problem that actually causes it.
The failure mode itself is straightforward enough. Tests anchored to CSS classes, XPath expressions, or nth-child positional selectors break the moment the UI gets refactored, even when the application's actual behavior hasn't changed at all. A UI team swaps out a component library, and every test relying on a generated class name breaks in one afternoon. A developer reorganizes a form's layout, and XPath selectors built on DOM hierarchy positions snap. None of this reflects a real defect. It reflects a test that was, structurally, built to break the first time anyone touched the markup.
The fix is to anchor tests to meaning instead of structure: ARIA roles, data-testid attributes, accessible labels. These survive visual refactors and structural reshuffling because they're tied to what an element is for, not where it happens to sit in the DOM tree at a given moment.
DOM drift deserves its own mention here, distinct from selector fragility even though the two get lumped together constantly testbooster.ai TestDino Flaky Test Benchmark Report 2026 Gartner. Modern applications generate the DOM dynamically, so the element a test expects might not exist yet, might be off-screen, or might be tucked inside a shadow DOM. That's really a timing problem wearing a selector problem's clothes, and it's a large part of why the two categories get conflated so often in postmortems testbooster.ai TestDino Flaky Test Benchmark Report 2026 Gartner.
The 28% figure is a useful baseline, not a hard ceiling QA Wolf, January 2026 diffie.ai. Modern end-to-end suites built on component libraries with heavily generated class names, or products going through active design-system churn, likely see selector flake well above that academic number. The exact figure isn't knowable in general, but the direction is clear enough.
Selectors break because someone changed the UI. The next category is stranger, because it breaks when nobody changed anything at all.
Shared state and test order coupling: failures that only appear when tests run together
A test can be entirely correct on its own, yet fail intermittently when run in a suite because another test that ran first left something behind that it never expected to find.
The vectors here are as follows. Test A creates a user record; Test B expects to find exactly five users in the database and instead finds six, because Test A's leftover record is still sitting there. Test A authenticates and leaves a session cookie behind; Test B inherits that active session and silently skips the authentication steps it was actually written to exercise. Neither test did anything wrong on its own terms. They just weren't built to survive contact with each other.
Ordering makes this worse in a specific way. Modern test runners often randomize execution order or run tests in parallel for speed, which means a suite that passes cleanly in sequential order one day can flake the next simply because the order shifted. Nothing about the code changed. The schedule did.
There's a research finding here that reframes the whole category usefully: flaky tests tend to show up in clusters, with multiple tests sharing the same root cause simultaneously, which is evidence of shared infrastructure rather than isolated test bugs (Parry et al., April 2025). That clustering is itself evidence. It points toward shared infrastructure as the culprit, not a scattering of unrelated bugs across unrelated tests.
The structural fix is test independence, full stop: each test owns its own data lifecycle, creating what it needs, operating on it, and cleaning it up afterward, so that the outcome of one test never depends on whether another test happened to run first. Database transactions rolled back after each test, test-scoped fixtures, timestamp-based unique identifiers so parallel test runs don't collide over the same records. None of these are exotic techniques. They're discipline, applied consistently.
Marking an order-dependent test "known flaky" and quarantining it out of the suite feels like progress, but the coupling that caused the failure is still sitting there, untouched. Quarantine is triage. It buys time. It is not a fix, and treating it like one just means the underlying defect ships silently instead of loudly.
Environment and infrastructure inconsistency: when the test is right but the floor beneath it shifts
Here's the genuinely frustrating category, the one where the test code is correct, the application code is correct, and the failure still happens, because the ground both are standing on keeps shifting. No amount of rewriting test logic fixes this. The problem isn't in the code.
CI runners with variable CPU allocation are the clearest example: the exact same test times out on a resource-constrained runner and passes cleanly on a full-spec one, which loops directly back to that FSE 2025 finding that 46.5% of flaky tests are shaped by resource availability at execution time testbooster.ai Gartner.
Mobile testing amplifies all of this considerably. Device fragmentation, unexpected OS dialogs, permission prompts, a keyboard that covers the exact button a test needs to tap, animation timing, Appium locator drift across different OS versions, the surface area for environmental flakiness on mobile dwarfs what web testing deals with. A checkout test can fail not because checkout is broken, but because the on-screen keyboard happened to cover the submit button on one specific device configuration. The application worked. The test environment just wasn't accounted for.
Fixing this category means treating test infrastructure with the same seriousness as production infrastructure, which most organizations don't do by default. Pin CI runner specs and reserve dedicated capacity instead of pulling from a shared, variable-allocation pool. Mock or stub third-party APIs at the network layer so their latency never becomes the test's problem to begin with. Standardize timezone and locale settings across every CI environment variable so date logic behaves the same way everywhere.
The compute cost of ignoring all this is not trivial at scale. Google's own internal data put roughly 16% of its total testing compute toward running and rerunning tests affected by flakiness, and a meaningful chunk of that is environmental retries specifically, not engineers actually chasing down code defects diffie.ai. That's 16% of a company the size of Google's testing infrastructure, spent on a problem that pinned runner specs and better mocking would substantially reduce diffie.ai. DNS flaps or third-party API latency can surface as test failures with no code defect. Timezone or locale differences between local and CI environments can change date-based assertions.
Race conditions and concurrency: the root cause hiding inside the other four
Timing issues and race conditions get talked about as if they're the same thing, and they're related closely enough that the confusion is understandable, but the distinction matters for diagnosis. Timing issues are about waiting long enough for one operation to finish. Race conditions are about two or more operations contending for the same resource, or a test assuming an ordering between them that was never actually guaranteed.
Picture two parallel tests writing to the same database record at the same moment: one write wins, the other test asserts against data that's now stale. Or a concurrent scenario where one thread rotates a shared session token while another thread is mid-request using the old one. Neither of these is a timing problem in the wait-longer sense. There's no longer wait that fixes a contested resource. The operations need to not contend.
The scale here is large enough to matter on its own: the TestDino 2026 benchmark attributes 20% of all flakiness to concurrency and race conditions, trailing only async wait issues as the largest category testbooster.ai TestDino Flaky Test Benchmark Report 2026. Microsoft's own measurement across 2.4 million test executions on Windows builds found a 4.6% overall flake rate, with end-to-end integration tests flaking approximately 5x more often than unit tests diffie.ai. That gap between E2E and unit tests isn't a coincidence. E2E tests touch more services, more I/O, more shared resources, which means more surface area for concurrency issues diffie.ai.
Why exactly does this happen so unpredictably? The triggering condition for a race condition is a specific interleaving of operations, one that depends on scheduler timing at the operating-system or runtime level. That interleaving might never occur on a developer's quiet local machine, running one test at a time, and only show up on a CI runner under load, executing dozens of tests in parallel testbooster.ai TestDino Flaky Test Benchmark Report 2026 Gartner. Which is precisely why race conditions earn a reputation as nearly impossible to reproduce on demand: the conditions that trigger them are, by nature, conditions of contention, and a quiet machine has nothing to contend over.
The fixes track the fixes for shared state pretty closely, because the categories are cousins. Isolate parallel test workers at the data layer, giving each one its own schema, database, or tenant rather than a shared one. Avoid mutable singletons in test setup code that multiple workers might touch simultaneously. Use atomic operations and proper transactions anywhere a test genuinely must touch a shared resource.
Diagnosing the root cause of a given failure
An "element not found" or "element not visible" error should point first toward timing or DOM drift, and only second toward selector fragility. A test that passes reliably on a local machine but fails specifically and only in CI is a strong signal to look at environment or resource constraints before touching the test logic at all. A test that passes clean in isolation but fails once it's run inside the full suite points toward shared state or an order dependency, not a bug in that individual test. A test that fails more often specifically under parallel execution is pointing at a race condition. And a test that fails only on certain CI runner configurations, certain instance types, certain resource tiers, is very likely the resource-affected timing category the FSE 2025 research identified.
None of these signatures are proof on their own. But taken together, they turn "flaky test" from a shrug into a diagnosis with a short list of actual suspects, each with a known mechanism and a known fix. That's the entire premise this piece has been building toward: flakiness isn't noise to be tolerated or retried away, it's a symptom with a finite number of causes, and naming the right one is most of the work of making it stop. SOURCE PAGES (what the pages behind the outline's links say).
Sources
- The Flaky Test Problem: Root Cause and How AI Solves It for Good
- What Is a Flaky Test? Why Automated Tests Fail Randomly
- Why Tests Flake and How to Fix Them for Good
- Flaky Tests Are Killing Your Pipeline: The Complete Guide to Detection, Quarantine, and Prevention in 2026
- The Flaky Test Report 2026 | Diffie
- Flakiness in Automated Testing: Reasons and How to Reduce It
- pie.inc


