View all services
Talk to QA Advisor
/Blog/Playwright Performance Testing: What It Measures, and What It Cannot
Performance Testing8 min read

Playwright Performance Testing: What It Measures, and What It Cannot

Playwright measures real browser performance but cannot generate load. Here is what to measure, how to gate on it in CI, and the point where you need a load tool instead.

Published September 30, 2026Last updated September 30, 2026
On this page

Playwright is a functional testing tool, and the first thing to be clear about is that it is not a load testing tool. It drives a real browser, so it cannot generate thousands of concurrent users without a machine per few hundred of them. Anyone telling you to replace your load tool with Playwright is selling you an outage.

What Playwright does give you is the measurement a load tool cannot: real browser performance, as a real user experiences it, with JavaScript executing and the page actually rendering. A load tool tells you the server responded in 80ms. Playwright tells you the user saw content 2.4 seconds later because a render blocking script sat in the head.

This guide covers what to measure with Playwright, how to wire it into CI as a regression gate, and the point at which you need a different tool.

The two questions, and which tool answers each

Frontend performance asks: how fast is this page for one user. Load performance asks: how does the system behave with many users at once. They are different questions with different tools, and conflating them wastes weeks.

Browser performance testing versus load testing
QuestionRight toolWhat it measuresWhat it cannot see
How fast is this page for one user?PlaywrightReal browser paint and navigation timing, with JavaScript executingBehaviour under concurrency
How many users can we serve at once?k6, JMeter, GatlingServer throughput, error rate and latency under concurrent loadWhat the user actually sees in a browser
Did this release make the page slower?Playwright in CIRegression against a measured baselineWhether the cause is server or client side, without further work
Where does the system break?Load toolThe concurrency at which response time or errors degradeRendering and client-side execution cost
What is the experience at peak load?Both togetherBrowser timing measured while the load tool drives concurrencyNothing: this is the combination worth running
Source: QAble

The practical split is that Playwright owns the browser side and a load tool owns the server side. Teams that run both catch a class of problem neither finds alone, because a page can be fast under no load and collapse under concurrency for reasons the load tool cannot see in the browser.

Measuring navigation and paint timing

The browser already collects the numbers you want. Playwright's job is to read them out, using the same Performance API a real page has access to.

Reading real browser timing in a Playwright test
Capture the metrics typescript
1import { test, expect } from '@playwright/test';
2
3async function metrics(page) {
4  return page.evaluate(() => {
5    const nav = performance.getEntriesByType('navigation')[0];
6    const paints = performance.getEntriesByType('paint');
7    const fcp = paints.find(p => p.name === 'first-contentful-paint');
8    const lcpEntries = performance
9      .getEntriesByType('largest-contentful-paint');
10
11    return {
12      ttfb: nav.responseStart - nav.requestStart,
13      domContentLoaded: nav.domContentLoadedEventEnd - nav.startTime,
14      load: nav.loadEventEnd - nav.startTime,
15      fcp: fcp ? fcp.startTime : null,
16      lcp: lcpEntries.length
17        ? lcpEntries[lcpEntries.length - 1].startTime
18        : null,
19    };
20  });
21}
Gate on the median, not one run typescript
1// Browser timing is noisy. A single run will fail the
2// build often enough that somebody disables the gate.
3const RUNS = 5;
4const BUDGET = { fcp: 1800, lcp: 2500, load: 4000 };
5
6test('home page stays within its performance budget',
7  async ({ page }) => {
8    const samples = [];
9    for (let i = 0; i < RUNS; i++) {
10      await page.goto('/', { waitUntil: 'load' });
11      samples.push(await metrics(page));
12    }
13
14    const median = (key) => {
15      const v = samples.map(s => s[key]).sort((a, b) => a - b);
16      return v[Math.floor(v.length / 2)];
17    };
18
19    for (const [key, budget] of Object.entries(BUDGET)) {
20      expect(median(key), `${key} median`).toBeLessThan(budget);
21    }
22  });

The metrics worth capturing for most applications are first contentful paint, largest contentful paint, DOM content loaded, and the total load event. Cumulative layout shift matters too if your pages assemble progressively, though it needs an observer rather than a one-shot read.

Capture them as a set rather than individually. A regression in largest contentful paint alone might be a slow image; the same regression alongside a jump in DOM content loaded points at something blocking earlier in the pipeline.

Turning measurements into a gate

A number printed in a CI log is not a control. To catch regressions you need a threshold and a build that fails when it is crossed.

Set thresholds from your own measured baseline rather than from published targets. Run the measurement ten times against your current production build, take the median, and set the threshold somewhat above it. Published targets are useful as a direction of travel, but a gate calibrated to somebody else's application will either never fire or fire constantly.

Measure the median across runs, not a single run. Browser timing is noisy, and a single measurement will produce false failures often enough to get the gate disabled.

Pin the environment. Performance numbers from a shared CI runner under variable load are not comparable week to week. A dedicated runner, or at minimum consistent hardware, is what makes the trend line mean anything.

Parallelism, and why it distorts performance runs

Playwright runs tests in parallel by default, using several worker processes at once. Test files run in parallel while tests within a single file run in order in the same worker. Each worker is an independent OS process that starts its own browser.

That default is excellent for functional suites and wrong for performance measurement. Four browsers competing for CPU on one machine produce timings that reflect the contention, not the application.

Run performance tests with a single worker. From the command line that is --workers 1, or set it in the config for a dedicated performance project. Keep the functional suite parallel and the performance suite serial, as separate projects in the same configuration.

Why Playwright performance numbers mislead
SymptomCauseFix
Timings vary wildly between runsTests running in parallel: workers compete for CPURun the performance project with a single worker
CI numbers far worse than localShared runner under variable loadUse a dedicated runner, or compare only against its own history
Numbers improve after the first testBrowser and HTTP cache warm between runsDecide deliberately whether you measure cold or warm, then be consistent
LCP reported as nullRead before the element settledWait for the load state, and read the last entry not the first
Gate never firesThreshold copied from a published target rather than measuredBaseline against your own production build, then set the threshold above the measured median
Everything looks fast, users complainTesting an unrepresentative page or a warm cache on a fast networkThrottle CPU and network to match your real user population
Source: QAble

Throttling, and why unthrottled numbers flatter you

A test running on CI hardware over a fast network measures the best case your application will ever see. Your users are not on that connection.

Playwright can emulate slower conditions through the underlying browser protocol, constraining CPU and network so the measurement reflects a realistic device rather than a build agent. Applying a CPU slowdown factor and a bandwidth limit changes the numbers substantially, and it changes which problems surface: render blocking JavaScript that costs 80ms on a fast machine costs far more on a mid-range phone.

Pick the throttling profile from your own analytics rather than a default. If most of your traffic is desktop on broadband, throttling to a slow phone produces a gate that fails constantly and teaches people to ignore it. If most of your traffic is mobile, an unthrottled test is measuring an experience almost nobody has.

Whatever you choose, apply it consistently. A threshold calibrated on unthrottled runs means nothing once somebody enables throttling.

What to do with the results over time

Store every run. A performance gate tells you about today; a stored history tells you whether you have been slowly degrading for six months, which is the more common and more expensive pattern.

A simple approach that works: append the median metrics, the commit hash and the timestamp to a CSV or a small database in the CI job. You do not need a monitoring platform to spot a trend across fifty builds.

Review the trend monthly rather than watching every build. Individual build variation is noise. The useful signal is the shape of the line over weeks.

🔬 From our work
The mistake we see most often is a performance suite running with the default parallel workers, which means four browsers are competing for CPU on one machine and the numbers describe the contention rather than the application. It usually surfaces as timings that vary by a factor of two between runs with no code change. We have no measured figure for how common this is, and we are not going to make one up: check your own worker count against your performance project, which takes a minute and answers it for your setup.

Measuring the right pages

A performance suite that covers every page costs more than it returns. Pick the pages by traffic and by revenue rather than by convenience.

Start with the entry points that carry the most sessions, because a slow landing page affects everyone who arrives. Then add the steps in your primary conversion flow, where slowness costs money directly. Then add any page known to be heavy, such as a dashboard assembling several data sources.

Five to ten pages is usually enough to detect a systemic regression. Beyond that you are adding runtime for diminishing signal, because most performance regressions come from shared assets and affect many pages at once.

Measure each page in the state a real user meets it. A dashboard measured with an empty account is not the dashboard your customers load, and the difference is frequently the entire performance problem. Seed a realistic amount of data before measuring, and keep that amount stable so the numbers stay comparable.

When to reach for the load tool

Move to a dedicated load tool the moment your question involves concurrency. How many simultaneous users can we serve, what happens at twice our peak, where does the system break: none of these can be answered by a browser automation tool, and attempting it with Playwright produces numbers that describe your test machine rather than your application.

The two tools complement each other cleanly. Run the load test to find the concurrency at which server response degrades, then run the Playwright measurement against the system while it is under that load. That combination tells you what the user experience actually is at peak, which neither tool reports alone.

Fitting it into the release process

A performance check has to run somewhere that people will act on the result, otherwise it is data collection rather than a control.

Running it on every pull request gives the fastest feedback and the highest noise. It works when your measurement is stable and your budget is generous enough that only genuine regressions trip it. It fails badly when the numbers wander, because the team learns to override the gate.

Running it nightly against the main branch is the pragmatic default for most teams. It catches a regression within a day, the environment is quieter, and there is no pressure to skip it in order to merge.

Running it before release only is the minimum that is still worth doing. It catches the problem late, when the cause is buried among many changes, but it is far better than discovering it from customers.

Whichever cadence you pick, make the result visible in the same place the team already looks. A performance dashboard nobody opens is the usual end state, and it is why attaching the numbers to the build output tends to outperform a separate tool.

Attributing a regression once you have one

A failed gate tells you the page got slower. It does not tell you why, and the next hour determines whether the finding is useful.

Start by separating server time from browser time. Time to first byte comes from the server; everything after it is the browser's work. If time to first byte moved and the paint metrics moved by the same amount, the cause is server side and the browser measurement is just reporting it faithfully.

If time to first byte is unchanged and paint moved, the cause is on the client. The usual candidates are a new render blocking script, a larger bundle, an added web font, or an image that is no longer correctly sized.

Capture a trace on failure rather than only numbers. Playwright can record a trace including network activity and screenshots, and comparing a trace from the failing build against one from the last good build usually identifies the new request that caused the change within minutes.

Bisect if the trace is ambiguous. Performance regressions frequently arrive with a dependency upgrade rather than with application code, which makes them invisible in a diff of your own source.

Related reading

For the wider performance picture, see our guides on performance testing and Playwright end to end testing.

If you want help separating the browser side from the load side in your own programme, our performance testing team can scope it with you.

Frequently Asked Questions

Can Playwright replace a load testing tool?

No. It drives real browsers, so generating meaningful concurrency needs roughly a machine per few hundred users. Use k6, JMeter or Gatling for load, and Playwright for what the browser actually experiences.

Why do our Playwright performance numbers vary so much?

Usually parallelism. Playwright runs workers in parallel by default and each starts its own browser, so four browsers compete for CPU on one machine. Run the performance project with a single worker.

How many runs should we average?

Take the median of at least five. Browser timing is noisy, and gating on a single measurement produces false failures often enough that somebody disables the gate.

Where should performance thresholds come from?

Your own measured baseline, not a published target. Run against your current production build, take the median, and set the threshold above it. A threshold calibrated to another application never fires or always fires.

Should we throttle CPU and network?

Yes, to a profile matching your real traffic. Unthrottled numbers on CI hardware measure the best case your application will ever see, and hide render blocking work that costs far more on a mid-range phone.

How often should the performance suite run?

Nightly against the main branch suits most teams: fast enough to catch a regression within a day, quiet enough to keep the numbers stable. Per pull request only works if your measurement is stable.

Which pages should we measure?

Five to ten, chosen by traffic and revenue. Entry points carrying the most sessions, the steps of your primary conversion flow, and any page known to be heavy. Beyond that you add runtime for diminishing signal.

How do we find the cause of a regression?

Separate server time from browser time first. If time to first byte moved, the cause is server side. If it did not and paint moved, look for a new render blocking script, a larger bundle or an unsized image.

Free Assessment

Get a free QA audit for your project

Identify quality gaps before they become production bugs.

Get Free Audit

Ship software with confidence

Talk to a QA advisor and find out how QAble can help your team build quality in at every stage.

No sales pitch
Technical walkthrough
No lock-in commitment

Talk to QA Advisor

Direct access to QAble's QA specialists.

Response within 24 hours