Playwright is a functional testing tool, and the first thing to be clear about is that it is not a load testing tool. It drives a real browser, so it cannot generate thousands of concurrent users without a machine per few hundred of them. Anyone telling you to replace your load tool with Playwright is selling you an outage.
What Playwright does give you is the measurement a load tool cannot: real browser performance, as a real user experiences it, with JavaScript executing and the page actually rendering. A load tool tells you the server responded in 80ms. Playwright tells you the user saw content 2.4 seconds later because a render blocking script sat in the head.
This guide covers what to measure with Playwright, how to wire it into CI as a regression gate, and the point at which you need a different tool.
The two questions, and which tool answers each
Frontend performance asks: how fast is this page for one user. Load performance asks: how does the system behave with many users at once. They are different questions with different tools, and conflating them wastes weeks.
The practical split is that Playwright owns the browser side and a load tool owns the server side. Teams that run both catch a class of problem neither finds alone, because a page can be fast under no load and collapse under concurrency for reasons the load tool cannot see in the browser.
Measuring navigation and paint timing
The browser already collects the numbers you want. Playwright's job is to read them out, using the same Performance API a real page has access to.
The metrics worth capturing for most applications are first contentful paint, largest contentful paint, DOM content loaded, and the total load event. Cumulative layout shift matters too if your pages assemble progressively, though it needs an observer rather than a one-shot read.
Capture them as a set rather than individually. A regression in largest contentful paint alone might be a slow image; the same regression alongside a jump in DOM content loaded points at something blocking earlier in the pipeline.
Turning measurements into a gate
A number printed in a CI log is not a control. To catch regressions you need a threshold and a build that fails when it is crossed.
Set thresholds from your own measured baseline rather than from published targets. Run the measurement ten times against your current production build, take the median, and set the threshold somewhat above it. Published targets are useful as a direction of travel, but a gate calibrated to somebody else's application will either never fire or fire constantly.
Measure the median across runs, not a single run. Browser timing is noisy, and a single measurement will produce false failures often enough to get the gate disabled.
Pin the environment. Performance numbers from a shared CI runner under variable load are not comparable week to week. A dedicated runner, or at minimum consistent hardware, is what makes the trend line mean anything.
Parallelism, and why it distorts performance runs
Playwright runs tests in parallel by default, using several worker processes at once. Test files run in parallel while tests within a single file run in order in the same worker. Each worker is an independent OS process that starts its own browser.
That default is excellent for functional suites and wrong for performance measurement. Four browsers competing for CPU on one machine produce timings that reflect the contention, not the application.
Run performance tests with a single worker. From the command line that is --workers 1, or set it in the config for a dedicated performance project. Keep the functional suite parallel and the performance suite serial, as separate projects in the same configuration.
Throttling, and why unthrottled numbers flatter you
A test running on CI hardware over a fast network measures the best case your application will ever see. Your users are not on that connection.
Playwright can emulate slower conditions through the underlying browser protocol, constraining CPU and network so the measurement reflects a realistic device rather than a build agent. Applying a CPU slowdown factor and a bandwidth limit changes the numbers substantially, and it changes which problems surface: render blocking JavaScript that costs 80ms on a fast machine costs far more on a mid-range phone.
Pick the throttling profile from your own analytics rather than a default. If most of your traffic is desktop on broadband, throttling to a slow phone produces a gate that fails constantly and teaches people to ignore it. If most of your traffic is mobile, an unthrottled test is measuring an experience almost nobody has.
Whatever you choose, apply it consistently. A threshold calibrated on unthrottled runs means nothing once somebody enables throttling.
What to do with the results over time
Store every run. A performance gate tells you about today; a stored history tells you whether you have been slowly degrading for six months, which is the more common and more expensive pattern.
A simple approach that works: append the median metrics, the commit hash and the timestamp to a CSV or a small database in the CI job. You do not need a monitoring platform to spot a trend across fifty builds.
Review the trend monthly rather than watching every build. Individual build variation is noise. The useful signal is the shape of the line over weeks.
Measuring the right pages
A performance suite that covers every page costs more than it returns. Pick the pages by traffic and by revenue rather than by convenience.
Start with the entry points that carry the most sessions, because a slow landing page affects everyone who arrives. Then add the steps in your primary conversion flow, where slowness costs money directly. Then add any page known to be heavy, such as a dashboard assembling several data sources.
Five to ten pages is usually enough to detect a systemic regression. Beyond that you are adding runtime for diminishing signal, because most performance regressions come from shared assets and affect many pages at once.
Measure each page in the state a real user meets it. A dashboard measured with an empty account is not the dashboard your customers load, and the difference is frequently the entire performance problem. Seed a realistic amount of data before measuring, and keep that amount stable so the numbers stay comparable.
When to reach for the load tool
Move to a dedicated load tool the moment your question involves concurrency. How many simultaneous users can we serve, what happens at twice our peak, where does the system break: none of these can be answered by a browser automation tool, and attempting it with Playwright produces numbers that describe your test machine rather than your application.
The two tools complement each other cleanly. Run the load test to find the concurrency at which server response degrades, then run the Playwright measurement against the system while it is under that load. That combination tells you what the user experience actually is at peak, which neither tool reports alone.
Fitting it into the release process
A performance check has to run somewhere that people will act on the result, otherwise it is data collection rather than a control.
Running it on every pull request gives the fastest feedback and the highest noise. It works when your measurement is stable and your budget is generous enough that only genuine regressions trip it. It fails badly when the numbers wander, because the team learns to override the gate.
Running it nightly against the main branch is the pragmatic default for most teams. It catches a regression within a day, the environment is quieter, and there is no pressure to skip it in order to merge.
Running it before release only is the minimum that is still worth doing. It catches the problem late, when the cause is buried among many changes, but it is far better than discovering it from customers.
Whichever cadence you pick, make the result visible in the same place the team already looks. A performance dashboard nobody opens is the usual end state, and it is why attaching the numbers to the build output tends to outperform a separate tool.
Attributing a regression once you have one
A failed gate tells you the page got slower. It does not tell you why, and the next hour determines whether the finding is useful.
Start by separating server time from browser time. Time to first byte comes from the server; everything after it is the browser's work. If time to first byte moved and the paint metrics moved by the same amount, the cause is server side and the browser measurement is just reporting it faithfully.
If time to first byte is unchanged and paint moved, the cause is on the client. The usual candidates are a new render blocking script, a larger bundle, an added web font, or an image that is no longer correctly sized.
Capture a trace on failure rather than only numbers. Playwright can record a trace including network activity and screenshots, and comparing a trace from the failing build against one from the last good build usually identifies the new request that caused the change within minutes.
Bisect if the trace is ambiguous. Performance regressions frequently arrive with a dependency upgrade rather than with application code, which makes them invisible in a diff of your own source.
Related reading
For the wider performance picture, see our guides on performance testing and Playwright end to end testing.
If you want help separating the browser side from the load side in your own programme, our performance testing team can scope it with you.