Your regression suite grows every sprint. Nobody deletes tests, because deleting a test feels like removing a safety net. Eventually the suite takes four hours, the team stops trusting it, and someone starts merging on a yellow build.
The problem is rarely the tests. It is the absence of a rule for deciding which ones matter for a given change.
That rule has an evidence base, and it is more striking than most teams expect. Google measured their own continuous integration at scale, and published the numbers.
Source: Memon et al., *Taming Google-Scale Continuous Testing*, ICSE-SEIP 2017.
The consequence for strategy is the useful part. If almost every run confirms what you already believed, the problem is not insufficient testing. It is that the cost of finding the failing slice is spread evenly across a suite where it is not evenly distributed.
This guide covers how to choose what to re-run, how to keep the suite honest as it grows, and what the decision costs you either way.
Regression testing and confirmation testing are not the same activity
These two get used interchangeably, and the conflation is expensive. The ISTQB Certified Tester Foundation Level syllabus v4.0.1, section 2.2.3, separates them precisely:
> Confirmation testing confirms that an original defect has been successfully > fixed.
> Regression testing confirms that no adverse consequences have been caused by a > change, including a fix that has already been confirmation tested.
Why the distinction decides your scope
Confirmation testing is scoped by the defect. You know the target: run the tests that failed, or add tests covering the fix. ISTQB allows a legitimate reduction here, noting that when time or money is short, confirmation testing can be restricted to exercising the steps that reproduced the failure and checking that it no longer occurs.
Regression testing is scoped by the blast radius, which is unknown by definition. It cannot be reduced the same way. ISTQB's guidance is to perform impact analysis first, to recognise the extent of the regression testing required.
The asymmetry that forces mechanical selection
The person who made a change is the worst-placed person to judge what else it touched. They know what they intended to affect. The regression question is what they affected without intending to.
That asymmetry is the argument for a selection rule that does not depend on the author's judgement, which is the rest of this article.
The question worth answering
Most teams ask "did anything break". That question has no bounded answer, so the honest response is to run everything, which is why suites bloat.
A better question: which parts of the system could this change plausibly affect, and how confident do we need to be before release?
That reframes regression from a ritual into a risk decision. The answer changes per change, which means your suite needs tiers rather than one monolithic run.
Where regression sits in the test process
Searchers ask for "the 7 stages of testing" constantly. The honest answer is that there is no canonical numbered list of seven stages. ISTQB defines a test process as a set of activities: planning, monitoring and control, analysis, design, implementation, execution, and completion. That is seven activities, which is likely where the question comes from.
Regression testing is not one of them. It is a reason for testing that recurs inside execution, and the syllabus is explicit that confirmation and regression testing are needed at all test levels whenever defects are fixed or changes are made.
Why treating it as a stage creates the four-hour suite
If regression is a stage, it becomes a phase near the end of a release. That is exactly how teams arrive at a two-week regression window. If it is a property of execution at every level, it distributes across the whole cycle.
Four selection strategies, and when each applies
These are not informal categories. ISTQB's Certified Tester syllabus names regression test selection and test case prioritisation as distinct techniques, and the distinction decides your tooling.
Retest all
Run everything. Correct before a major release, after a framework or language runtime upgrade, and whenever the change touches shared infrastructure. Slow and expensive, so reserve it rather than defaulting to it.
Regression test selection
Run only the subset reaching the changed code. This requires a mapping from tests to code, which you get from coverage instrumentation rather than from guesswork. The mapping decays: every refactor that moves code without moving tests silently breaks a link, which is why this technique fails quietly rather than loudly.
Test case prioritisation
Run everything, ordered so the highest-risk tests execute first. Scope is unchanged, so nothing is missed, and you get signal in minutes rather than hours. This is the safest of the four and the most underused, because it requires no dependency map.
Risk-based pruning
Permanently retire tests that no longer earn their place. Culturally the hardest, and the only one that stops the suite growing without limit.
Most teams use only the first. Prioritisation is where the fastest improvement sits, because it needs no new infrastructure.
What actually decides scope
Three inputs, in order of usefulness.
Change surface
Which modules did this commit touch, and what imports them? A change to a shared authentication library has a far wider blast radius than a copy edit on one screen. Compute this from the dependency graph, not from memory.
Google's finding gives this a measurable shape: modelling the codebase as a dependency graph, they found that test targets more than a distance of 10 dependency edges from the changed code hardly ever break. Distance to the change predicts failure better than any label a human applies.
Defect history
Areas that broke before break again. Pull closed defects by module for the last four quarters and rank. This is the most useful single input and the one most teams never extract, because it lives in the issue tracker rather than the codebase.
Business consequence
A broken checkout costs revenue within minutes. A broken admin export annoys one person on Friday. Equal coverage of both is how suites become unaffordable.
The tiering rule, applied without tooling
- Tier 1, always runs: revenue path, authentication, anything that has broken in production in the last two quarters. Budget ten minutes.
- Tier 2, runs on related change: everything the dependency graph links to the diff. Budget thirty minutes.
- Tier 3, runs nightly: the remainder, plus the full suite regardless of diff.
The ten-minute budget on tier 1 is the number that matters. Past roughly that point developers context-switch, and a suite people wait for is a suite people read.
A worked example: 2,400 tests, 52 minutes, 40 merges a day
Here is the decision made concrete. A team has a 2,400-test suite taking 52 minutes end to end. Developers merge around 40 times a day. Running everything on every merge costs 34.7 hours of compute per day and puts a 52-minute gate in front of every merge.
Step one: set the budget before choosing the tests
Pick the number first. Ten minutes, for the context-switching reason above. The budget then constrains the test count rather than the other way round.
At an average 1.3 seconds per test, ten minutes buys about 460 tests. That is the budget. Now the question is which 460.
Step two: select by proximity and by history
Most teams cannot compute dependency distance directly, so the practical proxy is the three-part rule: a test is tier 1 if it covers a revenue path, covers authentication, or covers anything that broke in production in the last two quarters.
That last clause does the work. It is empirical, specific to your system, and it shrinks on its own as defects stop recurring.
Step three: measure whether the selection is correct
The metric is escaped-defect rate per tier, not coverage.
If tier 1 never catches anything, the tests in it are not the right ones. A suite that only ever passes is telling you it is testing the wrong code, which is Google's 99:1 finding applied locally.
Automating the selection
Selection by hand does not survive contact with a real sprint. It needs to run in CI.
A practical starting point: tag every test with the modules it exercises, then have your pipeline read the diff and select by tag.
# select tests by the modules a commit touched
CHANGED=$(git diff --name-only origin/main...HEAD \
| awk -F/ '{print $2}' | sort -u | paste -sd, -)
pytest -m "module_$CHANGED or critical" --maxfail=3Two rules keep this honest. Always run the critical tier regardless of the diff, because dependency maps are never complete. And run the full suite nightly, so anything the selection missed surfaces within a day rather than at release.
The failure mode in that script
If CHANGED comes back empty, the marker expression selects only critical. That is the correct behaviour, and it is worth stating explicitly because the obvious alternative is worse: a selector that runs nothing when it finds nothing waves through the riskiest change class you have.
Can you automate regression testing, and should all of it be automated?
Yes, and regression is the strongest candidate for automation of any test type. ISTQB's reasoning is worth stating exactly: regression test suites are run many times, and the number of regression test cases generally increases with each iteration or release. The syllabus also advises that test automation should start early in the project.
That is the economic argument. A test that runs once has an automation payback period that never arrives. A test that runs 40 times a day pays for itself in days.
Three cases where automating is a net loss
- Code about to be rewritten. The automation cost lands, then the code changes and the test is rewritten too. Wait.
- Output needing human judgement. Layout, tone, whether a chart reads correctly. A pixel-diff assertion fails on every legitimate design change.
- Tests run once. One-off migration verification is cheaper by hand.
Manual or automated: the honest split
The question has a real answer: both, with a shifting ratio. Automated regression carries the repeatable path checks. Manual exploratory passes carry the cases nobody wrote a test for, which is where novel defects live. A suite that is 100% automated stops finding anything it was not already designed to find.
Seven regression variants, and when each earns its cost
Retest-all and selective are not competing options. They are the nightly run and the commit run, and a working strategy uses both.
Choosing a tool: what actually differentiates them
Tool comparisons usually list features, and the features are broadly equivalent across the serious options. What differs is maintenance cost under change.
Ask for per-test failure history specifically. Without it every selection decision is a guess, and the Google finding says most of your tests have never failed and never will.
Keeping the suite from rotting
Selection controls what you run today. It does nothing about a suite that grows without limit.
Retire on evidence, not opinion
A test that has passed every run for two years and covers code nobody has touched is documentation, not a safety net. Track pass and fail history per test, and review anything that has never failed.
Fix or delete flaky tests within a sprint
A test that fails randomly trains the team to ignore red builds. That habit costs more than the test protects. Martin Fowler puts the consequence plainly: left uncontrolled, non-deterministic tests can completely destroy the value of an automated regression suite.
Quarantine it immediately, then fix it or remove it. Fowler is careful about the limit of quarantine: it helps reduce the damage to other tests, but you still have to fix them soon. A quarantine folder nobody empties is just a slower silencer.
Configuring three retries is not a fix either. The non-determinism is still there and it will mask a genuine regression on the run where it matters.
Cap the runtime, not the count
Give the suite a time budget, for example fifteen minutes in CI. When it exceeds the budget, something must be parallelised, moved to nightly, or retired. A budget forces the conversation that test counts never do.
The escalation rule is the part that is always missing. When tier 1 drifts to twelve minutes the default is to accept it, then fifteen, and then the budget means nothing. Exceeding the budget should force a removal, not an extension.
Where the effort goes wrong
Automating the wrong layer
End-to-end tests are the slowest and most brittle way to catch most regressions. If a bug could be caught by a unit or integration test, catching it through the browser costs you time on every run forever.
Treating coverage percentage as the goal
Coverage tells you which lines executed, not whether the assertions were meaningful. A suite at 90% coverage with weak assertions catches less than one at 60% with sharp ones.
Running everything because selection feels risky
It is a real risk, and the mitigation is the nightly full run, not permanent over-testing.
What this costs
Be honest about the trade. Selective regression means accepting a small chance of missing a defect that the full suite would have caught.
The mitigation is layered: critical tier always runs, nightly full suite, production monitoring for what both miss. Together these cost far less than a four-hour pipeline that nobody waits for.
Handling the cases that resist selection
Some changes defeat diff-based selection entirely, and pretending otherwise is how teams get burned.
Configuration and environment changes
A change to a feature flag default or an environment variable touches no application code, so a diff-based rule selects nothing. Treat configuration files as high blast radius and map them to the critical tier explicitly. ISTQB makes the same point: regression testing may not be restricted to the test object itself but can also relate to the environment.
Dependency upgrades
A minor version bump on a shared library can alter behaviour across every module that imports it. Your own diff is empty, so change-based selection finds nothing to run. Run the full suite on any dependency change, no exceptions.
Data migrations
Schema changes affect anything that reads the altered tables, and that relationship rarely appears in the code diff. Tests that assert on shape pass, tests that assert on values fail, and the ones that matter are usually the ones nobody wrote. Pair every migration with the tests covering the affected data paths, and keep that mapping in the migration itself.
Shared component libraries
A change to a button component in a design system can affect every screen. Component-level visual tests catch this far more cheaply than end-to-end runs across every page.
The pattern across all four: when the change surface is genuinely unknowable from the diff, stop trying to be clever and run everything.
Regression testing when the system has no tests
Plenty of teams inherit a system with a regression suite that does not exist. The advice to prune and prioritise assumes something to prune.
Start from defects, not from code
Take the last two quarters of production incidents and write a test for each one. You get immediate protection against failures that have actually occurred, and a suite whose value nobody questions.
Cover the revenue path first
Whatever sequence produces money: signup, checkout, renewal. One reliable end-to-end test of that path is worth more than fifty unit tests of a utility function.
Write characterisation tests before refactoring
When you do not know what the correct behaviour is, capture the current behaviour. The test then tells you when a change altered something, which is the question that matters during a refactor.
Accept that early coverage will be uneven
A suite that covers the dangerous 20% properly is more useful than one that covers everything shallowly.
The strategy document: five things, and nothing else
A regression strategy that lives in someone's head is not a strategy. The written version needs:
- The tier definitions, with the budget in minutes for each
- The membership rule for tier 1, specific enough that two engineers would select the same tests
- The trigger table: which event runs which tier
- The escalation rule: what happens when a tier exceeds its budget
- The review cadence: when tier 1 membership gets re-derived from recent production defects
A sequence that works
- Extract defect counts by module for the last four quarters
- Tag your tests with the modules they exercise
- Define a critical tier that always runs, and keep it under ten minutes
- Wire diff-based selection into CI, with the critical tier always included
- Schedule the full suite nightly
- Review flaky and never-failing tests every sprint
- Set a runtime budget and hold to it
Steps 1 and 3 give you most of the benefit. The rest is refinement.
Getting help with the first pass
The analysis in step 1 is the part teams skip, because it needs someone to sit with the defect data rather than the test code. It is also the part that makes every later decision defensible.
If you want an outside read on your suite, QAble's automated regression testing service starts with exactly that analysis: what has broken, what is covered, and what the suite costs you per run.