View all services
Talk to QA Advisor
/Blog/How to Choose What to Re-Run: A Regression Testing Strategy for 2026
QA Strategy8 min read

How to Choose What to Re-Run: A Regression Testing Strategy for 2026

Most regression strategies fail at one decision: which tests run on this commit and which wait until tonight. Google's own CI data shows 99 runs in 100 confirm what you already believed. Here is the selection rule, the tier budgets, and the four change types that defeat selection entirely.

Last updated October 5, 2026
On this page

Your regression suite grows every sprint. Nobody deletes tests, because deleting a test feels like removing a safety net. Eventually the suite takes four hours, the team stops trusting it, and someone starts merging on a yellow build.

The problem is rarely the tests. It is the absence of a rule for deciding which ones matter for a given change.

That rule has an evidence base, and it is more striking than most teams expect. Google measured their own continuous integration at scale, and published the numbers.

Ninety-nine runs in a hundred confirm what you already believed
Across 5.5 million affected tests in one period, only 63,000 ever failed. The rest never failed once. The ratio of passed to failed test targets per code change is 99:1, and just 1.23% of all test executions actually caught a breakage or a fix being introduced. The entire purpose of a regression cycle is to find that 1.23%. Source: Memon et al., Taming Google-Scale Continuous Testing, ICSE-SEIP 2017.

Source: Memon et al., *Taming Google-Scale Continuous Testing*, ICSE-SEIP 2017.

The consequence for strategy is the useful part. If almost every run confirms what you already believed, the problem is not insufficient testing. It is that the cost of finding the failing slice is spread evenly across a suite where it is not evenly distributed.

This guide covers how to choose what to re-run, how to keep the suite honest as it grows, and what the decision costs you either way.

Regression testing and confirmation testing are not the same activity

These two get used interchangeably, and the conflation is expensive. The ISTQB Certified Tester Foundation Level syllabus v4.0.1, section 2.2.3, separates them precisely:

> Confirmation testing confirms that an original defect has been successfully > fixed.

> Regression testing confirms that no adverse consequences have been caused by a > change, including a fix that has already been confirmation tested.

Why the distinction decides your scope

Confirmation testing is scoped by the defect. You know the target: run the tests that failed, or add tests covering the fix. ISTQB allows a legitimate reduction here, noting that when time or money is short, confirmation testing can be restricted to exercising the steps that reproduced the failure and checking that it no longer occurs.

Regression testing is scoped by the blast radius, which is unknown by definition. It cannot be reduced the same way. ISTQB's guidance is to perform impact analysis first, to recognise the extent of the regression testing required.

The asymmetry that forces mechanical selection

The person who made a change is the worst-placed person to judge what else it touched. They know what they intended to affect. The regression question is what they affected without intending to.

That asymmetry is the argument for a selection rule that does not depend on the author's judgement, which is the rest of this article.

The question worth answering

Most teams ask "did anything break". That question has no bounded answer, so the honest response is to run everything, which is why suites bloat.

A better question: which parts of the system could this change plausibly affect, and how confident do we need to be before release?

That reframes regression from a ritual into a risk decision. The answer changes per change, which means your suite needs tiers rather than one monolithic run.

Tier the suite by consequence, not by test type
TierWhat it containsWhen it runsTime budget
Tier 1Revenue path, authentication, anything broken in production in the last two quartersEvery push10 minutes
Tier 2Everything the dependency graph links to the diffOn related change30 minutes
Tier 3The remainder, plus the full suite regardless of diffNightlyUnbounded

Where regression sits in the test process

Searchers ask for "the 7 stages of testing" constantly. The honest answer is that there is no canonical numbered list of seven stages. ISTQB defines a test process as a set of activities: planning, monitoring and control, analysis, design, implementation, execution, and completion. That is seven activities, which is likely where the question comes from.

Regression testing is not one of them. It is a reason for testing that recurs inside execution, and the syllabus is explicit that confirmation and regression testing are needed at all test levels whenever defects are fixed or changes are made.

Why treating it as a stage creates the four-hour suite

If regression is a stage, it becomes a phase near the end of a release. That is exactly how teams arrive at a two-week regression window. If it is a property of execution at every level, it distributes across the whole cycle.

Test levelRegression triggerBudget
ComponentEvery commitUnder 10 minutes
IntegrationEvery merge to mainUnder 30 minutes
SystemNightly, and before releaseUnbounded, runs overnight
AcceptanceBefore release sign-offScoped to changed journeys

Four selection strategies, and when each applies

These are not informal categories. ISTQB's Certified Tester syllabus names regression test selection and test case prioritisation as distinct techniques, and the distinction decides your tooling.

Retest all

Run everything. Correct before a major release, after a framework or language runtime upgrade, and whenever the change touches shared infrastructure. Slow and expensive, so reserve it rather than defaulting to it.

Regression test selection

Run only the subset reaching the changed code. This requires a mapping from tests to code, which you get from coverage instrumentation rather than from guesswork. The mapping decays: every refactor that moves code without moving tests silently breaks a link, which is why this technique fails quietly rather than loudly.

Test case prioritisation

Run everything, ordered so the highest-risk tests execute first. Scope is unchanged, so nothing is missed, and you get signal in minutes rather than hours. This is the safest of the four and the most underused, because it requires no dependency map.

Risk-based pruning

Permanently retire tests that no longer earn their place. Culturally the hardest, and the only one that stops the suite growing without limit.

Most teams use only the first. Prioritisation is where the fastest improvement sits, because it needs no new infrastructure.

What actually decides scope

Three inputs, in order of usefulness.

Change surface

Which modules did this commit touch, and what imports them? A change to a shared authentication library has a far wider blast radius than a copy edit on one screen. Compute this from the dependency graph, not from memory.

Google's finding gives this a measurable shape: modelling the codebase as a dependency graph, they found that test targets more than a distance of 10 dependency edges from the changed code hardly ever break. Distance to the change predicts failure better than any label a human applies.

Distance predicts failure better than any label you apply
Google modelled their codebase as a dependency graph and found that test targets more than a distance of 10 dependency edges from the changed code hardly ever break. This is the finding that makes selection defensible: proximity to the diff is measurable, whereas a human judgement about which tests are 'related' is not. If you cannot compute dependency distance, the practical proxy is defect history, because a module that broke before is a module that sits close to code people keep changing. Source: Memon et al., Taming Google-Scale Continuous Testing, ICSE-SEIP 2017.

Defect history

Areas that broke before break again. Pull closed defects by module for the last four quarters and rank. This is the most useful single input and the one most teams never extract, because it lives in the issue tracker rather than the codebase.

Business consequence

A broken checkout costs revenue within minutes. A broken admin export annoys one person on Friday. Equal coverage of both is how suites become unaffordable.

The tiering rule, applied without tooling

  • Tier 1, always runs: revenue path, authentication, anything that has broken in production in the last two quarters. Budget ten minutes.
  • Tier 2, runs on related change: everything the dependency graph links to the diff. Budget thirty minutes.
  • Tier 3, runs nightly: the remainder, plus the full suite regardless of diff.

The ten-minute budget on tier 1 is the number that matters. Past roughly that point developers context-switch, and a suite people wait for is a suite people read.

A worked example: 2,400 tests, 52 minutes, 40 merges a day

Here is the decision made concrete. A team has a 2,400-test suite taking 52 minutes end to end. Developers merge around 40 times a day. Running everything on every merge costs 34.7 hours of compute per day and puts a 52-minute gate in front of every merge.

Step one: set the budget before choosing the tests

Pick the number first. Ten minutes, for the context-switching reason above. The budget then constrains the test count rather than the other way round.

At an average 1.3 seconds per test, ten minutes buys about 460 tests. That is the budget. Now the question is which 460.

Set the budget first, then choose the tests that fit it Set the budget first, then choose the tests that fit it Minutes 0 10 20 30 40 50 60 Tier 1, every commit: 10 10 Tier 1, every commit Tier 2, every merge: 30 30 Tier 2, every merge Tier 3, nightly full suite: 52 52 Tier 3, nightly full suite Source: Tier 3 shown at the worked example's full-suite runtime of 52 minutes

Step two: select by proximity and by history

Most teams cannot compute dependency distance directly, so the practical proxy is the three-part rule: a test is tier 1 if it covers a revenue path, covers authentication, or covers anything that broke in production in the last two quarters.

That last clause does the work. It is empirical, specific to your system, and it shrinks on its own as defects stop recurring.

Step three: measure whether the selection is correct

The metric is escaped-defect rate per tier, not coverage.

MetricWhat it tells youAction threshold
Tier 1 runtimeWhether developers will wait for itOver 10 min: cut tests
Defects caught in tier 1Whether selection is correctNear zero: selection is wrong
Defects caught only nightlyCost of the selection gambleRising: promote those tests
Defects escaping to productionWhether the model holds at allAny: trace back to a tier

If tier 1 never catches anything, the tests in it are not the right ones. A suite that only ever passes is telling you it is testing the wrong code, which is Google's 99:1 finding applied locally.

The four strategies differ in what they demand and how they fail
StrategyNeedsFails byMissed-defect risk
Retest allNothingTaking too long to be runNone from selection
Regression test selectionCoverage-derived test-to-code mapMap decays silently on refactorReal, needs nightly full run
Test case prioritisationA risk orderingNothing: scope is unchangedNone
Risk-based pruningPass and fail history per testCultural resistanceReal, and permanent

Automating the selection

Selection by hand does not survive contact with a real sprint. It needs to run in CI.

A practical starting point: tag every test with the modules it exercises, then have your pipeline read the diff and select by tag.

bash
# select tests by the modules a commit touched
CHANGED=$(git diff --name-only origin/main...HEAD \
  | awk -F/ '{print $2}' | sort -u | paste -sd, -)

pytest -m "module_$CHANGED or critical" --maxfail=3

Two rules keep this honest. Always run the critical tier regardless of the diff, because dependency maps are never complete. And run the full suite nightly, so anything the selection missed surfaces within a day rather than at release.

The failure mode in that script

If CHANGED comes back empty, the marker expression selects only critical. That is the correct behaviour, and it is worth stating explicitly because the obvious alternative is worse: a selector that runs nothing when it finds nothing waves through the riskiest change class you have.

Can you automate regression testing, and should all of it be automated?

Yes, and regression is the strongest candidate for automation of any test type. ISTQB's reasoning is worth stating exactly: regression test suites are run many times, and the number of regression test cases generally increases with each iteration or release. The syllabus also advises that test automation should start early in the project.

That is the economic argument. A test that runs once has an automation payback period that never arrives. A test that runs 40 times a day pays for itself in days.

Three cases where automating is a net loss

  • Code about to be rewritten. The automation cost lands, then the code changes and the test is rewritten too. Wait.
  • Output needing human judgement. Layout, tone, whether a chart reads correctly. A pixel-diff assertion fails on every legitimate design change.
  • Tests run once. One-off migration verification is cheaper by hand.

Manual or automated: the honest split

The question has a real answer: both, with a shifting ratio. Automated regression carries the repeatable path checks. Manual exploratory passes carry the cases nobody wrote a test for, which is where novel defects live. A suite that is 100% automated stops finding anything it was not already designed to find.

Seven regression variants, and when each earns its cost

TypeWhat it runsWhen it is the right choice
CorrectiveExisting tests, unchangedCode changed but specifications did not
Retest-allEvery test in the suiteBefore release, or after a dependency upgrade
SelectiveSubset chosen by impact analysisRoutine merges with a small, local diff
ProgressiveNew tests for changed specsThe specification itself changed
CompleteFull suite plus new testsMajor version, or several merged branches
PartialTests around integrated codeAfter integrating a feature branch
UnitOne component, dependencies stubbedFastest signal, belongs in tier 1

Retest-all and selective are not competing options. They are the nightly run and the commit run, and a working strategy uses both.

Choosing a tool: what actually differentiates them

Tool comparisons usually list features, and the features are broadly equivalent across the serious options. What differs is maintenance cost under change.

ConsiderationQuestion to askWhy it decides the outcome
Selector stabilityWhat happens when a developer renames a CSS class?The largest single source of suite maintenance
Parallel executionCan it shard without test interdependence?Determines whether a 10-minute budget is reachable
Flake handlingDoes it quarantine, or just retry?Retrying hides non-determinism rather than surfacing it
CI integrationCan it fail a build on a tier, not just the suite?Tiering is unenforceable without this
Reporting granularityPer-test historical failure rate?Selection without failure history is guesswork

Ask for per-test failure history specifically. Without it every selection decision is a guess, and the Google finding says most of your tests have never failed and never will.

Keeping the suite from rotting

Selection controls what you run today. It does nothing about a suite that grows without limit.

Retire on evidence, not opinion

A test that has passed every run for two years and covers code nobody has touched is documentation, not a safety net. Track pass and fail history per test, and review anything that has never failed.

Fix or delete flaky tests within a sprint

A test that fails randomly trains the team to ignore red builds. That habit costs more than the test protects. Martin Fowler puts the consequence plainly: left uncontrolled, non-deterministic tests can completely destroy the value of an automated regression suite.

Quarantine it immediately, then fix it or remove it. Fowler is careful about the limit of quarantine: it helps reduce the damage to other tests, but you still have to fix them soon. A quarantine folder nobody empties is just a slower silencer.

Configuring three retries is not a fix either. The non-determinism is still there and it will mask a genuine regression on the run where it matters.

Cap the runtime, not the count

Give the suite a time budget, for example fifteen minutes in CI. When it exceeds the budget, something must be parallelised, moved to nightly, or retired. A budget forces the conversation that test counts never do.

The escalation rule is the part that is always missing. When tier 1 drifts to twelve minutes the default is to accept it, then fifteen, and then the budget means nothing. Exceeding the budget should force a removal, not an extension.

Diff-based selection, and why the critical tier is never excluded
Naive: selects only what the diff touches bash
1CHANGED=$(git diff --name-only origin/main...HEAD \
2  | awk -F/ '{print $2}' | sort -u | paste -sd, -)
3
4pytest -m "module_$CHANGED"
Safe: critical tier always runs, full suite nightly bash
1CHANGED=$(git diff --name-only origin/main...HEAD \
2  | awk -F/ '{print $2}' | sort -u | paste -sd, -)
3
4# tier 1 runs regardless: dependency maps are never complete
5pytest -m "critical or module_$CHANGED" --maxfail=3
6
7# and nightly, on a schedule, with no selection at all
8# pytest --durations=25
The left version trusts the dependency map. Every refactor that moves code without moving tests breaks a link in that map silently.

Where the effort goes wrong

Automating the wrong layer

End-to-end tests are the slowest and most brittle way to catch most regressions. If a bug could be caught by a unit or integration test, catching it through the browser costs you time on every run forever.

Treating coverage percentage as the goal

Coverage tells you which lines executed, not whether the assertions were meaningful. A suite at 90% coverage with weak assertions catches less than one at 60% with sharp ones.

Running everything because selection feels risky

It is a real risk, and the mitigation is the nightly full run, not permanent over-testing.

What this costs

Be honest about the trade. Selective regression means accepting a small chance of missing a defect that the full suite would have caught.

The mitigation is layered: critical tier always runs, nightly full suite, production monitoring for what both miss. Together these cost far less than a four-hour pipeline that nobody waits for.

From our work
Across QAble's automation engagements the first useful action is almost always the same: extract closed defects by module for the last four quarters and let that ranking set the tier 1 contents. It usually takes a day and it changes the priority order of everything that follows. We have not measured the average across our client base, so this is a qualitative observation from our engineers rather than a statistic.

Handling the cases that resist selection

Some changes defeat diff-based selection entirely, and pretending otherwise is how teams get burned.

Configuration and environment changes

A change to a feature flag default or an environment variable touches no application code, so a diff-based rule selects nothing. Treat configuration files as high blast radius and map them to the critical tier explicitly. ISTQB makes the same point: regression testing may not be restricted to the test object itself but can also relate to the environment.

Dependency upgrades

A minor version bump on a shared library can alter behaviour across every module that imports it. Your own diff is empty, so change-based selection finds nothing to run. Run the full suite on any dependency change, no exceptions.

Data migrations

Schema changes affect anything that reads the altered tables, and that relationship rarely appears in the code diff. Tests that assert on shape pass, tests that assert on values fail, and the ones that matter are usually the ones nobody wrote. Pair every migration with the tests covering the affected data paths, and keep that mapping in the migration itself.

Shared component libraries

A change to a button component in a design system can affect every screen. Component-level visual tests catch this far more cheaply than end-to-end runs across every page.

The pattern across all four: when the change surface is genuinely unknowable from the diff, stop trying to be clever and run everything.

Regression testing when the system has no tests

Plenty of teams inherit a system with a regression suite that does not exist. The advice to prune and prioritise assumes something to prune.

Start from defects, not from code

Take the last two quarters of production incidents and write a test for each one. You get immediate protection against failures that have actually occurred, and a suite whose value nobody questions.

Cover the revenue path first

Whatever sequence produces money: signup, checkout, renewal. One reliable end-to-end test of that path is worth more than fifty unit tests of a utility function.

Write characterisation tests before refactoring

When you do not know what the correct behaviour is, capture the current behaviour. The test then tells you when a change altered something, which is the question that matters during a refactor.

Accept that early coverage will be uneven

A suite that covers the dangerous 20% properly is more useful than one that covers everything shallowly.

Common regression failures, and what each actually means
SymptomUsual causeFix
Suite passes, defect reaches productionCoverage gap in an area with defect historyWrite a test per production incident from the last two quarters
Same test fails intermittentlyShared state or timing, not a product defectQuarantine within the sprint, then fix or delete
Suite time grows every sprintNo retirement policySet a wall-clock budget and hold it
Selection misses a regressionDependency map stale after a refactorAlways run tier 1, and run the full suite nightly
Team merges on a red buildFlake rate high enough to train distrustTreat flake rate as a release blocker, not a nuisance

The strategy document: five things, and nothing else

A regression strategy that lives in someone's head is not a strategy. The written version needs:

  1. The tier definitions, with the budget in minutes for each
  2. The membership rule for tier 1, specific enough that two engineers would select the same tests
  3. The trigger table: which event runs which tier
  4. The escalation rule: what happens when a tier exceeds its budget
  5. The review cadence: when tier 1 membership gets re-derived from recent production defects

A sequence that works

  1. Extract defect counts by module for the last four quarters
  2. Tag your tests with the modules they exercise
  3. Define a critical tier that always runs, and keep it under ten minutes
  4. Wire diff-based selection into CI, with the critical tier always included
  5. Schedule the full suite nightly
  6. Review flaky and never-failing tests every sprint
  7. Set a runtime budget and hold to it

Steps 1 and 3 give you most of the benefit. The rest is refinement.

Getting help with the first pass

The analysis in step 1 is the part teams skip, because it needs someone to sit with the defect data rather than the test code. It is also the part that makes every later decision defensible.

If you want an outside read on your suite, QAble's automated regression testing service starts with exactly that analysis: what has broken, what is covered, and what the suite costs you per run.

Frequently Asked Questions

How long should a tier 1 regression run take?

Ten minutes or less. Beyond that developers stop waiting for it and start merging on hope, which removes the only benefit a fast tier has. Set the budget first, then choose the tests that fit inside it.

Which tests belong in tier 1?

The revenue path, authentication, and anything that broke in production in the last two quarters. That last criterion is the one teams skip, and it is the one grounded in evidence rather than opinion. It also shrinks on its own as defects stop recurring.

Can test selection be trusted to pick the right tests?

Only with a nightly full run behind it. Dependency maps decay quietly: every refactor that moves code without moving tests breaks a link, and nothing reports it. Selection plus a nightly safety net is safe; selection alone is not.

Which changes should trigger running everything?

Configuration, dependency, migration and shared-component changes. Their blast radius is not knowable from the diff, so a minor version bump can change every module that imports it while touching one line of your code.

Does a growing test suite mean better coverage?

No, and suite size appears on none of the measures worth tracking. Escaped defect rate, mean time to feedback and flake rate tell you whether the suite works; test count tells you only how long it takes.

Should regression testing be run manually or automated?

Both, with a shifting ratio. Automation carries the repeatable path checks, because ISTQB notes regression suites run many times and grow with each release, which is what makes the automation pay back. Manual exploratory passes carry the cases nobody wrote a test for, and a fully automated suite stops finding anything it was not already designed to find.

When does automating a regression test lose money?

Three cases: code about to be rewritten, output that needs human judgement such as layout or tone, and checks that will only ever run once. In each the automation cost lands before the payback period starts.

How do you tell a flaky test from a real regression?

By whether it reproduces on the same commit. If it does not, you have non-determinism, and configuring retries hides it rather than fixing it. Martin Fowler's position is worth holding to: quarantine limits the damage, but the test still has to be fixed soon or deleted.

What belongs in a written regression strategy?

Five things: the tier definitions with budgets in minutes, the membership rule for tier 1, the trigger table mapping events to tiers, the escalation rule for when a tier exceeds its budget, and the review cadence for re-deriving tier 1 from recent defects. The escalation rule is the one usually missing, and without it the budget erodes quietly.

Which metric shows the regression strategy is working?

Escaped defect rate per tier, not coverage percentage. If tier 1 never catches anything, the selection is wrong rather than the code being clean: Google's measurement found only 1.23% of test executions caught a breakage or fix, so a suite that only ever passes is telling you it is watching the wrong code.

Free Assessment

Get a free QA audit for your project

Identify quality gaps before they become production bugs.

Get Free Audit

Ship software with confidence

Talk to a QA advisor and find out how QAble can help your team build quality in at every stage.

No sales pitch
Technical walkthrough
No lock-in commitment

Talk to QA Advisor

Direct access to QAble's QA specialists.

Response within 24 hours