View all services
Talk to QA Advisor
/Blog/Python Automation Testing: A Practical Guide for QA Teams
Automation Testing6 min read

Python Automation Testing: A Practical Guide for QA Teams

How to structure a Python automation suite that stays maintainable: pytest fixtures and scope, parametrisation, page objects, project layout and the CI wiring that turns scripts into a suite people trust.

Published December 5, 2023Last updated September 22, 2026
On this page

Python is the default choice for test automation in most QA teams, and the reason is rarely the language itself. It is that the surrounding tooling, pytest in particular, removes enough boilerplate that tests get written at all.

This guide covers how to structure a Python automation suite that stays maintainable: fixtures, parametrisation, page objects, and the CI wiring that turns a folder of scripts into a suite people trust.

Choose the runner before the framework

Most Python automation decisions follow from one choice: which test runner you standardise on. Everything else plugs into it.

Python test runners for QA automation
RunnerSetup and teardownBest forWatch out for
pytestFixtures, composed by requestMost QA automation work, API and browserFixture scope mistakes leak state between tests
unittestsetUp and tearDown methods, inheritanceStandard-library-only constraints, existing large suitesClass hierarchies grow hard to follow
Robot FrameworkKeywords and library importsTeams where non-programmers author testsA second syntax to maintain alongside Python
behaveGiven/When/Then stepsSuites where BDD is genuinely used by the businessBDD ceremony with no business reader is pure overhead

pytest is the pragmatic default for QA work. Its fixture model handles setup and teardown without inheritance, its assertion rewriting means you write plain assert statements and still get useful failure output, and the plugin ecosystem covers parallelism, retries and reporting.

Stay with unittest when your team already has a large suite in it, or when you are constrained to the standard library.

Fixtures are the part that matters

The single biggest difference between a Python suite that lasts and one that rots is how test dependencies are supplied. Fixtures are pytest's answer, and they replace the setup and teardown methods that tie older suites in knots.

import pytest
from selenium import webdriver


@pytest.fixture(scope="session")
def driver():
    """One browser for the whole run. Closed even if tests fail."""
    d = webdriver.Chrome()
    d.implicitly_wait(0)          # explicit waits only, never implicit
    yield d
    d.quit()


@pytest.fixture
def logged_in(driver, base_url):
    """Fresh session per test, so tests cannot leak state into each other."""
    driver.get(f"{base_url}/login")
    LoginPage(driver).sign_in("[email protected]", "…")
    yield driver
    driver.delete_all_cookies()

Three things that pay off later.

Scope deliberately. A session fixture runs once for the whole suite, function runs per test. Browser startup is expensive, so share it; authentication state is dangerous to share, so do not.

Clean up after `yield`. Everything after the yield runs even when the test fails, which is what stops one broken test poisoning the next twenty.

Compose rather than inherit. A fixture can request other fixtures, as logged_in requests driver. This replaces the base-class hierarchies that older suites accumulate.

Parametrise instead of copying

The commonest source of unmaintainable Python suites is the copied test. One test body, ten near-identical variants, and a change to the behaviour means ten edits.

@pytest.mark.parametrize(
    "email,password,expected",
    [
        ("[email protected]", "correct-horse", "dashboard"),
        ("[email protected]", "wrong", "Invalid credentials"),
        ("", "correct-horse", "Email is required"),
        ("not-an-email", "correct-horse", "Enter a valid email"),
    ],
    ids=["happy", "bad-password", "missing-email", "malformed-email"],
)
def test_login(page, email, password, expected):
    page.sign_in(email, password)
    assert expected in page.result_text()

The ids argument matters more than it looks. Without it a failure reports test_login[[email protected] credentials]. With it, the failure says test_login[bad-password], and whoever reads the CI output at 9am knows what broke.

Keep selectors out of tests

Page objects are not ceremony. They exist so that a change to the UI touches one file rather than forty.

class LoginPage:
    EMAIL = (By.CSS_SELECTOR, "[data-test=email]")
    PASSWORD = (By.CSS_SELECTOR, "[data-test=password]")
    SUBMIT = (By.CSS_SELECTOR, "[data-test=submit]")

    def __init__(self, driver):
        self.driver = driver
        self.wait = WebDriverWait(driver, 10)

    def sign_in(self, email, password):
        self.wait.until(EC.visibility_of_element_located(self.EMAIL)) \
            .send_keys(email)
        self.driver.find_element(*self.PASSWORD).send_keys(password)
        self.driver.find_element(*self.SUBMIT).click()
        return self

Two rules that prevent most flakiness.

Target test-specific attributes, not CSS classes. [data-test=submit] survives a restyle; .btn.btn-primary.pull-right does not. This needs agreement from your developers, and it is worth asking for.

Use explicit waits, never implicit ones. An implicit wait applies globally and interacts badly with explicit waits, producing timeouts that are hard to reason about. Wait for the condition you actually need.

Structure the project so CI is straightforward

tests/
  conftest.py          # shared fixtures, discovered automatically
  pages/               # page objects, no assertions in here
  api/                 # request-level tests, fast
  e2e/                 # browser tests, slow
pytest.ini             # markers, default options
requirements.txt

conftest.py is discovered by pytest without imports, which is why shared fixtures belong there. Splitting api/ from e2e/ lets CI run the fast tests on every commit and the slow ones on a schedule.

Register your markers in pytest.ini so a typo in a marker name fails loudly instead of silently selecting nothing:

[pytest]
markers =
    smoke: minimal set that must pass before anything else
    e2e: full browser journey, slow
addopts = --strict-markers -ra

--strict-markers turns an unregistered marker into an error. -ra prints a summary of everything that did not pass, which is the output you want in a CI log.

Running in CI

Three flags do most of the work:

  • -n auto with pytest-xdist distributes tests across available cores, which is the cheapest speed improvement available
  • --junitxml=results.xml produces the format nearly every CI system parses for test reporting
  • -m "not e2e" selects the fast subset for commit builds

The failure mode to watch for is tests that pass alone and fail in parallel. That is almost always shared state: a session-scoped fixture holding data one test mutates. When it happens, narrow the fixture scope before reaching for a retry plugin.

Retries deserve a caution. pytest-rerunfailures will make a flaky suite green, and a green flaky suite is worse than a red one, because it stops reporting the instability that is costing you confidence. Use retries to survive a known infrastructure problem, not to hide a test defect.

From our work

🔬 From our work
An unquantified observation from our automation practice rather than a measured figure. When we review an existing Python suite, the most common cause of unreliability is not the browser library or the application under test. It is shared state: a session-scoped fixture holding data one test mutates, which only surfaces when the suite starts running in parallel. The usual response is to add a retry plugin, which turns a visible problem into an invisible one. Narrowing the fixture scope fixes the cause.

Common mistakes

Assertions inside page objects. The page object describes the page; the test decides what is correct. Mixing them means a page object cannot be reused by a test with different expectations.

Sleeping instead of waiting. time.sleep(5) is slow when the app is fast and still flaky when the app is slow. Wait for a condition.

One enormous conftest.py. Fixtures used by three tests do not belong in the root conftest. pytest discovers conftest.py at every level of the tree, so scope them to the directory that uses them.

Testing through the UI what an API call could verify. Browser tests are the slowest and most brittle tests you own. Use them for journeys that genuinely need a browser, and push everything else down to the API layer.

Related reading

As a software testing agency we work with teams running Python suites across mobile applications, functional testing and QA as a service engagements, and we run a QA audit when a suite needs assessing before it is extended.

If you would like a review of your existing Python suite before expanding it, talk to our automation testing team.

Frequently Asked Questions

Should we use pytest or unittest for a new Python test suite?

pytest for most QA automation. Its fixture model supplies test dependencies without inheritance, assertion rewriting gives useful failure output from plain assert statements, and the plugin ecosystem covers parallelism and reporting. Stay with unittest when you are constrained to the standard library or already maintain a large suite in it.

What fixture scope should we use for a browser?

Session scope for the browser itself, because startup is expensive and sharing it saves real time. Function scope for anything carrying authentication or user state, because sharing that lets one test leak into the next. Getting this wrong is the most common cause of suites that pass alone and fail in parallel.

Why do our tests pass individually but fail when run in parallel?

Almost always shared state, usually a session-scoped fixture holding data that one test mutates. The fix is narrowing the fixture scope so each test gets its own copy. Adding a retry plugin makes the symptom disappear while the cause remains, which turns a visible problem into an invisible one.

How do we stop selector changes breaking the whole suite?

Keep selectors in page objects rather than tests, so a UI change touches one file. Target test-specific attributes like data-test rather than CSS classes, since a restyle changes classes but not test hooks. That second part needs agreement from your developers and is worth asking for.

When should we use parametrize instead of separate tests?

Whenever the test body is the same and only the inputs differ. It keeps one implementation to maintain instead of ten near-identical copies. Always pass the ids argument, or a failure reports the raw parameter values instead of a readable case name, which makes CI output hard to act on.

Should we use implicit or explicit waits?

Explicit waits only. An implicit wait applies globally and interacts badly with explicit waits, producing timeouts that are difficult to reason about. Wait for the specific condition you need, and never use a fixed sleep, which is slow when the application is fast and still flaky when it is slow.

How should we structure the project for CI?

Keep shared fixtures in conftest.py, which pytest discovers without imports, and split fast API tests from slow browser tests in separate directories. That lets CI run the fast subset on every commit and the slow ones on a schedule. Register markers in pytest.ini with strict-markers so a typo fails loudly instead of silently selecting nothing.

Is it worth testing through the UI when an API call would do?

Usually not. Browser tests are the slowest and most brittle tests you own, so reserve them for journeys that genuinely need a browser and push everything else down to the API layer. A suite weighted towards the UI is the commonest reason test runs become too slow to run on every commit.

Free Assessment

Get a free QA audit for your project

Identify quality gaps before they become production bugs.

Get Free Audit

Ship software with confidence

Talk to a QA advisor and find out how QAble can help your team build quality in at every stage.

No sales pitch
Technical walkthrough
No lock-in commitment

Talk to QA Advisor

Direct access to QAble's QA specialists.

Response within 24 hours