Python is the default choice for test automation in most QA teams, and the reason is rarely the language itself. It is that the surrounding tooling, pytest in particular, removes enough boilerplate that tests get written at all.
This guide covers how to structure a Python automation suite that stays maintainable: fixtures, parametrisation, page objects, and the CI wiring that turns a folder of scripts into a suite people trust.
Choose the runner before the framework
Most Python automation decisions follow from one choice: which test runner you standardise on. Everything else plugs into it.
pytest is the pragmatic default for QA work. Its fixture model handles setup and teardown without inheritance, its assertion rewriting means you write plain assert statements and still get useful failure output, and the plugin ecosystem covers parallelism, retries and reporting.
Stay with unittest when your team already has a large suite in it, or when you are constrained to the standard library.
Fixtures are the part that matters
The single biggest difference between a Python suite that lasts and one that rots is how test dependencies are supplied. Fixtures are pytest's answer, and they replace the setup and teardown methods that tie older suites in knots.
import pytest
from selenium import webdriver
@pytest.fixture(scope="session")
def driver():
"""One browser for the whole run. Closed even if tests fail."""
d = webdriver.Chrome()
d.implicitly_wait(0) # explicit waits only, never implicit
yield d
d.quit()
@pytest.fixture
def logged_in(driver, base_url):
"""Fresh session per test, so tests cannot leak state into each other."""
driver.get(f"{base_url}/login")
LoginPage(driver).sign_in("[email protected]", "…")
yield driver
driver.delete_all_cookies()Three things that pay off later.
Scope deliberately. A session fixture runs once for the whole suite, function runs per test. Browser startup is expensive, so share it; authentication state is dangerous to share, so do not.
Clean up after `yield`. Everything after the yield runs even when the test fails, which is what stops one broken test poisoning the next twenty.
Compose rather than inherit. A fixture can request other fixtures, as logged_in requests driver. This replaces the base-class hierarchies that older suites accumulate.
Parametrise instead of copying
The commonest source of unmaintainable Python suites is the copied test. One test body, ten near-identical variants, and a change to the behaviour means ten edits.
@pytest.mark.parametrize(
"email,password,expected",
[
("[email protected]", "correct-horse", "dashboard"),
("[email protected]", "wrong", "Invalid credentials"),
("", "correct-horse", "Email is required"),
("not-an-email", "correct-horse", "Enter a valid email"),
],
ids=["happy", "bad-password", "missing-email", "malformed-email"],
)
def test_login(page, email, password, expected):
page.sign_in(email, password)
assert expected in page.result_text()The ids argument matters more than it looks. Without it a failure reports test_login[[email protected] credentials]. With it, the failure says test_login[bad-password], and whoever reads the CI output at 9am knows what broke.
Keep selectors out of tests
Page objects are not ceremony. They exist so that a change to the UI touches one file rather than forty.
class LoginPage:
EMAIL = (By.CSS_SELECTOR, "[data-test=email]")
PASSWORD = (By.CSS_SELECTOR, "[data-test=password]")
SUBMIT = (By.CSS_SELECTOR, "[data-test=submit]")
def __init__(self, driver):
self.driver = driver
self.wait = WebDriverWait(driver, 10)
def sign_in(self, email, password):
self.wait.until(EC.visibility_of_element_located(self.EMAIL)) \
.send_keys(email)
self.driver.find_element(*self.PASSWORD).send_keys(password)
self.driver.find_element(*self.SUBMIT).click()
return selfTwo rules that prevent most flakiness.
Target test-specific attributes, not CSS classes. [data-test=submit] survives a restyle; .btn.btn-primary.pull-right does not. This needs agreement from your developers, and it is worth asking for.
Use explicit waits, never implicit ones. An implicit wait applies globally and interacts badly with explicit waits, producing timeouts that are hard to reason about. Wait for the condition you actually need.
Structure the project so CI is straightforward
tests/
conftest.py # shared fixtures, discovered automatically
pages/ # page objects, no assertions in here
api/ # request-level tests, fast
e2e/ # browser tests, slow
pytest.ini # markers, default options
requirements.txtconftest.py is discovered by pytest without imports, which is why shared fixtures belong there. Splitting api/ from e2e/ lets CI run the fast tests on every commit and the slow ones on a schedule.
Register your markers in pytest.ini so a typo in a marker name fails loudly instead of silently selecting nothing:
[pytest]
markers =
smoke: minimal set that must pass before anything else
e2e: full browser journey, slow
addopts = --strict-markers -ra--strict-markers turns an unregistered marker into an error. -ra prints a summary of everything that did not pass, which is the output you want in a CI log.
Running in CI
Three flags do most of the work:
-n autowithpytest-xdistdistributes tests across available cores, which is the cheapest speed improvement available--junitxml=results.xmlproduces the format nearly every CI system parses for test reporting-m "not e2e"selects the fast subset for commit builds
The failure mode to watch for is tests that pass alone and fail in parallel. That is almost always shared state: a session-scoped fixture holding data one test mutates. When it happens, narrow the fixture scope before reaching for a retry plugin.
Retries deserve a caution. pytest-rerunfailures will make a flaky suite green, and a green flaky suite is worse than a red one, because it stops reporting the instability that is costing you confidence. Use retries to survive a known infrastructure problem, not to hide a test defect.
From our work
Common mistakes
Assertions inside page objects. The page object describes the page; the test decides what is correct. Mixing them means a page object cannot be reused by a test with different expectations.
Sleeping instead of waiting. time.sleep(5) is slow when the app is fast and still flaky when the app is slow. Wait for a condition.
One enormous conftest.py. Fixtures used by three tests do not belong in the root conftest. pytest discovers conftest.py at every level of the tree, so scope them to the directory that uses them.
Testing through the UI what an API call could verify. Browser tests are the slowest and most brittle tests you own. Use them for journeys that genuinely need a browser, and push everything else down to the API layer.
Related reading
- Selenium: an introductory approach for the browser library most of these examples drive
- Test automation frameworks compared if you are still choosing rather than refining
- Choosing an AI unit testing approach for where generated tests fit alongside a hand-written suite
- Static testing for the checks that run before anything executes
- Cloud testing for running suites against environments you do not host
- Salesforce testing for the platform-specific case
- Visual testing for the assertion type a Python suite handles worst
As a software testing agency we work with teams running Python suites across mobile applications, functional testing and QA as a service engagements, and we run a QA audit when a suite needs assessing before it is extended.
If you would like a review of your existing Python suite before expanding it, talk to our automation testing team.