View all services
Talk to QA Advisor
/Blog/Software Security Testing: A Practical Methodology for QA Teams
Compliance Testing6 min read

Software Security Testing: A Practical Methodology for QA Teams

How to run security testing as part of delivery: what to test against under OWASP Top 10:2025, which technique catches which flaw class, how to wire checks into a pipeline, and how to triage the findings.

Published February 12, 2024Last updated September 22, 2026
On this page

Security testing fails most often for an organisational reason, not a technical one: the findings arrive after the release branch is cut, so they become next quarter's backlog instead of this release's fixes.

This guide sets out how to run security testing as part of delivery: what to test against, which techniques catch which flaw classes, how to wire it into a pipeline, and how to triage what comes back.

Test against the current OWASP Top 10

The OWASP Top 10 was revised in 2025. Plenty of security testing material still references the 2021 list, so it is worth checking which revision a tool, report or proposal is working from.

OWASP Top 10:2025
RankCategoryWhat to test for it
A01Broken Access ControlEvery role pair, horizontal and vertical. Manual, not scanned.
A02Security MisconfigurationHeaders, defaults, permissions, exposed management surfaces
A03Software Supply Chain FailuresDependency and build-integrity checks, not application scanning
A04Cryptographic FailuresTransport, storage, key handling, algorithm choice
A05InjectionSQL, command, template and interpreter injection
A06Insecure DesignThreat modelling at design time. No tool finds this later.
A07Authentication FailuresLogin, reset, MFA enrolment, session expiry, alternative routes
A08Software or Data Integrity FailuresUnverified updates, deserialisation, CI/CD integrity
A09Security Logging and Alerting FailuresRepeat an attack and check whether anything was recorded
A10Mishandling of Exceptional ConditionsError paths, failure modes, partial states
Source: owasp.org/Top10/2025

Three of these change how you scope a test programme.

Software Supply Chain Failures at A03 covers your build pipeline, dependency resolution and distribution process. Scanning the running application will not reach it. You need dependency and build-integrity checks as separate activities.

Security Misconfiguration at A02 is the category that rewards environment access. A tester working only from outside finds a fraction of what someone with configuration visibility finds.

Mishandling of Exceptional Conditions at A10 is new. Error paths, failure modes, partial states. These are hard to reach without knowing how the application is meant to behave, which makes documentation a security testing asset rather than a nice-to-have.

Match the technique to the flaw class

Teams often buy one tool and expect coverage. Each technique sees a different part of the problem, and the gaps between them are where flaws survive.

Which technique catches which flaw class
TechniqueSeesCannot seeRuns
SAST (static analysis)Source code, data flow, unsafe patternsRuntime behaviour, configuration, authorisation logicEvery commit, fast
DAST (dynamic scanning)A running app from the outsideCode paths never exercised, who owns which recordNightly or per release
SCA (dependency scanning)Known CVEs in declared dependenciesYour own code, transitive build-time riskEvery commit, fast
Manual testingBusiness logic, authorisation, chained flawsNothing structural, but bounded by tester timePer release or quarterly
Threat modellingDesign flaws before code existsImplementation defectsAt design, and on major change

The practical consequence is that SAST and DAST do not overlap enough to substitute for one another. SAST reads code it cannot run; DAST exercises a running system it cannot see inside. Authorisation flaws evade both, because neither knows which user should own which record.

A test sequence you can actually run

Ordered so that each step feeds the next, rather than as a list of activities to pick from.

  1. Enumerate roles and trust boundaries. Write down every role, including ones that exist only in configuration, and every point where data crosses from less trusted to more trusted. This document drives everything after it.
  2. Inventory the attack surface. Endpoints, parameters, file uploads, webhook receivers, admin functionality, third-party integrations. You cannot test what you have not listed.
  3. Run dependency and configuration checks first. These are cheap, automated, and clear the noise before manual work starts. A03 and A02 findings surface here.
  4. Run SAST against the code, DAST against a running instance. Triage both before manual testing, so testers are not rediscovering what tooling already found.
  5. Test authentication flows by hand. Login, logout, password reset, multi-factor enrolment, session expiry, and every alternative route in. The alternative routes carry the flaws.
  6. Test authorisation between every role pair. Two accounts per role, attempt horizontal access between peers and vertical access upward. This is the step most often skipped and it maps to A01, the top category.
  7. Test business logic against the money path. Checkout, refunds, credit transfers, quota enforcement. No scanner understands your domain rules.
  8. Verify logging and alerting. Repeat two or three earlier attacks and check whether anything was recorded. A09 exists because most applications fail this.

Authorisation testing deserves its own budget

Broken Access Control has stayed at the top of the list across revisions, and it is the flaw class automated tooling handles worst.

The reason is structural. A scanner can see that an endpoint returned data. It cannot know the data belonged to another customer, because it has no model of who owns what. Only someone who understands your domain can make that judgement.

To test it properly you need every role enumerated, two accounts per role, a clear statement of which objects belong to which tenant, and a list of endpoints whose behaviour changes based on a claim in a token.

If a security testing proposal does not mention authorisation testing explicitly, assume it is not being done.

Wiring it into the pipeline

Security testing that runs quarterly produces findings nobody scheduled. Security testing that runs on every commit and blocks the build produces developers who disable it. The split that survives contact with a delivery team:

Where each check belongs in the pipeline
StageWhat runsBlocks the build?
Pre-commit hookSecret detectionYes. Cheap and unambiguous.
Pull requestSAST on changed files, SCA on the lockfileYes, on high-confidence rules only
Merge to mainFull SAST, full dependency scanNo. Raises tickets.
NightlyDAST against a deployed environmentNo
Per releaseManual authorisation and business logic testingRelease gate, by judgement
QuarterlyExternal penetration testNo

The principle is that fast, low-false-positive checks gate the build, while slow or judgement-heavy work runs on a schedule and feeds the backlog with tickets rather than blocking a merge.

Triage what comes back

A scanner run produces more findings than any team can action. Triage is where security testing either delivers value or becomes a report nobody opens.

For each finding, three questions in order:

Is it real? Reproduce it manually. Static and dynamic tooling both produce false positives, and a queue full of them trains people to ignore the queue.

Is it reachable? A vulnerable function that no request path reaches is a lower priority than a weaker flaw on your login page. Severity scores do not know your architecture.

What is the blast radius here? A high-severity finding on an internal tool behind a VPN may matter less than a medium on your public checkout. Rate impact in your context, not against the generic score.

Then fix in reachability order, not severity order. A queue worked strictly by CVSS leaves exploitable medium findings untouched while people chase unreachable criticals.

From our work

🔬 From our work
We do not publish client vulnerability data, so this is an unquantified observation from our security testing practice rather than a measured statistic. On a team's first structured security testing cycle, the findings that matter most are usually authorisation flaws rather than injection flaws. Injection is well understood and frameworks defend against it by default. Authorisation is bespoke to every application, no framework enforces it, and it is typically implemented one endpoint at a time by different people over several years.

The second pattern is that findings cluster at the oldest and newest parts of a codebase: the parts written before the team settled on conventions, and the parts written most recently, before those conventions were applied consistently.

Common failures

Testing only what is easy to test. Unauthenticated scanning of public pages is cheap and covers the smallest, best-defended part of your attack surface.

Treating the report as the deliverable. The deliverable is remediation. Budget engineering time for fixes in the same planning cycle as the testing, or you will hold a document describing flaws you have not fixed.

Running tools without tuning them. Default rule sets produce findings nobody is required to fix, which makes the whole report easy to dismiss.

No retest. A finding is not closed because a ticket is closed. Agree retest scope when you commission the work, not afterwards.

Related reading

If you would like help building security testing into your delivery pipeline, talk to our security testing team.

Frequently Asked Questions

Where should security testing sit in a delivery pipeline?

Split it by speed and false-positive rate. Secret detection belongs in a pre-commit hook, SAST on changed files and dependency scanning at pull request, full scans on merge to main, dynamic scanning nightly, and manual authorisation testing per release. Fast unambiguous checks block the build; slow or judgement-heavy work raises tickets instead.

Do SAST and DAST cover the same ground?

No, and treating them as interchangeable leaves a gap. SAST reads source it cannot execute, so it misses runtime and configuration problems. DAST exercises a running system it cannot see inside, so it misses code paths no request reaches. Neither finds authorisation flaws, because neither knows which user should own which record.

Why does authorisation testing need people rather than tools?

A scanner can tell that an endpoint returned data. It cannot tell the data belonged to a different customer, because it has no model of ownership in your domain. Testing it properly needs every role enumerated, two accounts per role so peer-level access can be attempted, and a statement of which objects belong to which tenant.

How do we triage more findings than we can fix?

Ask three questions in order. Is it real, reproduced by hand rather than trusted from a scanner. Is it reachable, since a vulnerable function no request path reaches matters less than a weaker flaw on your login page. What is the blast radius in your context. Then fix in reachability order rather than strictly by severity score.

How often should security testing run?

Automated checks run continuously, on every commit and every night. Manual testing runs per release or quarterly. External penetration testing runs when something material changes: a new authentication mechanism, a new payment flow, a significant architectural change, or a new integration handling sensitive data.

What changed in the OWASP Top 10 for 2025?

Software Supply Chain Failures moved to A03 and now covers your build pipeline and dependency resolution rather than only outdated components. Security Misconfiguration rose to A02. Mishandling of Exceptional Conditions is new at A10, covering error paths and partial states. Material still citing the 2021 list is two revisions behind.

Should security testing block a build?

Only where the check is fast and the finding is unambiguous. Secret detection and high-confidence SAST rules are reasonable gates. Blocking a merge on a full dynamic scan or a noisy rule set teaches developers to disable the gate, which is worse than not having one.

What makes a security finding actionable?

Reproduction steps with the exact request and response, impact rated against your architecture rather than a generic score, and remediation naming the specific endpoint and check to add. "Implement proper access control" is a category, not a fix. Agree retest scope when you commission the work rather than negotiating it afterwards.

Free Assessment

Get a free QA audit for your project

Identify quality gaps before they become production bugs.

Get Free Audit

Ship software with confidence

Talk to a QA advisor and find out how QAble can help your team build quality in at every stage.

No sales pitch
Technical walkthrough
No lock-in commitment

Talk to QA Advisor

Direct access to QAble's QA specialists.

Response within 24 hours