Security testing fails most often for an organisational reason, not a technical one: the findings arrive after the release branch is cut, so they become next quarter's backlog instead of this release's fixes.
This guide sets out how to run security testing as part of delivery: what to test against, which techniques catch which flaw classes, how to wire it into a pipeline, and how to triage what comes back.
Test against the current OWASP Top 10
The OWASP Top 10 was revised in 2025. Plenty of security testing material still references the 2021 list, so it is worth checking which revision a tool, report or proposal is working from.
Three of these change how you scope a test programme.
Software Supply Chain Failures at A03 covers your build pipeline, dependency resolution and distribution process. Scanning the running application will not reach it. You need dependency and build-integrity checks as separate activities.
Security Misconfiguration at A02 is the category that rewards environment access. A tester working only from outside finds a fraction of what someone with configuration visibility finds.
Mishandling of Exceptional Conditions at A10 is new. Error paths, failure modes, partial states. These are hard to reach without knowing how the application is meant to behave, which makes documentation a security testing asset rather than a nice-to-have.
Match the technique to the flaw class
Teams often buy one tool and expect coverage. Each technique sees a different part of the problem, and the gaps between them are where flaws survive.
The practical consequence is that SAST and DAST do not overlap enough to substitute for one another. SAST reads code it cannot run; DAST exercises a running system it cannot see inside. Authorisation flaws evade both, because neither knows which user should own which record.
A test sequence you can actually run
Ordered so that each step feeds the next, rather than as a list of activities to pick from.
- Enumerate roles and trust boundaries. Write down every role, including ones that exist only in configuration, and every point where data crosses from less trusted to more trusted. This document drives everything after it.
- Inventory the attack surface. Endpoints, parameters, file uploads, webhook receivers, admin functionality, third-party integrations. You cannot test what you have not listed.
- Run dependency and configuration checks first. These are cheap, automated, and clear the noise before manual work starts. A03 and A02 findings surface here.
- Run SAST against the code, DAST against a running instance. Triage both before manual testing, so testers are not rediscovering what tooling already found.
- Test authentication flows by hand. Login, logout, password reset, multi-factor enrolment, session expiry, and every alternative route in. The alternative routes carry the flaws.
- Test authorisation between every role pair. Two accounts per role, attempt horizontal access between peers and vertical access upward. This is the step most often skipped and it maps to A01, the top category.
- Test business logic against the money path. Checkout, refunds, credit transfers, quota enforcement. No scanner understands your domain rules.
- Verify logging and alerting. Repeat two or three earlier attacks and check whether anything was recorded. A09 exists because most applications fail this.
Authorisation testing deserves its own budget
Broken Access Control has stayed at the top of the list across revisions, and it is the flaw class automated tooling handles worst.
The reason is structural. A scanner can see that an endpoint returned data. It cannot know the data belonged to another customer, because it has no model of who owns what. Only someone who understands your domain can make that judgement.
To test it properly you need every role enumerated, two accounts per role, a clear statement of which objects belong to which tenant, and a list of endpoints whose behaviour changes based on a claim in a token.
If a security testing proposal does not mention authorisation testing explicitly, assume it is not being done.
Wiring it into the pipeline
Security testing that runs quarterly produces findings nobody scheduled. Security testing that runs on every commit and blocks the build produces developers who disable it. The split that survives contact with a delivery team:
The principle is that fast, low-false-positive checks gate the build, while slow or judgement-heavy work runs on a schedule and feeds the backlog with tickets rather than blocking a merge.
Triage what comes back
A scanner run produces more findings than any team can action. Triage is where security testing either delivers value or becomes a report nobody opens.
For each finding, three questions in order:
Is it real? Reproduce it manually. Static and dynamic tooling both produce false positives, and a queue full of them trains people to ignore the queue.
Is it reachable? A vulnerable function that no request path reaches is a lower priority than a weaker flaw on your login page. Severity scores do not know your architecture.
What is the blast radius here? A high-severity finding on an internal tool behind a VPN may matter less than a medium on your public checkout. Rate impact in your context, not against the generic score.
Then fix in reachability order, not severity order. A queue worked strictly by CVSS leaves exploitable medium findings untouched while people chase unreachable criticals.
From our work
The second pattern is that findings cluster at the oldest and newest parts of a codebase: the parts written before the team settled on conventions, and the parts written most recently, before those conventions were applied consistently.
Common failures
Testing only what is easy to test. Unauthenticated scanning of public pages is cheap and covers the smallest, best-defended part of your attack surface.
Treating the report as the deliverable. The deliverable is remediation. Budget engineering time for fixes in the same planning cycle as the testing, or you will hold a document describing flaws you have not fixed.
Running tools without tuning them. Default rule sets produce findings nobody is required to fix, which makes the whole report easy to dismiss.
No retest. A finding is not closed because a ticket is closed. Agree retest scope when you commission the work, not afterwards.
Related reading
- Penetration testing of web applications for the full engagement methodology
- Penetration testing costs for what engagements typically cost and what drives the range
- Cyber security testing techniques for the wider technique catalogue
- ChatGPT in testing for where language models help and where they do not
- Desktop application testing and IoT application testing for platform-specific security considerations
- How Google tests software for the organisational side of testing at scale
- Software testing in logistics and transportation for a regulated-domain example
- Cross-browser testing using Selenium for the client-side testing surface
- Choosing a software testing partner if you are evaluating external help
If you would like help building security testing into your delivery pipeline, talk to our security testing team.