In April 2026, the NCSC and the US Cybersecurity and Infrastructure Security Agency jointly warned that a backdoor named Firestarter had been used to keep access to compromised Cisco firewall appliances. It worked even after those devices had been fully patched and updated. That was a blunt reminder. A firewall being up to date and a firewall being secure are not the same claim, since one is about patching and the other is about configuration. Firewall penetration testing is how you find out which one you actually have.

Firewall penetration testing is a controlled, hands-on attempt to bypass, break through or misuse a firewall. A tester does this the way a real attacker would, so it goes well beyond simply confirming the device is switched on and running the latest firmware.
Table of Contents
What does firewall penetration testing actually involve?
A firewall penetration testing engagement works through several angles at once:
- Rule base analysis – hunting for overly broad “allow” rules, contradictory rules, and temporary access that was never removed.
- Default-deny enforcement – checking the firewall genuinely blocks everything it isn’t explicitly told to allow.
- Management interface exposure – checking whether admin access to the firewall can be reached from the internet, or from network segments it shouldn’t be visible to.
- NAT and VPN configuration – a misconfigured port forward or VPN setting can quietly open a route straight past the rule base.
- The firewall’s own software – an unpatched or poorly hardened firewall is itself a target, not just a defence.
The NCSC’s network security guidance sets a clear standard for the rule base, because it calls for a final “deny all” rule, after permitting only the minimum traffic genuinely needed. Access should be the exception, not the default. Firewall penetration testing exists to confirm that standard is actually being met, not just written down somewhere.
It’s worth being clear about what a firewall test is not, too. A network firewall filters traffic at the network boundary. A web application firewall works one layer up, filtering requests to a specific web application. We’ve written before about why a web application firewall buys you time without fixing the underlying vulnerability. The same principle applies here. Either kind of firewall is a control worth testing. Neither is a substitute for fixing what it’s covering for.
Why doesn’t a rule review catch everything?
Reading through a rule base is a useful first step. It’s worth doing regularly, regardless of anything else. But it’s a static exercise. A reviewer flags what looks wrong to the eye, and nothing more. Two rules that each look reasonable alone can combine into a route nobody intended, and that kind of gap only shows up once someone actually tries to use it. Firewall penetration testing does exactly that. A document review can’t.
How is this different from a rule review or a vulnerability scan?
| Method | What it checks | Its limit |
|---|---|---|
| Rule-base review | A person reads through the configuration for anything that looks wrong. | Static; catches only what’s visible on inspection. |
| Vulnerability scan | An automated tool flags known, published weaknesses. | Misses rule combinations that only become a problem in context. |
| Firewall penetration test | A tester actively tries to bypass or misuse the firewall. | A point-in-time result; needs repeating as the setup changes. |
None of the three replaces the others. Reviews happen on every change. Scans run on a set cycle. A full penetration test is done annually, or whenever something significant shifts.
External and internal testing cover different ground
Firewall penetration testing usually runs from two directions:
- External testing simulates an attacker on the internet, probing exposed services, VPN endpoints and management ports from outside.
- Internal testing simulates a threat already inside the network. That might be a compromised device, an infected supplier laptop, or a malicious insider. It checks whether the firewall properly segments internal zones from one another.
External testing gets most of the attention, because that’s where the obvious threat seems to come from. Internal testing tends to reveal the more damaging failure: whether a breach that starts small can spread freely once it’s inside.
How often should a firewall be tested?
PCI DSS requirement 11.4 gives most UK businesses a concrete benchmark. It calls for internal and external penetration testing at least once every 12 months, and again after any significant change to segmentation controls, including firewall rule changes. Service providers face a shorter cycle. They must validate segmentation every six months instead of annually. This requirement trips up more businesses than you’d expect. We’ve covered why most businesses get PCI DSS penetration testing requirements wrong in more detail elsewhere.
Beyond formal compliance, firewall rule bases change constantly in ordinary business. A new supplier needs access. A project adds a temporary rule that outlives the project. Each change is a fresh chance for something to open that shouldn’t have. An annual firewall penetration test, plus a retest after any material change, is a sensible baseline for most organisations, whatever sector they’re in.
What drives the cost of a firewall penetration test?
Prices differ between providers, but the scope drivers stay the same. How many firewalls and interfaces need testing? Does the work cover external testing, internal testing, or both? How complicated has the rule base become, and how many sites are involved? Testing one firewall from the outside is a much smaller job than testing segmentation across several connected sites. A sensible quote reflects that difference, rather than quoting a flat fee regardless of scope. Aardwolf Security’s penetration testing service scopes firewall penetration testing this way, against exactly the kind of rule-base and segmentation checks covered above.
What does the testing process look like in practice?
Established methodology, including NIST’s Technical Guide to Information Security Testing and Assessment (SP 800-115), treats rule-set review and active testing as complementary stages. Neither one substitutes for the other. In practice, a firewall penetration test typically runs as:
- Scoping – agreeing which firewalls, interfaces and segments are in scope, plus the rules of engagement.
- Rule analysis – comparing the live rule base against what it was actually meant to allow.
- Active testing – attempting to bypass filtering and reach restricted services or management access, from both outside and inside the network.
- Reporting – documenting findings in enough detail to reproduce them, prioritised by what an attacker could actually do.
- Remediation and retesting – fixing what was found, then confirming the fix genuinely closes the gap.
Frequently asked questions
Isn’t a vulnerability scan enough?
No. A scan checks automatically for known, published issues. A penetration test is manual and adversarial instead. The tester actively tries to exploit what they find, including problems that only emerge in context and that no automated scan would ever flag.
Do next-generation firewalls still need this?
Yes. Application awareness and intrusion prevention add useful capability, but they also add more configuration to get wrong. A badly maintained rule base is exploitable no matter how advanced the underlying platform is.
How does this differ from a full network penetration test?
A firewall penetration test stays focused on the device’s rules, management access and boundary enforcement. A full network penetration test is broader. It covers the firewall plus servers, workstations and other infrastructure. Many businesses fold firewall testing into that wider engagement, rather than running it alone.
Can our IT team just do this themselves?
Ongoing in-house review of the rule base is good practice on its own. Independent testing still matters, because an outside tester brings no assumptions about how the network is “supposed” to behave, and that’s usually exactly where the real gap has been sitting unnoticed.
What should happen after the test finds something?
A useful report prioritises findings by real-world impact, rather than just listing them, because that’s what tells you what to fix first. A retest afterwards confirms the fix actually closed the gap, instead of simply changing what it looked like.
What should bring a test forward outside the usual schedule?
Several events are good reasons to test sooner than planned. A new site or data centre connection going live is one. So is a merger that joins two networks together, or a cyber insurance renewal that asks for evidence. Any incident, or a significant rule-base change, should also bring the next test forward.
If you’re not sure whether your firewall configuration would hold up to a proper test, or it’s simply been longer than a year since the last one, get in touch and we can talk through scope and timing.
Subscribe to our newsletter
Honest updates, straight to your inbox. Unsubscribe any time.