Skip to main content
B-52 · Agentic Penetration Testing Platform

Continuous penetration testing between the deep engagements

A booked engagement measures the estate in the week it runs. Everything that ships after it is untested until the next one is scoped. B-52 runs in that interval, at whatever frequency your releases actually go out.

The security teams who use it this way already buy manual penetration tests and carry on buying them. What changes is the interval on either side of one.

What happens on a run

What the interval holds

The estate does not hold still between two engagements

This is an argument about the weeks after the engagement, where every change ships on the assumption that the last test still describes the system it was run against.

When a change gets tested, under a booked engagement and under a standing cadence
What changed in the estateTested under a booked engagementTested under a continuous cadence
A release that alters an authenticated flow At the next engagement, however many releases later that turns out to be. On the run that follows the release, against the flow as it now behaves.
An endpoint added for one integration If it is still there, and still in scope, when the next engagement is scoped. On the next run against that application, because the scope is standing rather than renegotiated each time.
A mobile build shipped to the stores At the next engagement that covers mobile. On the run that follows the build. B-52 decompiles and analyses the APK or IPA itself, so a release build is testable without a source drop.
A host that an infrastructure change exposed When it appears in the scope of a later engagement. When the next run reaches it, provided it sits inside the scope already authorised.

When a change gets tested, under a booked engagement and under a standing cadence

A release that alters an authenticated flow

Tested under a booked engagement
At the next engagement, however many releases later that turns out to be.
Tested under a continuous cadence
On the run that follows the release, against the flow as it now behaves.

An endpoint added for one integration

Tested under a booked engagement
If it is still there, and still in scope, when the next engagement is scoped.
Tested under a continuous cadence
On the next run against that application, because the scope is standing rather than renegotiated each time.

A mobile build shipped to the stores

Tested under a booked engagement
At the next engagement that covers mobile.
Tested under a continuous cadence
On the run that follows the build. B-52 decompiles and analyses the APK or IPA itself, so a release build is testable without a source drop.

A host that an infrastructure change exposed

Tested under a booked engagement
When it appears in the scope of a later engagement.
Tested under a continuous cadence
When the next run reaches it, provided it sits inside the scope already authorised.

Why a run in the gap is worth reading

What a run in the gap inherits from the engagement

Security Brigade’s full assessment record sits underneath this. The firm has been CERT-In empanelled since 2008, and every engagement it has run since it started in 2006 was worked inside Lemon, our own assessment platform, rather than dispersed across the auditors who ran them.

6,700+ assessments of test cases, vulnerabilities and threat models is the corpus behind our models. The harness was then written to work through what a Security Brigade auditor works through, which is what a run in the interval inherits.

The harness

The four practices every run works through

Mindmap creation
The target is mapped as it stands on the day of the run — entry points, roles, and which role can reach whose data. A map drawn three releases ago describes an application that no longer exists.
Test-case generation
The checks come from what this application does, so a flow added last month is tested as itself rather than against a list written before anyone had seen it.
Comprehensive JavaScript analysis
The client-side routes, parameters and endpoints nothing in the interface links to. In an interval nobody scoped, this is where the change that was never mentioned turns up.
Functional flow analysis
Multi-step flows and the state they carry, taken out of order and taken by the wrong role. Business logic is the part a release breaks quietly.

What autonomy adds on top of them

The generated test set is worked through in full, because there is no timebox to sample it down to. The same practices run the same way against the same target in March and in September, so a difference between two reports is a difference in the application rather than a difference in who was free that week.

And a run can happen at a cadence no human schedule supports, which is the only reason the interval can be covered at all.

What the cadence is worth

We measured it against our own assessors before we offered it as a schedule

B-52 was run in parallel with Security Brigade’s own expert assessment team, on the same targets, and each side’s findings were counted. Against the combined set — everything either party found, counted once — B-52 reached 90–95%, and it surfaced issues the human team did not.

Comparable coverage, different blind spots. That is the whole of this page: a run of that standard can happen on the Tuesday because a release went out on the Monday, and it needs a scope that was already signed off rather than a place in anyone’s calendar.

How the parallel benchmark was run
  • B-52 and Security Brigade’s own expert assessment team worked the same targets in parallel.
  • The denominator is the combined findings set: everything either party found, counted once.
  • B-52 reached 90–95% of that combined set.
  • Part of what B-52 found was absent from the human team’s own set, which is why the combined set is larger than either side’s.

Why this is the honest case for both models

The result reads two ways and both are true. A platform working alone covers most of what an expert team covers, which is what makes an unattended cadence worth running. The platform and the auditor together cover more than either does alone, which is what makes the verified engagement worth booking.

A continuous programme is the first reading. The engagement you file is the second. They are different jobs, they are bought separately, and neither answer makes the other unnecessary.

The benchmark in full

What a programme consists of

Sign off the scope, then read the reports

Four things, and only the first one asks anything of your team on an ordinary week.

Three actions still stop and wait for your written approval wherever they arise: destructive or state-changing actions in production, persistence and lateral movement beyond the entry host, and anything touching live credentials or real customer data. Persistence, once approved, is established by the platform rather than handed to an operator.

Where the targets come from

ShadowMap discovers. B-52 tests what it finds.

ShadowMap is Security Brigade’s other product: it watches an organisation’s external surface and reports what is exposed. Where a customer runs both, the flow between them goes one way — something ShadowMap surfaces can become a target B-52 is authorised to test. Nothing travels back in the other direction.

ShadowMap: what exists

The external surface as it is today, including the parts nobody registered with your team. Its output is an inventory, and an inventory is a list of things nobody has tested yet.

B-52: what happens when it is attacked

A target added to an authorised scope is tested on the next run, and what comes back is a finding with the request, the response and the steps that reproduce it.

The two are subscribed separately. B-52 is an adjacent product rather than a ShadowMap module, it sits outside a ShadowMap licence, and a continuous B-52 cadence runs perfectly well for a customer who has never bought ShadowMap.

What you can file it as

Which delivery model a regulator will accept

All three delivery models test the same eleven classes. What separates them is where the human sits, and that is what decides whether the output can go to a regulator.

The three delivery models and what a regulated filing can be built from
Delivery modelWhere the human isWhat you can file it as
Fully autonomous The model the interval runs in. Scope is authorised, then nobody acts until the report arrives. Not signable under Security Brigade’s CERT-In empanelment.
Autonomous, expert verified A senior auditor verifies every finding before it reaches you. This is the model a filed assessment runs in. Signable under Security Brigade’s CERT-In empanelment.
Human led A senior auditor runs the engagement with B-52 underneath. Signable under Security Brigade’s CERT-In empanelment.
Key
  • An empanelled auditor is in the engagement, so the output is signable for a regulated filing
  • No empanelled auditor in the engagement

A continuous programme surrounds the filed assessment rather than standing in for it. Where the output is going to a regulator, run it expert-verified or human-led: an empanelled auditor has to be in the engagement for what comes out of it to be signable, and CERT-In empanelment attaches to the testing itself and not only to the vendor selling it.

Compare the three delivery models

The cadence is a number you choose

One scan is one application or one target, and the ladder starts at $500. A programme is priced on three things: how many targets, how often they are tested, and which delivery model runs them.

See what a run covers