Skip to main content
Delivery model · Autonomous, expert verified

A senior auditor verifies every finding

B-52 works the scope exactly as it does unattended. Then a Security Brigade auditor goes through the findings one by one, and only what has been through both reaches you.

This is where an engagement starts unless something argues it somewhere else. Two reasons for that are set out below, and the second one is not ours.

Why this one first

B-52 and our own assessors found different things

We measured that rather than assumed it. B-52 was run in parallel with Security Brigade’s own expert assessment team, on the same targets, and what each of them found was counted.

B-52 reached 90–95% of the combined findings set, and it also surfaced issues the human team did not. Both halves of that sentence matter here. The first says the platform holds its own against the assessors it was built from; the second says the combined set is larger than either party produced on its own.

Comparable coverage, different blind spots. An engagement that puts both in the same run is the arrangement that result argues for, and this model is that arrangement.

And the market moved the same way

In Cobalt’s 2026 survey of 455 respondents, the share who would rely entirely on automation fell from 29% to 9% year on year, while preference for a hybrid of automation and human testers rose 22 points to 47%.

Cobalt, 455 respondents, published 25 June 2026.

How the parallel benchmark was measured
  • B-52 and Security Brigade’s own expert assessors worked the same targets at the same time.
  • The denominator is the combined findings set: everything either party found, counted once.
  • B-52 reached 90–95% of that set.
  • Issues the human team did not find are part of why the combined set is larger than either party’s own tally.

Deliberately excluded

  • Any public leaderboard or third-party score. The comparison that mattered was against the assessors whose work the platform was built from.

Where the auditor sits

The auditor verifies findings, and does nothing else

The auditor does not scope the engagement, direct the run or take work off the platform. One thing happens, at one point, to everything.

Verification

What passes through the auditor, and what comes out

What the run produces
Findings from all eleven coverage classes, each one carrying the request, the response and the steps to reproduce it. Candidates that could not be proved never became findings.
What the auditor does with them
Verifies every finding before any of it is released. Not a sample, not the high-severity ones, and not a skim of the summary — each finding in the set.
What reaches you
A report in which every line has been produced by the platform and confirmed by a named senior auditor, and which Security Brigade can sign under its CERT-In empanelment.

If you want the auditor earlier than this

Verification happens after the run, on the output. Where you need a person setting the direction rather than checking the result — an unusual scope, an application whose business logic has to be understood before it can be attacked — the human-led model puts the auditor at the front instead.

The human-led model

What verification changes

Four of these five rows are unchanged

Set against the fully autonomous model, most of the engagement is the same engagement. Verification is a reading step, and it moves one thing.

The expert-verified model set against the fully autonomous one
Against the fully autonomous modelIn this model
What is tested The same. Eleven coverage classes on the six phases, against the scope you signed off. Verification changes the reading of the output, not the reach of the run.
What a finding carries The same reproducible exploit artefact — request, response and reproduction steps. It is the unit of evidence in every B-52 model.
Who reads it before you do A senior Security Brigade auditor, who verifies every finding in the set. That reading is the whole of the difference between this model and the unattended one.
What Security Brigade can sign The assessment, under its CERT-In empanelment. The auditor has been in the engagement, and that involvement is what produces the signature.
The same holds for the human-led model. It does not hold where nobody at Security Brigade acted after scope sign-off.
Where the approval gates sit Unchanged. Destructive or state-changing production actions, persistence and movement beyond the entry host, and anything touching live credentials or real customer data each wait on your written approval.

The expert-verified model set against the fully autonomous one

What is tested

In this model
The same. Eleven coverage classes on the six phases, against the scope you signed off. Verification changes the reading of the output, not the reach of the run.

What a finding carries

In this model
The same reproducible exploit artefact — request, response and reproduction steps. It is the unit of evidence in every B-52 model.

Who reads it before you do

In this model
A senior Security Brigade auditor, who verifies every finding in the set. That reading is the whole of the difference between this model and the unattended one.

What Security Brigade can sign

In this model
The assessment, under its CERT-In empanelment. The auditor has been in the engagement, and that involvement is what produces the signature.

The same holds for the human-led model. It does not hold where nobody at Security Brigade acted after scope sign-off.

Where the approval gates sit

In this model
Unchanged. Destructive or state-changing production actions, persistence and movement beyond the entry host, and anything touching live credentials or real customer data each wait on your written approval.