Loading...

AI-assisted penetration testing: a workflow, not a tool

Someone on the team wants to point an agent at a client's estate, and the answer has to be yes or no.

AI-assisted penetration testing means using AI inside a test that still has an authorized scope and a defined methodology. Nobody argues with that definition. Two completely different things are sold under it, and which one is on the table decides whether the result can be defended afterwards.

One is a tool that runs the test. The other is a way of working that structures how a person runs it. I built the second kind and took it through a client engagement end to end, scope to delivered report — what follows is what that turned out to require.

Two things are being sold as one

Search for agentic penetration testing and two categories come back wearing one name.

The first is a product. Autonomous agents map the attack surface, chain exploits, and hand back findings, and the pitch is scale — the honest version of which is that human testing does not scale. The second is not a product at all. A pentester uses agents inside a process they still own, and the process, not the model, is what makes the output defensible.

Both get called the same thing. They optimize for opposite outcomes. One works to take the tester out of the loop; the other works to make what the tester did inspectable afterwards. That is the difference between buying capability and building a process that can be audited, and only one of them belongs to the person who signs the deliverable.

A sealed capsule dropping finished shapes from a chute, beside an open scaffold with a lit node at its centre

The market has already chosen, and it chose recently.

In Cobalt's 2026 survey of 455 security professionals, the share of organizations relying entirely on AI automation for their testing fell from 29% to 9% in a single year. Some 47% now prefer a hybrid model, one in which human expertise supports the AI. A drop like that is not drift. It is a category losing its argument in public, on a vendor's own data.

So the question is not whether to allow AI into a penetration test. It is which of the two things is on offer, and what you would need to see before trusting it. Readers still placing penetration testing among the other kinds of security test should start there; this one assumes the function is already running.

What you are actually risking

The fear people voice is that an agent will invent a vulnerability. The risk with evidence behind it runs the other way.

In the same survey, 78% said fully automated scanning had missed critical vulnerabilities and return false negatives. Those two failures are not symmetrical. A fabricated finding is embarrassing, and it gets caught the moment someone tries to reproduce it. A missed one ships. The report says the application was tested, the client files it, and nobody learns otherwise until somebody else finds the thing.

That asymmetry decides where your scrutiny belongs. Reviewing what an agent claims to have found is the easy half. Establishing what it never looked at is the half that protects you.

There is a second risk, and I ran into it on my own engagement. Offensive work with a language model needs a documented authorization context and the right terms with the vendor; without them the model declines the task outright. It reads as friction the first time: a scope exists, a signed authorization exists, and the tool will not proceed.

Then the constraint shows what it is doing. A model that refuses to act until the authorization is in front of it enforces what the methodology already claimed to require. Scope and rules of engagement stop being paperwork finished before the real work starts and become a precondition the machinery will not let anyone skip.

An agent session on a lab engagement, with the model provider's safeguards flagging the offensive-security task mid-run
The provider's safeguards, not the framework's gates — offensive work trips them even inside an authorized lab engagement

Who is accountable when the agent is wrong

Accountability is not a policy sentence. It is a property of where the decisions sit.

In the framework I built, findings carry one of three tiers — Verified, Potential, and Potential-Unverifiable — and one rule governs all of them: a finding is never Verified on the model's say-so. The agent may propose. It may assemble evidence, reproduce steps, write the case. It may not be the thing that decides the case is proven. That belongs to a person, every time.

The same principle is wired in at the edges. A readiness gate at scope refuses to let an engagement start before authorization and rules of engagement are recorded; a hard gate before the report refuses to let it finish while coverage is unproven. Neither is advisory.

Contrast that with reviewing the output once it exists. Approval at the end is review, and review scales badly against volume: give a reviewer forty findings and a deadline and they will spot-check. A rule that no agent may promote a finding is not review. It holds whether or not anyone is paying attention that afternoon.

The difference is commercial too. Platforms in this market are sold on autonomy — agents that execute the test and hand back findings — and the buyer gets the outcome, not the rule that produced it. A constraint nobody can read before buying is a promise, not a control. The contracts here are public: twenty skills and nine contracts, open to inspection rather than description. Read the rule, disagree with it, fork it — that option is the difference.

How you know what it did not test

Coverage separates a defensible engagement from a plausible one. A report answers it by implication: here is what was found, therefore here is what was examined. The inference does not hold.

Here coverage is a computed property of a file rather than a sentence in a document. Test objectives are recorded as a living ledger during threat modelling, and every dynamic test case writes its verdict back against the objective it was meant to satisfy, so coverage is proven from the ledger rather than asserted in prose. An objective with no test case attached shows up as a row that never closed — not as an absence nobody noticed.

Engagement folder expanded in a file tree, one directory per stage, the coverage ledger among the files it leaves behind
A lab engagement, not the client run — every stage writes a file, and coverage is one of them

That property does something recognisable from long before anyone mentioned AI. A fixed set of steps and artifacts makes a forgotten test case visible, and makes a superficially executed test harder to leave that way. The failure it closes is mundane: a tester glances at the code, then at one role, then at another, finishes none of them, and after a weekend has no record that anything was left open.

None of that requires an agent to go wrong. The structure is not there to supervise a model — it is there because human attention is interruptible and a ledger is not.

That it also makes an agent safe to involve is a real benefit, and it is the second one. On a scope with several roles and a broad API attack surface, the distance between a ledger and a memory is the distance between knowing and hoping.

What a report has to carry to survive an audit

Start with the concession, because the opposite claim is easy to disprove and gets made constantly.

Platforms in this market do ship evidence: proof of exploitability, with step-by-step reproduction guides and video demonstrations of the chain. An article arguing that the industry ignores evidence would be wrong on its first page.

The narrow point survives the concession. That evidence lives inside a platform, produced by a process the buyer does not define, cannot read, and does not keep when the contract ends. It answers did this vendor do the work. It does not answer the question a lead actually has, which is what must my own process produce before agent-assisted work reaches a client.

The gap shows in how the market talks about speed. A widely circulated figure for AI-driven acceleration in penetration testing — a 35% reduction in average time-to-report — appears in a vendor's 2026 round-up of AI in pentesting, attributed to a well-known offensive security firm and carrying no link to anything that firm published. I could not find it anywhere on that firm's own site either. It may well be real. Nobody outside can check, which is the problem with aggregate numbers assembled by people selling the thing being measured. The line is not vendor or not-vendor — it is whether the number can be traced. A named report with a stated sample, read back by people who were not paid to like it, can be argued with. An attribution with nothing behind it cannot.

An engagement that has to survive scrutiny needs what auditors ask for: the standards it was bound by, findings a third party can reproduce, a trail of what was attempted and when, and a stated methodology. None of that is new, and none of it is what AI changed. What AI changed is how much of it arrives as a by-product of the work instead of being assembled afterwards from memory.

Whether two of your testers now agree

Ask a security lead what makes two engagements differ and the answer is usually scope. The uncomfortable answer is the tester.

In 2021, three researchers commissioned two independent providers to test the same IT environment and compared what came back. There was some overlap. There was also enough divergence to conclude that the human component has a profound impact on the outcome of a penetration test. Two providers is a small experiment: it shows the variance exists and matters rather than measuring how common it is. It also appears not to have been repeated.

In my experience, execution differs between individual testers — craft rather than process, each person working a scope slightly differently. Craft is not a criticism. It is what a client is buying at the top end, and what makes a result impossible to promise at the median.

Structure changes what craft is spent on. Nobody has run the same scope twice with two operators, and that experiment is not practically available to anyone in this field — a limit on the evidence, not a reason to pretend otherwise. My position, argued from the design and one engagement rather than from a controlled result, is that this raises quality rather than merely protecting it. Fix the sequence of steps and the artifacts each one produces, and the variable part shrinks to judgement — where a good tester is genuinely worth the money — while the mechanical part stops depending on who is holding it.

Whether it is faster, and what the speed costs

Every security leader I talk to is stuck on the same thing. Not whether AI can help — they assume it can — but whether anyone can show a team working this way going several times faster without quality falling over. The demos do not answer it. Neither do the case studies.

So: the number, with its limits attached.

The engagement I ran took roughly 30% of the time the same scope would previously have cost me — not thirty per cent faster, but finished in under a third of the time. That figure is my own before-and-after, not a study. I came out of it with a strong sense that nothing had been skipped, and a sense is exactly what that is: no coverage comparison against a conventional run exists.

Why the number is not more precise is the interesting part. A clean measurement would need the same targets tested twice, once with the framework and once without, by testers of comparable skill. Nobody is going to fund that, so nobody in this field can honestly produce a figure like "4.7× faster". A vendor quoting a precise multiplier is quoting something assembled from engagements that differed in scope, staffing and target — which is how a 35% average ends up in circulation with no traceable source.

An estimate that names its own method is worth more than a precise number that hides one.

The speed came with a bill nobody advertises. Tooling the runbook called for was missing from the machine and had to be installed mid-engagement. One tool would not run in that environment at all: it went on a separate system, with its output carried back into the pipeline by hand. The workflow tolerated that — the artifact model does not care whether a human or an agent produced the evidence, as long as it lands in the right place — but the hours went somewhere, and a first engagement should budget for them.

The part nobody mentions: what goes in the context window

There is a mechanical reason structure helps a model, and it is not the one usually given.

A 2023 study of how language models use long contexts found performance often highest when the relevant information sits at the beginning or the end of the input. Performance degraded significantly when the model had to retrieve it from the middle of a long context, including in models built for long contexts. That study tested the models of its own generation, not today's. The design lesson outlives them: what occupies the window is a decision, and leaving that decision to accumulation is still a decision.

A long row of panels going dark across its middle, beside four short stacks each lit end to end

A pentest is a long piece of work. Run as one conversation, the scope agreed on Monday sits tens of thousands of tokens behind the exploit attempt on Thursday. Run as phases that read and write files, each stage loads what that stage needs — the scope when scoping, the attack surface when planning, the evidence when writing findings — and nothing else. The material is not summarised into a thinner and thinner recollection. It is on disk, in full, and pulled back when it matters again.

Cost follows the same shape. Working this way consumes fewer tokens than the impulsive alternative — what I have taken to calling vibe-hacking, where an operator chats at an agent until something happens and every message drags the whole history along. I have not measured that difference and will not quote a figure for it.

What to require before allowing any of this

The honest summary of all of this is a checklist, and it belongs to whoever signs the deliverable. If that is you, four things are worth insisting on.

Require that the engagement leaves a folder, not a conversation: one directory per engagement, one file per finding, artifacts written by each stage and readable without the tool that produced them. Require that coverage is provable from a ledger rather than implied by the findings list. Require a tier system where the strongest label is reserved for reproduced evidence and no agent may apply it — the rule matters far more than which tiers get chosen. Require a gate at the start that refuses to begin without recorded authorization, and a gate at the end that refuses to ship with coverage unproven.

Any process meeting those four conditions is defensible whether or not an agent touched it, and you can check all four from the outside — on your own process, or on the one a vendor is selling you. That is the point. None of this is AI governance; it is what a penetration test should have produced all along, made cheap enough to insist on because the agent generates the paperwork as a by-product of doing the work.

10x-Pentest is one implementation — twenty skills, nine contracts, Apache-2.0, public — and I ran a client engagement on it end to end. Take the four requirements and ignore the implementation if it suits you. They are the part that transfers.


18 min read
Share this post:

Related Posts

All posts

Get your three regular assessments for free now!

  • All available job profiles included
  • Start assessing your candidates' skills right away
  • No time restrictions - register now, use your free assessments later
Create free account
  • All available job profiles included
  • Start assessing your candidates' skills right away
  • No time restrictions - register now, use your free assessments later
Top Scroll to top