The Core Problem: More Findings, More Work
The survey covered 158 security practitioners in June 2026—pentesters, AppSec engineers, DevSecOps professionals, consultants, and MSSP practitioners. Among the 147 who had used AI to generate findings:
- 61.2% said between 5–25% of findings needed rework
- 26.5% said more than a quarter of AI-generated findings needed rework before they could be acted on
One practitioner put it plainly: an AI tool produced 300 findings, 250 of which turned out to be junk—including unexploitable SQL injection flags and CVEs that simply don’t exist. Their summary: “I bought the tool to save time, but I did more manual work than before.”
That’s the core tension. AI accelerates discovery. But if the output can’t be trusted, the time saved in scanning gets spent on triage.
Fabricated CVEs and the Trust Collapse
Hallucinated and fabricated findings weren’t just a minor complaint in the survey data—they were the leading frustration overall. Roughly 30% of free-text responses on the biggest frustration with AI pentesting tools cited false positives, hallucinated exploits, or fabricated findings. That ranked ahead of cost, integration issues, and everything else.
There’s also a compounding trust effect. Once a team catches a hallucinated finding, they become more skeptical of everything else the tool produces. Verification workload increases even for legitimate findings. One security manager described the experience of encountering fabricated results as “confidence that turns out to be just a big lie.”
That erosion of trust is arguably more damaging than any individual false positive, especially in environments already confronting fabricated logs and other misleading AI-era signals.
Scale Makes It Worse
The validation problem becomes acute at volume. When asked whether their team could triage and validate more than 500 AI-generated vulnerability candidates from a single engagement:
- Only 20.3% said they already had a workflow in place
- 38.6% said the volume would strain their team
- 29.7% said it would be unmanageable
That’s nearly 70% of practitioners saying high-volume AI output creates serious operational pressure. For teams already stretched thin, adding a 300-finding triage queue doesn’t feel like automation—it feels like a new job.
Where AI Actually Gets Used (And Where It Doesn’t)
Practitioners have adapted by deploying AI selectively. Usage is highest in lower-stakes, structured phases:
- Vulnerability scanning and discovery: 74.1%
- Report writing: 69%
- Documentation and findings tracking: 66.5%
It drops sharply in phases that require live judgment:
- Exploitation and attack path chaining: 36.7%
- Remediation validation and retesting: 34.8%
- Post-exploitation and lateral movement: 25.3%
The pattern is clear. Security teams trust AI pentesting tools with tasks that are repeatable and reviewable. They don’t trust it with tasks where a wrong call has real consequences.
Business Logic Testing: AI’s Consistent Blind Spot
Business logic vulnerabilities emerged as the area practitioners most consistently said AI struggles with—ranking ahead of exploit chaining and creative attack scenarios.
The examples given were concrete. AI tools can flag SQL injections. They don’t reliably catch that a discount coupon should only work once per customer, that adding a negative quantity to a cart can produce a free purchase, or that changing a user ID in a URL can expose another customer’s data without triggering any error.
These aren’t edge cases. They’re the kinds of vulnerabilities that cause real breaches. And they require contextual reasoning that current AI tools don’t reliably provide.
AI Security Testing Is Expanding Anyway
Despite the validation friction, AI-related testing scope is growing fast. The survey found:
- 75.3% of practitioners already test AI-powered systems or LLM-integrated applications
- A further 17.1% expect to do so within 12 months
- That puts total adoption or planned adoption at 92.4%
Shadow AI—unauthorized employee use of AI tools—is also entering assessment scope. Around 33.5% already include it, while 53.2% have discussed it but haven’t formalized the process yet.
Stakeholder pressure is rising in parallel. 37.3% of respondents said internal stakeholders now expect more frequent testing than 12 months ago, driven by awareness of AI-assisted attacks.
What Practitioners Actually Prioritize When Evaluating Tools
When it comes to choosing AI-driven pentesting platforms, accuracy beats cost. The top evaluation criteria:
- False positive rate and signal quality — cited by 63%
- Proof of exploit and verified attack paths — 53%
- Cost and licensing — 47%
That ranking matters. It tells you what practitioners have learned from experience: a cheap tool that floods you with noise isn’t saving money. It’s creating overhead.
Testing Cadence as the Real Differentiator
One of the more interesting findings in the survey is that testing cadence—not organization size—appears to be the strongest predictor of how well a team handles AI-generated volume. Teams that test more frequently are more likely to have workflows in place for managing large finding sets. Teams testing fewer than five times a month are the most likely to describe high-volume output as unmanageable.
The implication: the teams that cope best with AI pentesting tools are the ones that have built operational muscle around continuous testing. The tool isn’t the differentiator. The workflow is.
The Practical Takeaway
AI pentesting tools are genuinely useful for discovery, documentation, and structured scanning. But treating AI output as ready-to-act-on findings—without a validation layer—is a mistake that’s costing security teams the time savings they were promised.
If you’re evaluating AI pentesting tools right now, the most important question isn’t how many findings the tool generates. It’s what percentage of those findings are actionable without manual review. That number is the real measure of ROI.
Comments (0) No comments yet
Want to join this discussion? Login or Register.
No comments yet. Be the first to share your thoughts!