CyberSecurity SEE

AI Pentesting Tools Outpace Security Teams in Findings, New Survey Reveals

New Study Reveals Challenges in AI-Assisted Pentesting: Validation Backlog Reshapes Security Practices

In a groundbreaking study conducted by Pentest-Tools.com, findings indicate that AI-assisted penetration testing tools are yielding vulnerability results at a pace that surpasses the verification capabilities of most security teams. This rapid output creates a validation backlog that nullifies the time-saving advantages initially anticipated from AI integration within the cybersecurity landscape.

The research, which surveyed 158 security practitioners in June 2026, included a diverse group of professionals such as penetration testers, security engineers, application security (AppSec) and DevSecOps professionals, consultants, and managed security service providers (MSSP). All participants reported using AI-assisted tools in their vulnerability assessments and validations.

Among the 147 respondents who had experience utilizing AI for generating findings, a staggering 87.8% indicated that the outputs required substantial manual validation. The data further illuminated the validation challenges, with 61.2% of participants noting that between 5% and 25% of the AI-generated findings necessitated reworking. Alarmingly, 26.5% reported that more than a quarter of the findings generated were unreliable and needed significant re-evaluation before they could be deemed trustworthy.

As the volume of AI-generated vulnerabilities increased, the perception of capacity issues became more pronounced. When asked if their team could effectively triage and validate upwards of 500 AI-generated vulnerability candidates from a single engagement, merely 20.3% reported having an established workflow for managing such a caseload. In contrast, 38.6% felt that the volume would stretch their resources thin, while 29.7% deemed it unmanageable altogether.

Participants voiced concerns that the automated tools were frequently returning findings that included duplicates, false positives, unexploitable issues, or even fabricated Common Vulnerabilities and Exposures (CVE) that do not exist. One practitioner recounted a case in which an AI tool produced 300 findings, of which 250 were identified as invalid, describing issues such as “potential SQL injection vulnerabilities that are not exploitable” and “AI-generated CVEs that do not exist.” This individual lamented, “I invested in the tool to save time, yet I wound up performing more manual work than I did previously.”

The prevalence of fabricated and hallucinated findings emerged as the primary frustration among survey respondents. Approximately 30% of the free-text responses highlighted false positives, fictitious exploits, or concocted findings as their predominant complaints. This trend has instigated a “trust effect”; once a false finding was detected, teams became increasingly cautious with the tool’s outputs, inadvertently increasing the workload for verifying genuine findings. A security manager at a mid-sized company articulated the disillusionment by stating, “It’s like confidence that turns out to be just a big lie.”

The survey further revealed that practitioners are employing AI most effectively during phases that require scanning and discovery (74.1%), report writing (69%), and documentation or tracking of findings (66.5%). However, the use of AI diminishes during phases demanding live judgement, such as exploitation and attack path chaining (36.7%), remediation validation and retesting (34.8%), and post-exploitation lateral movement (25.3%).

Business logic testing emerged as the area where AI tools confront substantial challenges, outpacing even exploit chaining and creativity. Respondents shared clear examples: while AI systems can identify SQL injections, they often fail to detect scenarios where a discount coupon applies only once per customer or potential loopholes that might allow negative quantities to produce free transactions.

Interestingly, the survey indicates that the scope of testing related to AI is expanding as well. An impressive 75.3% of practitioners stated that they are currently testing AI-powered systems or applications integrated with large language models (LLM) as part of their ongoing engagements. An additional 17.1% plan to include these assessments within the next year, culminating in a total adoption or planned adoption rate of 92.4%. Moreover, approximately one-third (33.5%) of professionals already take into account the risks associated with unauthorized employee use of AI, often referred to as “shadow AI,” while 53.2% acknowledged discussions regarding this issue, even if they have not yet formalized a process.

Increasing stakeholder pressures were also evident, as 37.3% of respondents noted that internal stakeholders now anticipate more frequent testing than they did just a year ago, largely due to heightened awareness of AI-assisted cyber threats. Meanwhile, 31.6% acknowledged that stakeholders are aware of the shifting risks but have yet to alter their purchasing behaviors accordingly.

When evaluating AI-driven penetration testing platforms, criteria related to accuracy took precedence over cost considerations. The false positive rate and signal quality emerged as the top priorities for 63% of participants, closely followed by proof of exploit and verified attack paths at 53%. Cost and licensing factors were ranked third at 47%.

A spokesperson for Pentest-Tools.com commented on the findings: “AI is accelerating the rate at which practitioners unearth vulnerabilities. However, the real challenge lies in the subsequent validation process. When faced with a total of 300 findings where 250 are deemed invalid, any time savings achieved during discovery gets consumed by the triage process.”

The survey’s findings suggest that testing cadence, rather than organizational size, serves as a more accurate predictor of how well teams manage the influx of AI-generated data. Teams that conduct testing more frequently are better equipped to handle larger volumes of findings, while those testing less than five times a month find the volume overwhelming. Nonetheless, Pentest-Tools.com cautions that the sample size for those in the highest-frequency testing groups is small, urging readers to interpret this finding as preliminary rather than definitive.

The comprehensive survey report titled “AI Pentesting in 2026: Why Testing Cadence Decides Who Copes” can be accessed here.

This exploration into the efficacy and challenges of AI in penetration testing underscores a pivotal moment in cybersecurity, revealing the intricacies involved in merging advanced technology with the human element of validation.

Source link

Exit mobile version