CyberSecurity SEE

Harness Is More Important Than Model

Harness Is More Important Than Model

Ridge Security has recently made headlines with the release of a groundbreaking study that conducts the first public benchmark comparing eight leading large language models (LLMs) in the realm of autonomous penetration testing. This extensive evaluation aims to determine how well these diverse models can perform in the context of identifying vulnerabilities within systems. Contrary to what might be popularly believed, the findings reveal that the intricacies of system architecture are more influential than sheer model intelligence when it comes to effectiveness in security tasks.

The study conducted a total of 96 model-target tests against purposely crafted vulnerable environments. It meticulously tracked the ability of each model to complete full testing workflows, which included essential stages such as reconnaissance, hypothesis testing, payload adaptation, exploitation, and verification. The results unveiled a competitive landscape: Grok 4.5 stood out, achieving 77% coverage, while its counterpart, Claude Opus 4.6, reached a significant 63%, albeit with a cost of $217 per run. In contrast, Gemini 3 Flash showed a lower coverage rate of 52% at a relatively modest cost of $5.42, while GPT-OSS-120B demonstrated the most efficient performance at just $2.32 per run.

These findings challenge a long-held assumption within the industry that the highest-scoring LLM invariably makes for the best autonomous security agent. Lydia Zhang, President of Ridge Security, elaborated on this stance, highlighting that autonomous offensive security is predominantly characterized as a systems problem. Success in this area necessitates an architecture capable of managing task execution, accommodating new discoveries, and verifying findings.

Experts in the field echoed these sentiments, emphasizing that effective penetration testing relies significantly on the capability for state tracking, the dependable use of security tools, and the provision of proof for every finding. Notably, Li Zhao from Black Duck Software observed that a well-engineered agent, even if based on a smaller model, can outperform cutting-edge models when the surrounding architecture, tools, and workflows are finely tuned.

One of the critical challenges identified in the research revolves around frontier models, which sometimes refuse to carry out specific actions during authorized testing workflows, particularly in payload generation and exploitation tasks. Such refusals are typically a result of model alignment intended to prevent misuse, yet they can lead to operational setbacks. For instance, if a model halts its operations midway through an authorized test, this can create inadvertent gaps in the testing, which might be misinterpreted as clean results. This necessitates that teams disclose any skipped tests and rely on established testing tools to complete the work within an approved scope.

Gunter Ollmann from Cobalt Labs emphasized the necessity of ensuring that authorization for offensive security activities should not reside solely within the LLM. Instead, he argued, the surrounding system should dictate what actions are executed. This brings to the forefront the essential concept of an “agent harness.”

The benchmark highlights this agent harness as a vital element that distinctly separates reasoning from execution and verification processes. This harness effectively maintains the model’s reasoning capabilities scoped and safe through mechanisms like scope enforcement, memory management, deterministic tooling, validation, and the generation of audit trails. Without such a harness in place, offensive security activities risk devolving into disorganized prompts rather than structured workflows. Experts depict an unharnessed frontier model as akin to a brilliant intern with root access but lacking supervision, underscoring that the true value lies in the expertise encapsulated within the harness and its ability to prove findings before they reach stakeholders.

Furthermore, the significance of Ridge Security’s findings aligns seamlessly with the framework proposed by the Agentic SOC Alliance, which advocates for a three-layer model concerning autonomous security operations. This model consists of a context layer—comprised of structured knowledge regarding devices and their behaviors, a harness layer intended for orchestration and governance, and an interchangeable model layer. Incorporating binary analysis emerges as a crucial aspect, providing the necessary context that stops agents from wasting resources rediscovering previously known information.

In light of these findings, it becomes imperative for organizations to assess outcomes by prioritizing validated findings along with false-positive rates, repeatability, and cost per validated finding instead of merely chasing model rankings. As offensive AI tools are increasingly able to unearth vulnerabilities more swiftly, the ultimate measure of success will hinge on whether organizations can validate, prioritize, and effectively remediate what is uncovered.

In conclusion, Ridge Security’s research provides a compelling narrative about the critical elements that govern the efficacy of large language models in penetration testing. Moving forward, organizations must adopt a more nuanced approach that emphasizes the architecture and processes supporting AI tools rather than solely relying on the capabilities of individual models. This transformative perspective holds the key to enhancing the overall effectiveness of autonomous security operations.

Source link

Exit mobile version