Full Report
A new benchmark of eight leading AI systems shows that an intelligent model still needs a good context-rich harness.
Analysis Summary
# Research: Why the smartest LLMs are not-so-smart pen testers
## Metadata
- **Authors:** John P. Mello Jr. (Reporting on research by Ridge Security)
- **Institution:** Ridge Security (with commentary from Hadrian, Black Duck Software, and ReversingLabs)
- **Publication:** ReversingLabs Blog
- **Date:** October 2024 (Implicit via 2026 report references/current industry context)
## Abstract
This research evaluates the efficacy of Large Language Models (LLMs) in autonomous offensive security roles. By benchmarking eight leading AI systems, the study demonstrates that raw model intelligence is secondary to the "harness"—the surrounding architecture that manages tool orchestration, state tracking, and context. The findings suggest that while high-end models achieve better coverage, they do so at a prohibitive cost, and smaller models can perform effectively if provided with superior environmental context and execution frameworks.
## Research Objective
The study aims to determine whether the highest-performing LLMs (based on standard benchmarks) translate directly into the most effective autonomous penetration testers. It specifically addresses the relationship between model intelligence, coverage, cost-efficiency, and the necessity of specialized security harnesses.
## Methodology
### Approach
Researchers conducted the first public benchmark comparing multiple leading LLMs within an autonomous penetration-testing workflow. The study focused on a "systems problem" approach, evaluating how models handle the full attack lifecycle.
### Dataset/Environment
- **Test Bed:** 96 model-target tests conducted against intentionally vulnerable environments.
- **Workflow Stages:** Reconnaissance, hypothesis testing, payload adaptation, exploitation, and verification.
- **Models Tested:** Eight leading systems, including Grok 4.5, Claude 3 Opus, Gemini 1.5 Flash, and GPT-OSS-120B.
### Tools & Technologies
- **Ridge Security Offensive-Security Harness:** Used to manage execution and provide the models with access to security tools.
- **Agentic SOC Framework:** Utilizing binary analysis to provide threat context.
## Key Findings
### Primary Results
1. **Harness Over Model:** The effectiveness of an autonomous pen tester depends more on the system architecture (harness) than the underlying LLM.
2. **The "Refusal" Barrier:** Frontier models often stall or refuse to complete authorized exploitation steps due to internal safety alignments, even in controlled environments.
3. **Cost vs. Coverage Disparity:** There is a non-linear relationship between cost and performance. High-end models like Claude Opus cost ~$217 per run for 63% coverage, whereas Gemini Flash costs ~$5.42 for 52% coverage.
4. **Efficiency Metrics:** GPT-OSS-120B was identified as the most efficient, yielding the highest number of findings per million tokens.
### Supporting Evidence
- **Grok 4.5:** Highest coverage at 77%.
- **Claude Opus 4.6:** 63% coverage at $217 per run.
- **Gemini 3 Flash:** 52% coverage at $5.42 per run.
- **Statistical Metric:** GPT-OSS-120B produced 16.9 findings per million tokens at a cost of $2.32 per run.
### Novel Contributions
- **First Public LLM Pen-Testing Benchmark:** Establishes a baseline for comparing how different LLMs perform in multi-stage offensive workflows.
- **Introduction of the "System-First" Philosophy:** Shifts the focus from LLM leaderboard rankings to the integration of binary analysis and state management.
## Technical Details
The research emphasizes that autonomous agents fail in security not because of a lack of "logic," but due to a lack of **State Tracking** and **Context Awareness**. The "harness" acts as the agent's memory and execution layer, preventing the model from wasting tokens on "rediscovering" information. The "Agentic SOC" model specifically uses binary analysis to supply pre-processed context, allowing the LLM to focus on high-level strategy rather than low-level parsing.
## Practical Implications
### For Security Practitioners
- **Procurement Strategy:** Do not select an AI security tool based solely on the underlying LLM (e.g., "Powered by GPT-4"). Evaluate the tool's ability to handle tool orchestration and error recovery.
- **Cost Management:** Be wary of "token spiral." High-intelligence models may be too expensive for routine scanning; smaller models with better context are often more viable.
### For Defenders
- **Verification is Key:** AI-generated findings must be verified by the harness. A "clean" result from an LLM may simply mean the model stalled or refused to perform an exploit, not that the system is secure.
- **Contextual Efficiency:** Use binary analysis and existing vulnerability data to prime AI agents, reducing the computational cost of autonomous discovery.
### For Researchers
- **Safety Alignment Conflicts:** Further research is needed to bypass "helpful assistant" refusals in authorized security testing without compromising the model's fundamental safety guardrails.
## Limitations
- **Alignment Interference:** Results are skewed by model-specific safety filters that categorize authorized pen-testing as malicious activity.
- **Static vs. Dynamic:** The benchmark focuses on intentionally vulnerable environments which may not fully replicate the complexity of modern, hardened enterprise networks.
## Comparison to Prior Work
Unlike standard LLM benchmarks (like MMLU or HumanEval) which test isolated logic or coding, this research focuses on **Long-Horizon Tasks**. It moves beyond the "chatbot" evaluation to "agentic" evaluation, emphasizing repeatability and tool-use over prose generation.
## Real-world Applications
- **Continuous Automated Pen-Testing:** Using smaller, cheaper models (like Gemini Flash) for daily regression testing while reserving larger models for complex, quarterly assessments.
- **Agentic SOC:** Integrating AI agents into Security Operations Centers to handle the initial validation of alerts.
## Future Work
- **Advanced Math & Cryptography:** Exploring how mathematical techniques can verify the claims and actions of AI agents to ensure they haven't "hallucinated" a vulnerability.
- **Tokenomics Optimization:** Developing specialized harnesses that minimize token usage by providing richer initial context.
## References
- Ridge Security Benchmark Study (2024)
- Gartner® Magic Quadrant™ for Software Supply Chain Security
- ReversingLabs Spectra Assure Community Analysis
- *Related Research:* Agentic SOC Alliance Framework (defanged: hxxps://www[.]reversinglabs[.]com/blog/smart-llm-pen-testing-fail)