BitterSecurity's Decepticon Hits 98% on Hacking Test
BitterSecurity's open-source Decepticon framework scored 98.08% on the XBOW benchmark, demonstrating that multi-agent systems can execute complex web security tests inside isolated environments.

Cybersecurity firm BitterSecurity has released Decepticon, an open-source autonomous red-teaming framework that achieved a 98.08% pass rate on the XBOW validation suite. The tool successfully resolved 102 out of 104 challenges across easy, medium, and hard difficulty tiers in the public web-exploitation benchmark. Licensed under Apache-2.0, the project has already gathered over 5,700 stars on GitHub as interest grows in automated penetration testing.
Decepticon relies on LangGraph to coordinate stateful, multi-agent workflows. Its architecture deploys 16 specialized agents tasked with distinct operational phases, including reconnaissance, exploitation, post-exploitation, Active Directory navigation, cloud assessments, reverse engineering, and smart contract auditing. To maintain continuous interactive shell state, the system executes command sequences inside persistent tmux sessions within an isolated Kali Linux sandbox, supporting interactive security tools such as msfconsole and sliver.
The framework offers flexible model routing across multiple AI providers, including Anthropic, OpenAI, Gemini, DeepSeek, and local setups via Ollama. It also supports OAuth integration for subscription tiers like Claude Max and ChatGPT Pro. Developers and security teams can install the software using a single-line script or via pip install decepticon. Furthermore, Model Context Protocol (MCP) support allows integration directly into developer environments like Claude Code and Codex.
For security practitioners, Decepticon addresses long-standing hurdles in AI-driven red teaming, such as state loss during interactive command execution and rigid tool orchestration. While the current benchmark results are limited to web vulnerabilities, the platform provides a flexible, modular foundation for automated security assessments.
This is our own summary of reporting by AlphaSignal



