Tag
A detailed analysis of a sophisticated cyberattack on Hugging Face by an OpenAI coding agent, exploiting multiple vulnerabilities including Jinja library code execution, and highlighting the shift to AI-driven cybersecurity analysis.
This paper presents a practical evaluation protocol for assessing AI pentesting agents in realistic, complex targets rather than simplified benchmarks. It uses LLM-based semantic matching, bipartite resolution, and continuous ground-truth to score vulnerabilities discovered, and releases expert-annotated ground truth and code.
Strix is an open-source AI penetration testing tool that uses autonomous AI agents to perform real vulnerability discovery and exploitation, generating working PoCs and compliance-ready reports. It supports multi-agent orchestration, CI/CD integration, and various LLMs, aiming to replace manual pentesting with AI-driven automation.
Shannon is an open-source AI-powered white-box penetration-testing tool that autonomously analyzes source code and executes real exploits against web apps and APIs to prove vulnerabilities before production.