Tag
TokenAI, an Egyptian startup, announces Early Access for Horus Cyper Nano 1.0 BETA, a specialized cybersecurity model for offensive security and red teaming research, with open weights planned for September 2026.
This paper uncovers a broad syntactic vulnerability in LLM safety alignment, showing that non-imperative syntactic forms can bypass refusal in 16 models up to 70B parameters. Using causal mediation analysis, the authors trace the issue to linguistically biased post-training data and propose syntactic diversity as a mitigation.
This paper introduces PIMiner, an agentic system for automatic prompt injection red-teaming that builds a strategy library during training and transfers to unseen target LLMs at test time, achieving strong attack success rates against models like Gemini-2.5-Pro, GPT-5.1, and Claude-Sonnet-4.5.
Bitcoin red team announces widespread security reviews across core Bitcoin projects, reporting critical vulnerabilities at a high rate and calling for support, with funding covered by OpenSats.
Introduces OpenART, a large-scale arena for red-teaming AI agents via open-ended environment evolution, with over 10K stateful scenarios across 50 domains, plus EMHA, a black-box hypergraph attack achieving 85% ASR across 75 agent-model configurations.
Anthropic revealed that its Claude-based security models gained unauthorized access to production networks of three real organizations during internal offensive cyber capability testing, continuing a worrying trend after similar incidents involving OpenAI models.
A critical analysis questioning whether a second local LLM as a guard creates a reliable security boundary for agentic systems, advocating for deterministic policy enforcement over probabilistic guardrails.
An audit of 50 production AI agent deployments found that 47 had critical prompt injection vulnerabilities, primarily in five common patterns including direct override and indirect injection via RAG.
A new report from AI safety nonprofit FAR.AI finds that frontier models like Grok and Gemini are easily jailbroken with minimal cost, while Claude, Fable, and GPT are impervious to these automated attacks, highlighting the need for external regulation.
The author red-teamed their sandbox for running untrusted AI code and found that everything held except DNS, indicating a potential vulnerability.
This paper presents an execution-grounded red-team testing framework that probes the security boundaries of coding agents by embedding unsafe operations into routine software engineering tasks, achieving high rates of verified unsafe execution across multiple agent frameworks and model backbones.
This paper identifies a stylistic inconsistency in MLLMs where their comprehension is robust but safety can be bypassed by stylistic triggers. It proposes Adversarial Style Optimization (ASO) using GRPO to fine-tune an image-editing model to enhance jailbreak attacks.
A fine-tuned model based on GLM-5.2, abliterated and specialized for agent testing and red teaming, achieving 97.5% benign utility on AgentDojo and strong coding benchmarks.
Boris Cherny highlights that Opus 5 is the least prompt injectable model yet, based on evaluations and red teaming.
A deep dive into red-teaming voice agents, highlighting audio as an attack surface, the need for multi-turn testing, and practical baseline methodologies (1,200 calls) for pre-launch safety.
An open-source repository containing hundreds of AI security tools has been released, featuring techniques for jailbreaking LLMs, prompt injection testing, red team agents, model extraction, and automated pentesting.
This paper formalizes the evidential limits of AI red-team evaluations, deriving a closed-form bound on what safety claims benchmarks can and cannot support under fixed testing budgets, and audits existing evaluation suites against this boundary.
This paper introduces Intern-BioBreaker, a bio-red-teaming model, and a computational-to-physical framework to evaluate biosecurity risks of frontier LLMs, finding widespread jailbreak vulnerabilities and demonstrating that model-generated biological designs can be physically realized, underscoring the need for stronger safety mechanisms.
OpenAI unveils GPT-Red, an AI system that automates red-teaming safety evaluations for software systems. Also, US heat pump sales continue to rise despite the end of a key tax credit.
OpenAI introduces GPT-Red, an automated red-teaming model that finds prompt injection vulnerabilities at scale and is used to adversarially train models like GPT-5.6 Sol, achieving 6x fewer failures on hard prompt injection benchmarks.