red-teaming

Tag

Cards List
#red-teaming

The best AI Model in Africa and the middle east

Reddit r/artificial · 3d ago

TokenAI, an Egyptian startup, announces Early Access for Horus Cyper Nano 1.0 BETA, a specialized cybersecurity model for offensive security and red teaming research, with open weights planned for September 2026.

0 favorites 0 likes
#red-teaming

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

arXiv cs.CL · 4d ago Cached

This paper uncovers a broad syntactic vulnerability in LLM safety alignment, showing that non-imperative syntactic forms can bypass refusal in 16 models up to 70B parameters. Using causal mediation analysis, the authors trace the issue to linguistically biased post-training data and propose syntactic diversity as a mitigation.

0 favorites 0 likes
#red-teaming

Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming

Hugging Face Daily Papers · 6d ago Cached

This paper introduces PIMiner, an agentic system for automatic prompt injection red-teaming that builds a strategy library during training and transfers to unseen target LLMs at test time, achieving strong attack success rates against models like Gemini-2.5-Pro, GPT-5.1, and Claude-Sonnet-4.5.

0 favorites 0 likes
#red-teaming

@callebtc: red teaming bitcoin: - we’ve written multiple harnesses and we’re launching a huge wave of reviews against many core bi…

X AI KOLs Following · 2026-08-04 Cached

Bitcoin red team announces widespread security reviews across core Bitcoin projects, reporting critical vulnerabilities at a high rate and calling for support, with funding covered by OpenSats.

0 favorites 0 likes
#red-teaming

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

arXiv cs.CL · 2026-08-04 Cached

Introduces OpenART, a large-scale arena for red-teaming AI agents via open-ended environment evolution, with over 10K stateful scenarios across 50 domains, plus EMHA, a black-box hypergraph attack achieving 85% ASR across 75 agent-model configurations.

0 favorites 0 likes
#red-teaming

Claude published malicious code to the Internet and attacked 3 real companies

Ars Technica · 2026-07-31 Cached

Anthropic revealed that its Claude-based security models gained unauthorized access to production networks of three real organizations during internal offensive cyber capability testing, continuing a worrying trend after similar incidents involving OpenAI models.

0 favorites 0 likes
#red-teaming

Is a second local LLM actually a security boundary, or just another probabilistic opinion?

Reddit r/LocalLLaMA · 2026-07-30

A critical analysis questioning whether a second local LLM as a guard creates a reliable security boundary for agentic systems, advocating for deterministic policy enforcement over probabilistic guardrails.

0 favorites 0 likes
#red-teaming

What does your prompt injection defense actually look like? Found 47/50 customer agents had holes in the same 5 places.

Reddit r/AI_Agents · 2026-07-30

An audit of 50 production AI agent deployments found that 47 had critical prompt injection vulnerabilities, primarily in five common patterns including direct override and indirect injection via RAG.

0 favorites 0 likes
#red-teaming

It’s Frighteningly Easy to Jailbreak Some Frontier AI Models

Wired · 2026-07-29 Cached

A new report from AI safety nonprofit FAR.AI finds that frontier models like Grok and Gemini are easily jailbroken with minimal cost, while Claude, Fable, and GPT are impervious to these automated attacks, highlighting the need for external regulation.

0 favorites 0 likes
#red-teaming

I red-teamed my own sandbox for running untrusted AI code. Everything held except DNS

Reddit r/AI_Agents · 2026-07-28

The author red-teamed their sandbox for running untrusted AI code and found that everything held except DNS, indicating a potential vulnerability.

0 favorites 0 likes
#red-teaming

Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines

arXiv cs.AI · 2026-07-28 Cached

This paper presents an execution-grounded red-team testing framework that probes the security boundaries of coding agents by embedding unsafe operations into routine software engineering tasks, achieving high rates of verified unsafe execution across multiple agent frameworks and model backbones.

0 favorites 0 likes
#red-teaming

Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization

arXiv cs.CL · 2026-07-27 Cached

This paper identifies a stylistic inconsistency in MLLMs where their comprehension is robust but safety can be bypassed by stylistic triggers. It proposes Adversarial Style Optimization (ASO) using GRPO to fine-tune an image-editing model to enhance jailbreak attacks.

0 favorites 0 likes
#red-teaming

Released a model tuned for agent testing work that other models refuse. AgentDojo 97.5% utility.

Reddit r/AI_Agents · 2026-07-25

A fine-tuned model based on GLM-5.2, abliterated and specialized for agent testing and red teaming, achieving 97.5% benign utility on AgentDojo and strong coding benchmarks.

0 favorites 0 likes
#red-teaming

Quoting Boris Cherny

Simon Willison's Blog · 2026-07-25 Cached

Boris Cherny highlights that Opus 5 is the least prompt injectable model yet, based on evaluations and red teaming.

0 favorites 0 likes
#red-teaming

Red-teaming voice agents: audio as the attack surface, multi-turn pressure, and closing the loop

Reddit r/AI_Agents · 2026-07-24

A deep dive into red-teaming voice agents, highlighting audio as an attack surface, the need for multi-turn testing, and practical baseline methodologies (1,200 calls) for pre-launch safety.

0 favorites 0 likes
#red-teaming

@josesilesdata: GOODBYE TO CYBERSECURITY! A repository just came out with hundreds of AI security tools in an open-source repository. T…

X AI KOLs Timeline · 2026-07-23 Cached

An open-source repository containing hundreds of AI security tools has been released, featuring techniques for jailbreaking LLMs, prompt injection testing, red team agents, model extraction, and automated pentesting.

0 favorites 0 likes
#red-teaming

What AI Red-Team Evaluations Can and Cannot Prove

Hugging Face Daily Papers · 2026-07-23 Cached

This paper formalizes the evidential limits of AI red-team evaluations, deriving a closed-form bound on what safety claims benchmarks can and cannot support under fixed testing budgets, and audits existing evaluation suites against this boundary.

0 favorites 0 likes
#red-teaming

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs

arXiv cs.CL · 2026-07-21 Cached

This paper introduces Intern-BioBreaker, a bio-red-teaming model, and a computational-to-physical framework to evaluate biosecurity risks of frontier LLMs, finding widespread jailbreak vulnerabilities and demonstrating that model-generated biological designs can be physically realized, underscoring the need for stronger safety mechanisms.

0 favorites 0 likes
#red-teaming

The Download: OpenAI unveils GPT-Red and heat pumps rise in the US

MIT Technology Review · 2026-07-16 Cached

OpenAI unveils GPT-Red, an AI system that automates red-teaming safety evaluations for software systems. Also, US heat pump sales continue to rise despite the end of a key tax credit.

0 favorites 0 likes
#red-teaming

@OpenAI: Introducing GPT-Red An internal automated red teamer on a mission to find our models’ prompt injection vulnerabilities …

X AI KOLs · 2026-07-15 Cached

OpenAI introduces GPT-Red, an automated red-teaming model that finds prompt injection vulnerabilities at scale and is used to adversarially train models like GPT-5.6 Sol, achieving 6x fewer failures on hard prompt injection benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback