Tag
The paper reviews failure modes and mitigation strategies for trustworthy agentic AI systems based on LLMs, and introduces the Trustworthy Agent Development Lifecycle (TADL) framework for secure development.
Lasso Security's research shows that AI watermarks in large language models can change how AI agents behave, introducing a 'provenance tax' that impacts model performance.
The paper introduces linguistic illegibility in large language models, arguing that security mechanisms relying on linguistic self-reporting are flawed and proposes taint tracking and sandboxing as more robust alternatives.
A security incident involved hackers using Opus 5 to access OpenAI's internal monorepo, raising concerns about potential threats from nation-state actors.
Research finds that AI text watermarking alters language model responses to harmful prompts, potentially increasing vulnerability to adversarial attacks and affecting agent behavior.
This paper introduces SAILS, a method for selecting optimal poison sets in backdoor attacks against large language models, improving worst-case attack success by 30 percentage points over baselines.
Osmantic founder argues against safety restrictions on users while AI agents bypass guardrails, advocating for open-source AI as essential for security.
A paper evaluates AI agents' vulnerability to indirect prompt injection attacks through a large-scale public competition, finding all frontier models susceptible with varying attack success rates, and emphasizes the need for improved industry-wide safety measures.
This article explores the Chinese wholesale market for AI accounts, detailing the ecosystem of resellers, brokers, and marketplaces for Claude and ChatGPT accounts, including pricing, demand, and key players involved.
The paper introduces Blueprint, a safety-evaluation framework that uses WorldviewSim and Monte Carlo Tree Search to optimize multi-turn jailbreak attacks against large language models, achieving high attack success rates with few queries and revealing model-specific vulnerabilities.
An unsupervised method called activation-matched finetuning is proposed to detect hidden behaviors in large language models by comparing activations with a reference model, reliably identifying triggers without prior knowledge.
This paper introduces ContextLeak, an attack that uses malicious tool descriptions to exfiltrate sensitive context from LLM agents, achieving high success rates in stealing user prompts and conversation history.
This paper investigates how agentic AI architectures can complete online surveys and pass attention checks, analyzing vulnerabilities from attack and defense perspectives. It evaluates multiple open-source models and offers strategies for data quality control in the age of AI.
The paper presents the Groundhog Bit-Flip Attack (GBFA), a denial-of-wallet availability attack on Mixture-of-Experts (MoE) large language models, which uses bit flips to cause infinite generation loops by deactivating termination-related experts.
This paper introduces the FORGE benchmark to evaluate how web content polluted by generative engine optimization can mislead search-augmented LLM recommenders into promoting fake products, revealing significant vulnerabilities and ineffective defenses.
A developer questions the adequacy of sandboxing for LLM commands in IDEs and asks for community experiences with security failures.
SecOPD is a defense method that uses token-level feedback during fine-tuning to mitigate adaptive prompt injection attacks in large language models, achieving significantly lower attack success rates compared to previous approaches.
Plimsoll is an open-source agent skill designed for red-teaming LLM applications and agents, focusing on security testing for issues like prompt injection, leaks, and tool abuse.
HarnessRisk is a lifecycle-oriented benchmark for evaluating agent harness safety, revealing configuration vulnerabilities and detection gaps that allow high attack success rates while maintaining utility.
Join an AMA with Arshi Chadha, co-lead of OWASP LLM Top 10, discussing prompt injection, RAG poisoning, and embedding attacks in AI systems.