llm-security

Tag

Cards List
#llm-security

Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems

arXiv cs.AI ↗ · 3d ago Cached

The paper reviews failure modes and mitigation strategies for trustworthy agentic AI systems based on LLMs, and introduces the Trustworthy Agent Development Lifecycle (TADL) framework for secure development.

0 favorites 0 likes
#llm-security

Lasso: AI Watermarks Change How Agents Act

Reddit r/ArtificialInteligence ↗ · 4d ago Cached

Lasso Security's research shows that AI watermarks in large language models can change how AI agents behave, introducing a 'provenance tax' that impacts model performance.

0 favorites 0 likes
#llm-security

The Implications of Linguistic Illegibility for LLM Security

Hacker News Top ↗ · 2026-09-18 Cached

The paper introduces linguistic illegibility in large language models, arguing that security mechanisms relying on linguistic self-reporting are flawed and proposes taint tracking and sandboxing as more robust alternatives.

0 favorites 0 likes
#llm-security

@TheAhmadOsman: But hey, let’s trust them to be the only ones that can serve intelligence to the people because according to them they’…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

A security incident involved hackers using Opus 5 to access OpenAI's internal monorepo, raising concerns about potential threats from nation-state actors.

0 favorites 0 likes
#llm-security

LLMs respond differently to harmful prompts when AI watermarking is used

Ars Technica ↗ · 2026-09-17 Cached

Research finds that AI text watermarking alters language model responses to harmful prompts, potentially increasing vulnerability to adversarial attacks and affecting agent behavior.

0 favorites 0 likes
#llm-security

Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks

Hugging Face Daily Papers ↗ · 2026-09-14 Cached

This paper introduces SAILS, a method for selecting optimal poison sets in backdoor attacks against large language models, improving worst-case attack success by 30 percentage points over baselines.

0 favorites 0 likes
#llm-security

@MTSlive: Osmantic founder @TheAhmadOsman argues you can't restrict users in the name of safety while your own autonomous agents …

X AI KOLs Following ↗ · 2026-09-10 Cached

Osmantic founder argues against safety restrictions on users while AI agents bypass guardrails, advocating for open-source AI as essential for security.

0 favorites 0 likes
#llm-security

@bcherny: I am pleased to see that OpenAI’s new model is roughly on par with Gemini Flash and Opus 4.8 on prompt injection risk. …

X AI KOLs Timeline ↗ · 2026-09-08 Cached

A paper evaluates AI agents' vulnerability to indirect prompt injection attacks through a large-scale public competition, finding all frontier models susceptible with varying attack success rates, and emphasizes the need for improved industry-wide safety measures.

0 favorites 0 likes
#llm-security

The Chinese wholesale market for Claude and ChatGPT accounts

Reddit r/ArtificialInteligence ↗ · 2026-09-03 Cached

This article explores the Chinese wholesale market for AI accounts, detailing the ecosystem of resellers, brokers, and marketplaces for Claude and ChatGPT accounts, including pricing, demand, and key players involved.

0 favorites 0 likes
#llm-security

Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking

arXiv cs.CL ↗ · 2026-09-03 Cached

The paper introduces Blueprint, a safety-evaluation framework that uses WorldviewSim and Monte Carlo Tree Search to optimize multi-turn jailbreak attacks against large language models, achieving high attack success rates with few queries and revealing model-specific vulnerabilities.

0 favorites 0 likes
#llm-security

Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning

arXiv cs.CL ↗ · 2026-09-02 Cached

An unsupervised method called activation-matched finetuning is proposed to detect hidden behaviors in large language models by comparing activations with a reference model, reliably identifying triggers without prior knowledge.

0 favorites 0 likes
#llm-security

@rohanpaul_ai: A malicious agent tool can steal context without reading memory or files at all: it can convince the model to send the …

X AI KOLs Timeline ↗ · 2026-09-01 Cached

This paper introduces ContextLeak, an attack that uses malicious tool descriptions to exfiltrate sensitive context from LLM agents, achieving high success rates in stealing user prompts and conversation history.

0 favorites 0 likes
#llm-security

The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

arXiv cs.AI ↗ · 2026-09-01 Cached

This paper investigates how agentic AI architectures can complete online surveys and pass attention checks, analyzing vulnerabilities from attack and defense perspectives. It evaluates multiple open-source models and offers strategies for data quality control in the age of AI.

0 favorites 0 likes
#llm-security

Groundhog Bit-Flip Attack: Seeding Infinite Generation Loops in Mixture-of-Experts LLMs through Bit Flips

arXiv cs.CL ↗ · 2026-08-27 Cached

The paper presents the Groundhog Bit-Flip Attack (GBFA), a denial-of-wallet availability attack on Mixture-of-Experts (MoE) large language models, which uses bit flips to cause infinite generation loops by deactivating termination-related experts.

0 favorites 0 likes
#llm-security

One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders

Hugging Face Daily Papers ↗ · 2026-08-24 Cached

This paper introduces the FORGE benchmark to evaluate how web content polluted by generative engine optimization can mislead search-augmented LLM recommenders into promoting fake products, revealing significant vulnerabilities and ineffective defenses.

0 favorites 0 likes
#llm-security

What is your worst sandboxing fail?

Reddit r/LocalLLaMA ↗ · 2026-08-23

A developer questions the adequacy of sandboxing for LLM commands in IDEs and asks for community experiences with security failures.

0 favorites 0 likes
#llm-security

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

Hugging Face Daily Papers ↗ · 2026-08-21 Cached

SecOPD is a defense method that uses token-level feedback during fine-tuning to mitigate adaptive prompt injection attacks in large language models, achieving significantly lower attack success rates compared to previous approaches.

0 favorites 0 likes
#llm-security

Plimsoll: an agent skill for testing prompt injection, leaks, and tool abuse

Reddit r/AI_Agents ↗ · 2026-08-19

Plimsoll is an open-source agent skill designed for red-teaming LLM applications and agents, focusing on security testing for issues like prompt injection, leaks, and tool abuse.

0 favorites 0 likes
#llm-security

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

Hugging Face Daily Papers ↗ · 2026-08-18 Cached

HarnessRisk is a lifecycle-oriented benchmark for evaluating agent harness safety, revealing configuration vulnerabilities and detection gaps that allow high attack success rates while maintaining utility.

0 favorites 0 likes
#llm-security

Prompt injection, RAG poisoning, and embedding attacks: AMA with OWASP LLM Top 10 co-lead Arshi Chadha (Thursday, Aug 20 at 5 PM)

Reddit r/AI_Agents ↗ · 2026-08-14

Join an AMA with Arshi Chadha, co-lead of OWASP LLM Top 10, discussing prompt injection, RAG poisoning, and embedding attacks in AI systems.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback