adversarial-detection

Tag

Cards List
#adversarial-detection

How are you detecting new prompt injection patterns after launch?

Reddit r/AI_Agents · 4d ago

The article discusses methods for detecting new prompt injection patterns in AI systems after launch, including semantic search, trace-level safety scores, and tools like Braintrust, while highlighting challenges with false positives and attack taxonomy.

0 favorites 0 likes
#adversarial-detection

We implemented a second-order early warning signal for multi-turn prompt injection based on information geometry

Reddit r/artificial · 2026-07-09

A second-order early warning signal for multi-turn prompt injection is introduced, based on information geometry and a statistical manifold. The method uses a meta rate derived from the second derivative of the stability parameter to predict adversarial trajectories before threshold crossing, providing proactive detection.

0 favorites 0 likes
#adversarial-detection

USAD: Uncertainty-aware Statistical Adversarial Detection

arXiv cs.LG · 2026-06-29 Cached

USAD proposes two new statistics, Variance Discrepancy and Perturbation-based Covariance Discrepancy, to capture global and local uncertainty patterns of adversarial examples, achieving superior detection performance over baseline methods.

0 favorites 0 likes
#adversarial-detection

If your AI agent can send emails, browse websites, or call tools, I want to test something with you

Reddit r/artificial · 2026-06-02

Arc Gate is a security tool for AI agents that tracks entire conversations to detect adversarial behavioral drift across multiple turns, unlike traditional per-message checks. The author seeks teams with real agent workflows to test it.

0 favorites 0 likes
#adversarial-detection

Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning

arXiv cs.LG · 2026-05-26 Cached

Proposes Agent-ToM, a learning-to-monitor framework using Theory-of-Mind reasoning to detect covert malicious behavior in autonomous LLM agents by inferring beliefs and intents, outperforming baseline monitors.

0 favorites 0 likes
← Back to home

Submit Feedback