Tag
This paper introduces Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using LLM agents, evaluating models across multiple safety metrics through adaptive adversarial interactions.
Multi2AV-Safety is the first benchmark for evaluating safety in multimodal-to-audio-Video generation, covering all 11 non-singleton conditioning configurations with 11,024 attack instances, and revealing compositional risks where harmful semantics emerge from benign inputs.
Introduces SurakshaEval, a safety benchmark for LLMs covering ten Indian languages and English, with human-written prompts spanning seven harm types. Benchmarks multilingual LLMs and finds issues like over-refusal and missed implicit bias in Indic contexts.
SafePyramid is a hierarchical benchmark with 1,000 multi-turn conversations across 10 domains and 3,000 policies to evaluate guardrail systems' ability to identify safety violations via in-context policy specification. Tests on 10 frontier LLMs show that even GPT-5.5 only correctly identifies all violated rules 54% of the time at the easiest level, highlighting the challenge of reliable in-context policy guardrailing.
This paper investigates whether LLMs can reliably self-report when their outputs have been compromised by adversarial prefills, finding that models often cannot distinguish between compromised and intentional outputs, and their limited recognition stems from normal refusal behavior rather than true self-awareness.