safety-benchmark

Tag

Cards List
#safety-benchmark

BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

arXiv cs.AI ↗ · 2026-09-16 Cached

This paper introduces Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using LLM agents, evaluating models across multiple safety metrics through adaptive adversarial interactions.

0 favorites 0 likes
#safety-benchmark

Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation

arXiv cs.AI ↗ · 2026-08-28 Cached

Multi2AV-Safety is the first benchmark for evaluating safety in multimodal-to-audio-Video generation, covering all 11 non-singleton conditioning configurations with 11,024 attack instances, and revealing compositional risks where harmful semantics emerge from benign inputs.

0 favorites 0 likes
#safety-benchmark

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

arXiv cs.CL ↗ · 2026-08-11 Cached

Introduces SurakshaEval, a safety benchmark for LLMs covering ten Indian languages and English, with human-written prompts spanning seven harm types. Benchmarks multilingual LLMs and finds issues like over-refusal and missed implicit bias in Indic contexts.

0 favorites 0 likes
#safety-benchmark

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

Hugging Face Daily Papers ↗ · 2026-06-29 Cached

SafePyramid is a hierarchical benchmark with 1,000 multi-turn conversations across 10 domains and 3,000 policies to evaluate guardrail systems' ability to identify safety violations via in-context policy specification. Tests on 10 frontier LLMs show that even GPT-5.5 only correctly identifies all violated rules 54% of the time at the easiest level, highlighting the challenge of reliable in-context policy guardrailing.

0 favorites 0 likes
#safety-benchmark

@rohanpaul_ai: LLMs often cannot tell when an attack made them say something unsafe. Asking an LLM whether its own previous answer was…

X AI KOLs Timeline ↗ · 2026-06-24 Cached

This paper investigates whether LLMs can reliably self-report when their outputs have been compromised by adversarial prefills, finding that models often cannot distinguish between compromised and intentional outputs, and their limited recognition stems from normal refusal behavior rather than true self-awareness.

0 favorites 0 likes
← Back to home

Submit Feedback