capability-safety-confound

Tag

Cards List
#capability-safety-confound

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

arXiv cs.AI ↗ · 2026-08-19 Cached

This paper evaluates automated safety benchmarks for small language models, finding high ambiguity in judgments that compromises reliability and reveals a capability-safety confound.

0 favorites 0 likes
← Back to home

Submit Feedback