Tag
HARDEN introduces a constrained evolutionary search to generate more challenging evaluation cases for language models while preserving expected outputs, demonstrating significant accuracy reductions across benchmarks.
Researchers propose an adversarial hacker-fixer loop using LLM agents to automatically patch brittle verifiers in agent benchmarks, reducing attack success rates from 62% to 0% on KernelBench and demonstrating that weaker defenders can neutralize much stronger attackers.