benchmark-hardening

Tag

Cards List
#benchmark-hardening

HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases

arXiv cs.AI ↗ · 4d ago Cached

HARDEN introduces a constrained evolutionary search to generate more challenging evaluation cases for language models while preserving expected outputs, demonstrating significant accuracy reductions across benchmarks.

0 favorites 0 likes
#benchmark-hardening

Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops

Hugging Face Daily Papers ↗ · 2026-06-08 Cached

Researchers propose an adversarial hacker-fixer loop using LLM agents to automatically patch brittle verifiers in agent benchmarks, reducing attack success rates from 62% to 0% on KernelBench and demonstrating that weaker defenders can neutralize much stronger attackers.

0 favorites 0 likes
← Back to home

Submit Feedback