Tag
The paper examines how clinical multi-agent systems can be compromised by socially plausible shortcuts and finds that independent referee oversight is essential for detecting such gaming.
Discusses how Goodhart's Law undermines trust in AI benchmarks, as metrics become targets and lose their validity as evaluation measures.