Tag
BenchMIRT introduces a method to audit LLM benchmarks at the individual prompt level using multidimensional item response theory, separating underlying capabilities like safety and general reasoning to reveal what benchmarks actually measure.
Claude Sonnet 5's upgrades are non-uniform; it underperforms Sonnet 4.6 on CyberGym vulnerability tasks because it wasn't deliberately trained for cyber tasks, relying instead on general reasoning. Anthropic's system card confirms this, while noting Sonnet 5's low pricing until August.