general-reasoning

Tag

Cards List
#general-reasoning

BenchMIRT: What are LLM benchmarks actually measuring?

Hugging Face Blog · 2026-09-01 Cached

BenchMIRT introduces a method to audit LLM benchmarks at the individual prompt level using multidimensional item response theory, separating underlying capabilities like safety and general reasoning to reveal what benchmarks actually measure.

0 favorites 0 likes
#general-reasoning

@rohanpaul_ai: Claude Sonnet 5 upgrades are not uniform across every skill. e.g. its weaker than Sonnet 4.6 on CyberGym Here, CyberGym…

X AI KOLs Timeline · 2026-06-30 Cached

Claude Sonnet 5's upgrades are non-uniform; it underperforms Sonnet 4.6 on CyberGym vulnerability tasks because it wasn't deliberately trained for cyber tasks, relying instead on general reasoning. Anthropic's system card confirms this, while noting Sonnet 5's low pricing until August.

0 favorites 0 likes
← Back to home

Submit Feedback