Tag
The article discusses a research paper revealing that many science-based LLM benchmarks have incorrect answers, and when corrected, LLM performance scores rise significantly, suggesting current evaluations may underestimate AI capabilities in physics.