Tag
A large-scale human evaluation of 10 frontier LLMs across 62 questions finds that all models exhibit response drift, with most converging to a 78-81% deviation ceiling, while two achieve lower deviation. Drift varies by domain and question, and automated metrics explain little of human judgments, highlighting the need for human evaluation.