model-fidelity

Tag

Cards List
#model-fidelity

Response drift across frontier large language models

arXiv cs.CL · 4d ago Cached

A large-scale human evaluation of 10 frontier LLMs across 62 questions finds that all models exhibit response drift, with most converging to a 78-81% deviation ceiling, while two achieve lower deviation. Drift varies by domain and question, and automated metrics explain little of human judgments, highlighting the need for human evaluation.

0 favorites 0 likes
← Back to home

Submit Feedback