Tag
This paper investigates weaknesses in emotion recognition in conversations (ERC) by analyzing LLM performance in zero-shot settings, revealing systematic failures due to annotation ambiguity, and proposes an LLM-as-Judge framework for more robust evaluation.
The article critiques the overreliance on benchmark scores in AI model evaluation and advocates for comprehensive assessment methods like execution traces and error analysis, with companies like Parsewave emphasizing deeper insights.
Highlights OpenAI researcher Noam Brown's argument: the true ceiling of LLM capabilities is far higher than current benchmarks show, due to insufficient test-time compute, and stronger models benefit more from additional computation. This poses a serious challenge for AI safety evaluation, as many dangerous capabilities may only emerge under long time and high compute budgets.
A free Stanford lecture by Percy Liang on AI generalization explains why models excel on benchmarks but fail on real codebases, covering benchmark memorization, bias-variance tradeoff, and hallucination.