Tag
This paper examines the limitations of general LLM rankings, including benchmark saturation and commercial incentives, and advocates for task-specific evaluation methods to improve reliability.
A new preprint called TEPA treats memory validity as a first-class state, revoking outdated precedents when new evidence conflicts while keeping audit trails. It outperforms append-only and last-write-wins in a complete-reversal experiment, though results are not yet independently reproduced.
This paper audits four agent-safety benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo) across many models, showing that their scores are confounded by capability, metrics have artifacts, and rankings disagree across benchmarks, undermining interchangeable safety claims.
This paper argues that standard RLHF's scalarization of human preferences collapses multiple valid interpretations into a single target, mis-measuring alignment in culturally plural societies. Analyzing a Malaysian dataset, they find 79% of prompts have multiple majority-supported responses that single-winner aggregation discards.
Introduces Generative-Evaluative Agreement (GEA), a validity criterion for LLM-enabled adaptive assessments, and measures it on a two-stage adaptive test, finding that the model recovers about half the intended variance with systematic bias.