Tag
This paper replicates a distributional-semantics extractive summarization method for Hindi and evaluates it on standard corpora, finding that sentence position is the only contributing feature and current Hindi benchmarks fail to incentivize advanced content selection.
This paper explores efficient recurring evaluation methods for production LLM agents, comparing techniques like adaptive testing and fixed subsets, and provides practical recommendations for deployment.
PersonalBench is a new benchmark that evaluates inference-time personalization methods in LLMs through authorship verification, LLM-as-judge, and stylometrics, finding that while methods produce author-differentiated output, they do not bridge the gap to human authorship.
The paper introduces THPT-Ladder, a benchmark that evaluates AI models using Vietnam's 2025 convex marking scheme, revealing how partial credit gaps lead to inaccurate assessments compared to standard accuracy metrics.
This paper geometrically analyzes why LLMs acting as judges agree strongly with each other but weakly with humans, finding that inter-LLM consensus reflects a collapsed subspace rather than true human alignment on subjective rubrics. Post-hoc calibration on human data improves alignment, but even calibrated LLMs fall short of human reliability.
This paper argues that traditional benchmarks both overestimate and underestimate frontier AI capabilities, and proposes 'open-world evaluations'—long-horizon, real-world tasks assessed qualitatively—as a complementary approach. The CRUX project is introduced, with a demonstration where an AI agent successfully published an iOS app to the App Store with minimal intervention.
This paper introduces Self-Supervised Prompt Optimization (SPO), a framework that optimizes prompts for LLMs without external references by using output comparisons, significantly reducing costs and data requirements.