evaluation-methods

Tag

Cards List
#evaluation-methods

Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study

arXiv cs.CL ↗ · 4d ago Cached

This paper replicates a distributional-semantics extractive summarization method for Hindi and evaluates it on standard corpora, finding that sentence position is the only contributing feature and current Hindi benchmarks fail to incentivize advanced content selection.

0 favorites 0 likes
#evaluation-methods

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

arXiv cs.AI ↗ · 2026-09-21 Cached

This paper explores efficient recurring evaluation methods for production LLM agents, comparing techniques like adaptive testing and fixed subsets, and provides practical recommendations for deployment.

0 favorites 0 likes
#evaluation-methods

PersonalBench: Measuring the Authorship Gap in LLM Personalization

arXiv cs.CL ↗ · 2026-08-21 Cached

PersonalBench is a new benchmark that evaluates inference-time personalization methods in LLMs through authorship verification, LLM-as-judge, and stylometrics, finding that while methods produce author-differentiated output, they do not bridge the gap to human authorship.

0 favorites 0 likes
#evaluation-methods

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

arXiv cs.AI ↗ · 2026-08-20 Cached

The paper introduces THPT-Ladder, a benchmark that evaluates AI models using Vietnam's 2025 convex marking scheme, revealing how partial credit gaps lead to inaccurate assessments compared to standard accuracy metrics.

0 favorites 0 likes
#evaluation-methods

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

arXiv cs.CL ↗ · 2026-06-03 Cached

This paper geometrically analyzes why LLMs acting as judges agree strongly with each other but weakly with humans, finding that inter-LLM consensus reflects a collapsed subspace rather than true human alignment on subjective rubrics. Post-hoc calibration on human data improves alignment, but even calibrated LLMs fall short of human reliability.

0 favorites 0 likes
#evaluation-methods

Open-World Evaluations for Measuring Frontier AI Capabilities

arXiv cs.AI ↗ · 2026-05-22 Cached

This paper argues that traditional benchmarks both overestimate and underestimate frontier AI capabilities, and proposes 'open-world evaluations'—long-horizon, real-world tasks assessed qualitatively—as a complementary approach. The CRUX project is introduced, with a demonstration where an AI agent successfully published an iOS app to the App Store with minimal intervention.

0 favorites 0 likes
#evaluation-methods

Self-Supervised Prompt Optimization

Papers with Code Trending ↗ · 2025-02-07 Cached

This paper introduces Self-Supervised Prompt Optimization (SPO), a framework that optimizes prompts for LLMs without external references by using output comparisons, significantly reducing costs and data requirements.

0 favorites 0 likes
← Back to home

Submit Feedback