Tag
This paper identifies blind spots in evaluating deep imbalanced regression, proposing balanced metrics and showing high tail-region instability across random seeds.
This blog post discusses how to benchmark LLMs for specific production use cases, proposing a comprehensive evaluation process that includes task-specific metrics, trade-offs between quality and cost, and analysis of failure cases.
This paper introduces WildfireSpreadBench to benchmark wildfire spread prediction models, revealing that evaluation metrics like AP versus F1 can lead to different model rankings and affect operational suitability.
This study benchmarks gender bias in machine translation evaluation metrics across occupations, revealing that masculine translations tend to score higher and that biases vary by language and evaluator.
This paper introduces methods to quantify how effectively a panel of language models approximates human judgments, using spectral diversity and distribution recovery metrics to measure effective representation.
This paper diagnoses the gap between next-turn evaluation metrics and autonomous workflow execution success for AI agents, showing that supervised fine-tuning improves turn-level performance but fails to enhance end-to-end workflow success.
This paper critiques conventional depth truncation methods for evaluating recursive language models and introduces the Depth Control Protocol (DCP) to disentangle and isolate factors affecting depth utilization, improving evaluation accuracy.
The article argues that three independent findings on LLM reliability problems converge on the need for calibrated abstention, where models can appropriately decline to answer when uncertain, and proposes evaluation reforms such as triple-scoring and calibration metrics to address this gap.
This paper describes a method for building pseudo-references in machine translation evaluation without human references, using multiple translation models, quality estimation, and GPT-5.5 post-editing to improve selection accuracy.
The paper introduces a generator that creates consistent fictional enterprise data without real datasets, using reference-free evaluation methods to ensure realism, and includes a hosted service for building relational databases from business questions.
The article explores how LLMs can generate correct code but often introduce sloppiness like unnecessary abstractions, and discusses methods to measure code quality, including using AI judges and human evaluation.
ProMediConv is a benchmarking framework for proactive conversational agents in legal dispute mediation, using real-world cases to model multi-stage dialogues and propose new evaluation metrics. It aims to advance AI-assisted conflict resolution by addressing gaps in current research.
This paper investigates how surface noise in text affects LLM judges' bias measurement, finding that it systematically overestimates bias, particularly in fairness-critical categories, and introduces the Fable benchmark to study this issue.
The article discusses the need for a new benchmark in AI to better evaluate model performance and address current limitations in existing standards.
SVG-Score introduces a human-aligned evaluation framework for text-to-SVG generation, addressing limitations of current metrics like CLIPScore by developing specialized evaluators and benchmarks.
CivBench is an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments using Civilization VI, introducing metrics like Proactive Monitoring Rate and RAG@10 to assess agent behavior.
UTP-Bench is a new benchmark for uncertainty-aware travel planning that evaluates LLMs on robustness against delays and crowd variability using real-world data from India.
The paper introduces Retrieval-Invoked Actual-Use Effect (RAE) to evaluate whether retrieved skills actually help LLM agents, showing that aggregate metrics can mislead by hiding negative effects on specific tasks.
The tweet questions the trustworthiness of an AI model ranking metric, asserting that GPT 5.6 Sol is fundamentally stronger than Opus 5 and implying the rest of the chart may be unreliable.
This paper tests whether decodable empathy directions in LLMs can reliably shift automated empathy scores, finding that affective facet control is partial and cognitive steering is inconsistent, highlighting that detection does not imply control.