evaluation-metrics

Tag

Cards List
#evaluation-metrics

Exposing Blind Spots in Deep Imbalanced Regression Evaluation

arXiv cs.LG ↗ · 2026-09-23 Cached

This paper identifies blind spots in evaluating deep imbalanced regression, proposing balanced metrics and showing high tail-region instability across random seeds.

0 favorites 0 likes
#evaluation-metrics

BLOG-2 of series AI SYSTEM DESIGN

Reddit r/ArtificialInteligence ↗ · 2026-09-22

This blog post discusses how to benchmark LLMs for specific production use cases, proposing a comprehensive evaluation process that includes task-specific metrics, trade-offs between quality and cost, and analysis of failure cases.

0 favorites 0 likes
#evaluation-metrics

WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction

arXiv cs.LG ↗ · 2026-09-22 Cached

This paper introduces WildfireSpreadBench to benchmark wildfire spread prediction models, revealing that evaluation metrics like AP versus F1 can lead to different model rankings and affect operational suitability.

0 favorites 0 likes
#evaluation-metrics

Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations

arXiv cs.CL ↗ · 2026-09-21 Cached

This study benchmarks gender bias in machine translation evaluation metrics across occupations, revealing that masculine translations tend to score higher and that biases vary by language and evaluator.

0 favorites 0 likes
#evaluation-metrics

How Many Humans Is a Judge Panel Worth?

arXiv cs.CL ↗ · 2026-09-21 Cached

This paper introduces methods to quantify how effectively a panel of language models approximates human judgments, using spectral diversity and distribution recovery metrics to measure effective representation.

0 favorites 0 likes
#evaluation-metrics

When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

arXiv cs.CL ↗ · 2026-09-21 Cached

This paper diagnoses the gap between next-turn evaluation metrics and autonomous workflow execution success for AI agents, showing that supervised fine-tuning improves turn-level performance but fails to enhance end-to-end workflow success.

0 favorites 0 likes
#evaluation-metrics

Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models

arXiv cs.AI ↗ · 2026-09-18 Cached

This paper critiques conventional depth truncation methods for evaluating recursive language models and introduces the Depth Control Protocol (DCP) to disentangle and isolate factors affecting depth utilization, improving evaluation accuracy.

0 favorites 0 likes
#evaluation-metrics

The Missing "I Don't Know": Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention

arXiv cs.LG ↗ · 2026-09-17 Cached

The article argues that three independent findings on LLM reliability problems converge on the need for calibrated abstention, where models can appropriately decline to answer when uncertain, and proposes evaluation reforms such as triple-scoring and calibration metrics to address this gap.

0 favorites 0 likes
#evaluation-metrics

In the Blind: Building Pseudo-References for MT Evaluation

arXiv cs.CL ↗ · 2026-09-15 Cached

This paper describes a method for building pseudo-references in machine translation evaluation without human references, using multiple translation models, quality estimation, and GPT-5.5 post-editing to improve selection accuracy.

0 favorites 0 likes
#evaluation-metrics

Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data

arXiv cs.AI ↗ · 2026-09-12 Cached

The paper introduces a generator that creates consistent fictional enterprise data without real datasets, using reference-free evaluation methods to ensure realism, and includes a hosted service for building relational databases from business questions.

0 favorites 0 likes
#evaluation-metrics

If coding is solved, what now?: Measuring the sloppiness of code

Hacker News Top ↗ · 2026-09-11 Cached

The article explores how LLMs can generate correct code but often introduce sloppiness like unnecessary abstractions, and discusses methods to measure code quality, including using AI judges and human evaluation.

0 favorites 0 likes
#evaluation-metrics

ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation

arXiv cs.CL ↗ · 2026-09-11 Cached

ProMediConv is a benchmarking framework for proactive conversational agents in legal dispute mediation, using real-world cases to model multi-stage dialogues and propose new evaluation metrics. It aims to advance AI-assisted conflict resolution by addressing gaps in current research.

0 favorites 0 likes
#evaluation-metrics

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

arXiv cs.CL ↗ · 2026-09-11 Cached

This paper investigates how surface noise in text affects LLM judges' bias measurement, finding that it systematically overestimates bias, particularly in fairness-critical categories, and introduces the Fable benchmark to study this issue.

0 favorites 0 likes
#evaluation-metrics

Time for a new benchmark

Reddit r/singularity ↗ · 2026-09-08

The article discusses the need for a new benchmark in AI to better evaluate model performance and address current limitations in existing standards.

0 favorites 0 likes
#evaluation-metrics

SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation

arXiv cs.AI ↗ · 2026-09-04 Cached

SVG-Score introduces a human-aligned evaluation framework for text-to-SVG generation, addressing limitations of current metrics like CLIPScore by developing specialized evaluators and benchmarks.

0 favorites 0 likes
#evaluation-metrics

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

arXiv cs.AI ↗ · 2026-09-03 Cached

CivBench is an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments using Civilization VI, introducing metrics like Proactive Monitoring Rate and RAG@10 to assess agent behavior.

0 favorites 0 likes
#evaluation-metrics

UTP-Bench: Uncertainty-aware Travel Planning Benchmark

arXiv cs.AI ↗ · 2026-09-03 Cached

UTP-Bench is a new benchmark for uncertainty-aware travel planning that evaluates LLMs on robustness against delays and crowd variability using real-world data from India.

0 favorites 0 likes
#evaluation-metrics

@dair_ai: Good measurement work on whether retrieved agent skills actually help. They report that agent skills that lift your agg…

X AI KOLs Timeline ↗ · 2026-09-03 Cached

The paper introduces Retrieval-Invoked Actual-Use Effect (RAE) to evaluate whether retrieved skills actually help LLM agents, showing that aggregate metrics can mislead by hiding negative effects on specific tasks.

0 favorites 0 likes
#evaluation-metrics

@antirez: How do you trust a metric where 1st and 2nd place are those? GPT 5.6 Sol is in the real world a *fundamentally* strong …

X AI KOLs Timeline ↗ · 2026-08-27 Cached

The tweet questions the trustworthiness of an AI model ranking metric, asserting that GPT 5.6 Sol is fundamentally stronger than Opus 5 and implying the rest of the chart may be unreliable.

0 favorites 0 likes
#evaluation-metrics

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

arXiv cs.CL ↗ · 2026-08-27 Cached

This paper tests whether decodable empathy directions in LLMs can reliably shift automated empathy scores, finding that affective facet control is partial and cognitive steering is inconsistent, highlighting that detection does not imply control.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback