evaluation-metrics

Tag

Cards List
#evaluation-metrics

Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning

arXiv cs.AI ↗ · 2026-08-25 Cached

This paper compares prior injection methods for sparse-reward reinforcement learning in vision-language math reasoning, finding that effectiveness depends on delivery and that certain evaluation slices can mislead generalization assessments.

0 favorites 0 likes
#evaluation-metrics

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

arXiv cs.LG ↗ · 2026-08-24 Cached

The paper introduces C-Score, a diagnostic framework for evaluating pseudo-label-based semi-supervised learning under open-world unlabeled contamination, demonstrating that clean accuracy is insufficient to detect hidden performance degradation.

0 favorites 0 likes
#evaluation-metrics

The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search

Hugging Face Daily Papers ↗ · 2026-08-24 Cached

This paper identifies flaws in RAG evaluation metrics by measuring context utilization causally, demonstrating that narrow sequential contexts improve recall over wide contexts, and introduces a submodular scheduler for optimized allocation.

0 favorites 0 likes
#evaluation-metrics

Transforming Heart Disease Prediction with Advanced Machine Learning Techniques

arXiv cs.LG ↗ · 2026-08-20 Cached

This research paper compares various machine learning classifiers for heart disease prediction, finding that Support Vector Machine and Simple Cart achieve the best performance on UCI and Kaggle datasets respectively, highlighting ML's potential to aid in early clinical diagnosis.

0 favorites 0 likes
#evaluation-metrics

Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions

arXiv cs.LG ↗ · 2026-08-20 Cached

This paper introduces an Interaction-based Prompt Sensitivity (IPS) metric to evaluate and explain prompt sensitivity in large language models by analyzing interactions. It applies IPS to 50 open-source LLMs, identifying factors like fine-tuning and model scale that reduce sensitivity through low-order interactions.

0 favorites 0 likes
#evaluation-metrics

Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting

arXiv cs.LG ↗ · 2026-08-19 Cached

The paper proposes Continuous-time Squared Error (CSE) as a less biased evaluation metric for irregular time series forecasting and introduces a systematic benchmark to assess models' continuous-time performance.

0 favorites 0 likes
#evaluation-metrics

Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents

arXiv cs.CL ↗ · 2026-08-19 Cached

This paper proposes that only AI agents capable of System 2 thinking should access untrusted documents in RAG systems to enhance security, introducing new metrics to evaluate robustness and showing reasoning models are more resistant to knowledge poisoning.

0 favorites 0 likes
#evaluation-metrics

Building World Models with Agent Swarms

Reddit r/AI_Agents ↗ · 2026-08-18

The author describes building a world model harness that coordinates research agents to maximize evaluation metrics, achieving a 25x score improvement in a multimodal masked reconstruction task using geospatial inputs.

0 favorites 0 likes
#evaluation-metrics

How Much Memory Does Your Agent Actually Need?

Hugging Face Blog ↗ · 2026-08-18 Cached

This research blog post discusses calibrating agentic memory based on model capability, showing that different AI models require tailored memory strategies for optimal performance, with ALTK-Evolve enabling self-distilled learning without human annotation.

0 favorites 0 likes
#evaluation-metrics

Held-out IC improved from 0.0613 to 0.0843 in this agent research paper

Reddit r/ArtificialInteligence ↗ · 2026-08-17

The paper reports an improvement in Information Coefficient from 0.0613 to 0.0843 using an agent-guided research process, compared to a GRU baseline, with simulated results on held-out data from 2021 to 2025.

0 favorites 0 likes
#evaluation-metrics

Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation

arXiv cs.CL ↗ · 2026-08-17 Cached

The paper proposes information satisfaction as a reader-centered axis for summarization evaluation, showing that current metrics fail to capture user-specific informational needs and agree poorly with human judgment.

0 favorites 0 likes
#evaluation-metrics

Forecast Collapse in Time-Series Foundation Models

Hugging Face Daily Papers ↗ · 2026-08-14 Cached

The paper identifies forecast collapse in time-series foundation models for hourly equity return prediction and introduces CalibRank to balance calibration and ranking, significantly improving cross-sectional correlation.

0 favorites 0 likes
#evaluation-metrics

When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper measures the self-referential evaluation loop when a knowledge base is used as the gold standard for entity-level machine translation in low-resource historical domains, showing that gains from KB injection are confined to overlapping segments and do not reflect true translation quality.

0 favorites 0 likes
#evaluation-metrics

The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom

arXiv cs.LG ↗ · 2026-08-12 Cached

This reproduction study independently reimplements LeWorldModel on the TwoRoom environment, reaching 94% goals rather than the reported 87%, and shows that four undocumented evaluation conventions determine the outcome. It also finds that one-step prediction error does not reliably predict long-horizon planning success and that batch normalization can inflate validation loss.

0 favorites 0 likes
#evaluation-metrics

Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper investigates whether automatic evaluation metrics for machine translation are reliable for Classical Chinese to English translation, using a diagnostic framework based on minimal pairs. It finds all metrics have blind spots, with MetricX24 performing best overall.

0 favorites 0 likes
#evaluation-metrics

UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation

arXiv cs.LG ↗ · 2026-08-04 Cached

UpliftBench is a benchmark paper showing that disagreements between uplift modeling evaluations often stem from metric choice rather than model quality, identifying specific mismatches between ranking metrics and deployment objectives across several dataset families.

0 favorites 0 likes
#evaluation-metrics

Representations from Pretrained Machine-Learning Interatomic Potentials as Coarse Coordinates for Material Generation and Evaluation

arXiv cs.LG ↗ · 2026-08-03 Cached

This paper proposes using atom-averaged features from pretrained MLIPs like MACE as coarse coordinates for evaluating and guiding inorganic crystal structure generation, introducing the Coarse-Fine Transport Distance (CFTD) metric that captures both quality and novelty in a distribution-based framework.

0 favorites 0 likes
#evaluation-metrics

VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P]

Reddit r/MachineLearning ↗ · 2026-08-01

This paper highlights that VLMs for chest x-ray report generation can score well on benchmarks while erasing clinically meaningful terms and introducing biased language, and proposes a framework to measure these failures.

0 favorites 0 likes
#evaluation-metrics

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting

arXiv cs.CL ↗ · 2026-07-30 Cached

This paper evaluates end-to-end trade-offs in moderation for conversational AI, comparing filter placement (input, response, both) and actions (blocking vs rewriting) using customer-outcome metrics like Usefulness and Harmful Exposure instead of component accuracy.

0 favorites 0 likes
#evaluation-metrics

SAFAARI: Schema-Aware Framework for Accelerated Advertiser Response Intelligence

arXiv cs.AI ↗ · 2026-07-29 Cached

This paper presents SAFAARI, a multi-agent framework that improves schema linking for NL-to-SQL systems in customer support, achieving an 81.66% SEAL score and 8x reduction in development time.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback