Tag
This paper compares prior injection methods for sparse-reward reinforcement learning in vision-language math reasoning, finding that effectiveness depends on delivery and that certain evaluation slices can mislead generalization assessments.
The paper introduces C-Score, a diagnostic framework for evaluating pseudo-label-based semi-supervised learning under open-world unlabeled contamination, demonstrating that clean accuracy is insufficient to detect hidden performance degradation.
This paper identifies flaws in RAG evaluation metrics by measuring context utilization causally, demonstrating that narrow sequential contexts improve recall over wide contexts, and introduces a submodular scheduler for optimized allocation.
This research paper compares various machine learning classifiers for heart disease prediction, finding that Support Vector Machine and Simple Cart achieve the best performance on UCI and Kaggle datasets respectively, highlighting ML's potential to aid in early clinical diagnosis.
This paper introduces an Interaction-based Prompt Sensitivity (IPS) metric to evaluate and explain prompt sensitivity in large language models by analyzing interactions. It applies IPS to 50 open-source LLMs, identifying factors like fine-tuning and model scale that reduce sensitivity through low-order interactions.
The paper proposes Continuous-time Squared Error (CSE) as a less biased evaluation metric for irregular time series forecasting and introduces a systematic benchmark to assess models' continuous-time performance.
This paper proposes that only AI agents capable of System 2 thinking should access untrusted documents in RAG systems to enhance security, introducing new metrics to evaluate robustness and showing reasoning models are more resistant to knowledge poisoning.
The author describes building a world model harness that coordinates research agents to maximize evaluation metrics, achieving a 25x score improvement in a multimodal masked reconstruction task using geospatial inputs.
This research blog post discusses calibrating agentic memory based on model capability, showing that different AI models require tailored memory strategies for optimal performance, with ALTK-Evolve enabling self-distilled learning without human annotation.
The paper reports an improvement in Information Coefficient from 0.0613 to 0.0843 using an agent-guided research process, compared to a GRU baseline, with simulated results on held-out data from 2021 to 2025.
The paper proposes information satisfaction as a reader-centered axis for summarization evaluation, showing that current metrics fail to capture user-specific informational needs and agree poorly with human judgment.
The paper identifies forecast collapse in time-series foundation models for hourly equity return prediction and introduces CalibRank to balance calibration and ranking, significantly improving cross-sectional correlation.
This paper measures the self-referential evaluation loop when a knowledge base is used as the gold standard for entity-level machine translation in low-resource historical domains, showing that gains from KB injection are confined to overlapping segments and do not reflect true translation quality.
This reproduction study independently reimplements LeWorldModel on the TwoRoom environment, reaching 94% goals rather than the reported 87%, and shows that four undocumented evaluation conventions determine the outcome. It also finds that one-step prediction error does not reliably predict long-horizon planning success and that batch normalization can inflate validation loss.
This paper investigates whether automatic evaluation metrics for machine translation are reliable for Classical Chinese to English translation, using a diagnostic framework based on minimal pairs. It finds all metrics have blind spots, with MetricX24 performing best overall.
UpliftBench is a benchmark paper showing that disagreements between uplift modeling evaluations often stem from metric choice rather than model quality, identifying specific mismatches between ranking metrics and deployment objectives across several dataset families.
This paper proposes using atom-averaged features from pretrained MLIPs like MACE as coarse coordinates for evaluating and guiding inorganic crystal structure generation, introducing the Coarse-Fine Transport Distance (CFTD) metric that captures both quality and novelty in a distribution-based framework.
This paper highlights that VLMs for chest x-ray report generation can score well on benchmarks while erasing clinically meaningful terms and introducing biased language, and proposes a framework to measure these failures.
This paper evaluates end-to-end trade-offs in moderation for conversational AI, comparing filter placement (input, response, both) and actions (blocking vs rewriting) using customer-outcome metrics like Usefulness and Harmful Exposure instead of component accuracy.
This paper presents SAFAARI, a multi-agent framework that improves schema linking for NL-to-SQL systems in customer support, achieving an 81.66% SEAL score and 8x reduction in development time.