Tag
The paper presents TW3Cast, a time-series forecasting system that uses a frozen router of lightly fine-tuned foundation models to achieve top performance on the GIFT-Eval benchmark without agents or language models.
This paper audits a multilingual affective generation benchmark and reveals that system rankings are driven by measurement artifacts rather than genuine performance differences, with annotator variability playing a key role.
Vals AI evaluated Grok 4.7, finding it ranks #24 on the Vals Index with a score of 54.2%, down from Grok 4.6, but shows improvements in legal and medical domains.
This paper presents RoboDawn, a method to transfer Vision-Language Model intelligence to robotic control, achieving state-of-the-art results on benchmarks with zero-shot and one-shot learning and successful real-world applications.
A Yale-led paper finds that frontier AI models' poor performance on physics benchmarks is largely due to benchmark flaws rather than model limitations, suggesting benchmark quality is a critical bottleneck for accurate evaluation.
The paper proposes CLEAR, an agentic framework for cross-source evidence adjudication to improve large language models in medicine by handling conflicts from multiple knowledge sources. It demonstrates competitive performance across benchmarks, with significant gains in settings where direct inference or retrieval is weak.
The paper introduces REALM, a framework for long-term memory in LLM agents that uses retrieval-driven reconsolidation to autonomously organize memories into a cognitive graph, achieving improved performance on long-term memory benchmarks.
This study examines the action-level reliability of clinical LLM agents by rerunning tasks with identical inputs and comparing orders, finding significant divergence that benchmarks may miss and proposing enhanced evaluation methods.
This paper proposes a Program-Solve interface where clinical language models generate Python code for a deterministic executor to perform math calculations, evaluating on MedCalc-Bench and finding improved accuracy for larger models like Qwen2.5-32B compared to direct arithmetic and hand-written libraries.
Apodex-1.1 is a reasoning-first AI model for complex, long-horizon research tasks, with end-to-end execution, adaptive agent teams, and built-in verification to deliver verifiable results.
This paper studies agent harness optimization to improve LLM tool agents without retraining, focusing on prompts and tool-boundary middleware, and introduces a protocol and the PRISM optimizer for measurable gains.
RecKAN introduces a learnable recursive polynomial basis for Kolmogorov-Arnold Networks, outperforming existing KAN variants on classification and forecasting tasks.
This paper proposes methods for detecting hallucinations in black-box LLMs by combining semantic entropy and token-level uncertainty signals, evaluating techniques like TopK, CoCoA, Gated, and Stacked across multiple benchmarks to find that no single method is universally strongest but Stacked often performs best.
This paper introduces FactoSR, a factorized reinforcement learning framework that enhances spatial reasoning in Vision-Language Models by decomposing 4D properties into orthogonal sub-objectives, achieving significant performance boosts on multi-view and video benchmarks.
This paper presents a method combining neural network means with exact Matérn kernel corrections for operator learning in PDEs, achieving competitive or improved performance on public benchmarks like structural mechanics and OCO-2 radiative transfer emulation.
The paper identifies strong-modality collapse in multimodal learning where fusion degrades the dominant modality's performance, and proposes Inverted Asymmetric Fusion (IAF) to preserve it, improving over unimodal baselines.
This paper reevaluates the imperfective paradox benchmark for large language models, identifying conceptual and evaluation mis-specifications, and introduces lexically matched minimal pairs to reveal sufficiency bias in models' semantic reasoning.
This paper proposes a weight-delta audit method to analyze medical specialization in language models, examining internal update changes beyond endpoint benchmark scores.
This paper proposes the Transition Complexity Profile (TCP), a set of metrics to quantify the difficulty of transition prediction in game world modeling, aiming to standardize benchmark evaluation in reinforcement learning and world model research.
This paper audits temporal leakage in financial news NLP benchmarks across multiple models, finding that random splits inflate performance metrics and identifies M&A events as a category with a localized positive signal under chronological evaluation.