benchmark-evaluation

Tag

Cards List
#benchmark-evaluation

TW3Cast: A Frozen Router of Lightly Fine-Tuned Foundation Models for Time-Series Forecasting on GIFT-Eval, Selected Entirely on the Training Split

arXiv cs.AI ↗ · 2d ago Cached

The paper presents TW3Cast, a time-series forecasting system that uses a frozen router of lightly fine-tuned foundation models to achieve top performance on the GIFT-Eval benchmark without agents or language models.

0 favorites 0 likes
#benchmark-evaluation

Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure

arXiv cs.CL ↗ · 2d ago Cached

This paper audits a multilingual affective generation benchmark and reveals that system rankings are driven by measurement artifacts rather than genuine performance differences, with annotator variability playing a key role.

0 favorites 0 likes
#benchmark-evaluation

@xeophon: Pacing the frontier

X AI KOLs Timeline ↗ · 6d ago Cached

Vals AI evaluated Grok 4.7, finding it ranks #24 on the Vals Index with a score of 54.2%, down from Grok 4.6, but shows improvements in legal and medical domains.

0 favorites 0 likes
#benchmark-evaluation

Transferring the Intelligence of VLMs to Robotic Control

Hugging Face Daily Papers ↗ · 2026-09-19 Cached

This paper presents RoboDawn, a method to transfer Vision-Language Model intelligence to robotic control, achieving state-of-the-art results on benchmarks with zero-shot and one-shot learning and successful real-world applications.

0 favorites 0 likes
#benchmark-evaluation

@rohanpaul_ai: New Yale Univ + other top lab paper shows frontier models are already close to maxing out today’s closed-ended physics …

X AI KOLs Following ↗ · 2026-09-16 Cached

A Yale-led paper finds that frontier AI models' poor performance on physics benchmarks is largely due to benchmark flaws rather than model limitations, suggesting benchmark quality is a critical bottleneck for accurate evaluation.

0 favorites 0 likes
#benchmark-evaluation

CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine

arXiv cs.AI ↗ · 2026-09-16 Cached

The paper proposes CLEAR, an agentic framework for cross-source evidence adjudication to improve large language models in medicine by handling conflicts from multiple knowledge sources. It demonstrates competitive performance across benchmarks, with significant gains in settings where direct inference or retrieval is weak.

0 favorites 0 likes
#benchmark-evaluation

Retrieval-Driven Memory Reconsolidation for Long-Term LLM Agents

arXiv cs.CL ↗ · 2026-09-16 Cached

The paper introduces REALM, a framework for long-term memory in LLM agents that uses retrieval-driven reconsolidation to autonomously organize memories into a cognitive graph, achieving improved performance on long-term memory benchmarks.

0 favorites 0 likes
#benchmark-evaluation

Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs

arXiv cs.CL ↗ · 2026-09-15 Cached

This study examines the action-level reliability of clinical LLM agents by rerunning tasks with identical inputs and comparing orders, finding significant divergence that benchmarks may miss and proposing enhanced evaluation methods.

0 favorites 0 likes
#benchmark-evaluation

Towards a Deterministic Math Solver for Clinical Language Models

arXiv cs.AI ↗ · 2026-09-12 Cached

This paper proposes a Program-Solve interface where clinical language models generate Python code for a deterministic executor to perform math calculations, evaluating on MedCalc-Bench and finding improved accuracy for larger models like Qwen2.5-32B compared to direct arithmetic and hand-written libraries.

0 favorites 0 likes
#benchmark-evaluation

Apodex-1.1-mini-GGUF*Hugging Face

Reddit r/LocalLLaMA ↗ · 2026-09-10 Cached

Apodex-1.1 is a reasoning-first AI model for complex, long-horizon research tasks, with end-to-end execution, adaptive agent teams, and built-in verification to deliver verifiable results.

0 favorites 0 likes
#benchmark-evaluation

Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses

arXiv cs.AI ↗ · 2026-09-10 Cached

This paper studies agent harness optimization to improve LLM tool agents without retraining, focusing on prompts and tool-boundary middleware, and introduces a protocol and the PRISM optimizer for measurable gains.

0 favorites 0 likes
#benchmark-evaluation

RecKAN: Kolmogorov-Arnold Networks with a Learnable Recursive Polynomial Basis

arXiv cs.LG ↗ · 2026-09-03 Cached

RecKAN introduces a learnable recursive polynomial basis for Kolmogorov-Arnold Networks, outperforming existing KAN variants on classification and forecasting tasks.

0 favorites 0 likes
#benchmark-evaluation

From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper proposes methods for detecting hallucinations in black-box LLMs by combining semantic entropy and token-level uncertainty signals, evaluating techniques like TopK, CoCoA, Gated, and Stacked across multiple benchmarks to find that no single method is universally strongest but Stacked often performs best.

0 favorites 0 likes
#benchmark-evaluation

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

Hugging Face Daily Papers ↗ · 2026-09-03

This paper introduces FactoSR, a factorized reinforcement learning framework that enhances spatial reasoning in Vision-Language Models by decomposing 4D properties into orthogonal sub-objectives, achieving significant performance boosts on multi-view and video benchmarks.

0 favorites 0 likes
#benchmark-evaluation

Neural means and kernel corrections for operator learning

arXiv cs.LG ↗ · 2026-09-02 Cached

This paper presents a method combining neural network means with exact Matérn kernel corrections for operator learning in PDEs, achieving competitive or improved performance on public benchmarks like structural mechanics and OCO-2 radiative transfer emulation.

0 favorites 0 likes
#benchmark-evaluation

Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion

arXiv cs.LG ↗ · 2026-08-28 Cached

The paper identifies strong-modality collapse in multimodal learning where fusion degrades the dominant modality's performance, and proposes Inverted Asymmetric Fusion (IAF) to preserve it, improving over unimodal baselines.

0 favorites 0 likes
#benchmark-evaluation

The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure

arXiv cs.CL ↗ · 2026-08-27 Cached

This paper reevaluates the imperfective paradox benchmark for large language models, identifying conceptual and evaluation mis-specifications, and introduces lexically matched minimal pairs to reveal sufficiency bias in models' semantic reasoning.

0 favorites 0 likes
#benchmark-evaluation

Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization

arXiv cs.AI ↗ · 2026-08-24 Cached

This paper proposes a weight-delta audit method to analyze medical specialization in language models, examining internal update changes beyond endpoint benchmark scores.

0 favorites 0 likes
#benchmark-evaluation

Position: Profiling Game Worlds by Transition Complexity

arXiv cs.AI ↗ · 2026-08-20 Cached

This paper proposes the Transition Complexity Profile (TCP), a set of metrics to quantify the difficulty of transition prediction in game world modeling, aiming to standardize benchmark evaluation in reinforcement learning and world model research.

0 favorites 0 likes
#benchmark-evaluation

Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal

arXiv cs.CL ↗ · 2026-08-19 Cached

This paper audits temporal leakage in financial news NLP benchmarks across multiple models, finding that random splits inflate performance metrics and identifies M&A events as a category with a localized positive signal under chronological evaluation.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback