Tag
Presents LitTraceQA, a benchmark for scientific question answering that requires systems to retrieve relevant papers, locate supporting evidence, and produce verified answers in multiple formats.
This paper introduces a concept-centric benchmark to probe LLMs' understanding of geo-spatial concepts like direction, distance, and topology, testing abstraction, compositionality, and grounding across various model architectures and scales. Findings reveal clear limitations in current LLMs' conceptual understanding.
This paper introduces PHASE-Tree, a multi-timescale character-state representation for long-horizon role-playing dialogue, along with the LongEvoRoleBench benchmark to evaluate evolved-state generation. The proposed approach outperforms baselines on character-level, semantic, and embedding metrics across long-dialogue corpora.
Introduces Ekphrasis, a 400-task benchmark for measuring visual creative ideation in text-only LLMs, separating usefulness, expressiveness, and novelty. The paper validates it with cross-modal grounding, showing text-level visual ideation ordering survives rendering.
Ask-E is a new benchmark and training environment that evaluates and trains models on generating questions calibrated to specific skill levels, defined by the capabilities of two existing language models. Frontier models score below 50% on calibration, and training on Ask-E improves downstream math benchmarks without new math data or correctness-based rewards.
LLMRouter presents a unified formulation of LLM routing as a sequential decision process, along with an open-source infrastructure and benchmark (xRouteBench) for developing, evaluating, and deploying LLM routers. Empirical results show learned routers achieve 14.6% relative improvement over the strongest fixed-model baseline.
Presents ArchEGraph, a large-scale graph dataset for building energy modeling with aligned geometry, topology, weather, and thermal loads, along with benchmark tasks for graph reconstruction and load prediction.
Introduces TradeVerse, a benchmark built from WTO meeting minutes that tests LLMs on longitudinal political trade negotiations, including tasks like predicting product categories, identifying responding countries, and generating statements.
This paper introduces the Cross-Lingual Comprehension Gap (CLCG) metric to measure how LLM response quality degrades when content is presented in non-English languages. Across 18 languages and multiple models, it finds a significant performance drop, especially for low-resource languages, questioning the assumption of English-centric capability transfer.
ED-CSP is a machine learning model that predicts crystal structures from electron diffraction patterns, achieving improved match rates over the PXRD-based PXRDGen and demonstrating the value of multi-view diffraction data.
This paper investigates whether large language models reproduce critical acclaim versus commercial success hierarchies in film preferences, finding across 20,000 pairwise comparisons that models consistently favor critically acclaimed but commercially obscure films, with this orientation growing with model scale.
This paper introduces SEE, a multimodal benchmark of expert-curated questions for scientific discovery in chemistry, biology, and materials science. Evaluation of 19 MLLMs shows the best model reaches only 48.7% accuracy, and even with tool use only 52.7%, revealing that current models lack reliable evidence-bounded scientific reasoning.
This paper introduces a unified benchmark and fine-grained annotation framework for long-horizon agent trajectory attribution, enabling evaluation of primary attribution localization and attribution-chain recovery across diverse settings. It provides over 1,300 annotated trajectories from existing agent benchmarks and releases a reusable annotation skill for standardizing future trajectories.
WebRider is a hierarchical framework that formalizes delegated web tasks as intent contracts, preserving persona-conditioned policies through every browsing step. It includes RiderBench, a benchmark of 4,096 live-web contracts, and an 8B action-policy model trained through its guarded interface.
TRACE is a new multi-layer benchmark for diagnosing drift and failures in human-AI controller coordination, built from ALFRED traces with 1,918 drifted samples annotated across five execution layers. Baseline results show drift identification and attribution well above random baselines across classical, recurrent, and attention-based model families.
A skeptical take on the BigBang-v1 finetune of Qwen 3.5, noting suspicious benchmark claims and potential test contamination despite Bartowski's GGUF conversion.
An experimental reasoning system at Orivael scored 100% on ARC-AGI-3 ft09 with zero model calls, revealing that its failures stem from incorrect environment representations rather than planning errors.
A community member benchmarked seven self-hostable memory providers for Hermes Agent, testing each with 71,060 conversation turns and 3,750 questions about changing facts; full results and GitHub repo are in the thread.
A user shares updated benchmark results for DeepSeek V4 Flash on SlopCodeBench using local quants (antirez imatrix quant) with the pi harness, showing improved performance over previous runs but still slower than the hosted API.
Enabling PCI-E peer-to-peer (P2P) for consumer Nvidia GPUs with patched drivers and vLLM environment variables yields roughly 25% prefill throughput improvement for free, as demonstrated by benchmarks.