benchmark

Tag

Cards List
#benchmark

LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering

arXiv cs.CL · 5h ago Cached

Presents LitTraceQA, a benchmark for scientific question answering that requires systems to retrieve relevant papers, locate supporting evidence, and produce verified answers in multiple formats.

0 favorites 0 likes
#benchmark

Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding

arXiv cs.CL · 5h ago Cached

This paper introduces a concept-centric benchmark to probe LLMs' understanding of geo-spatial concepts like direction, distance, and topology, testing abstraction, compositionality, and grounding across various model architectures and scales. Findings reveal clear limitations in current LLMs' conceptual understanding.

0 favorites 0 likes
#benchmark

PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue

arXiv cs.CL · 5h ago Cached

This paper introduces PHASE-Tree, a multi-timescale character-state representation for long-horizon role-playing dialogue, along with the LongEvoRoleBench benchmark to evaluate evolved-state generation. The proposed approach outperforms baselines on character-level, semantic, and embedding metrics across long-dialogue corpora.

0 favorites 0 likes
#benchmark

Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs

arXiv cs.CL · 5h ago Cached

Introduces Ekphrasis, a 400-task benchmark for measuring visual creative ideation in text-only LLMs, separating usefulness, expressiveness, and novelty. The paper validates it with cross-modal grounding, showing text-level visual ideation ordering survives rendering.

0 favorites 0 likes
#benchmark

Ask-E: An Environment for Calibrated Question Generation

arXiv cs.CL · 5h ago Cached

Ask-E is a new benchmark and training environment that evaluates and trains models on generating questions calibrated to specific skill levels, defined by the capabilities of two existing language models. Frontier models score below 50% on calibration, and training on Ask-E improves downstream math benchmarks without new math data or correctness-based rewards.

0 favorites 0 likes
#benchmark

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

arXiv cs.CL · 5h ago Cached

LLMRouter presents a unified formulation of LLM routing as a sequential decision process, along with an open-source infrastructure and benchmark (xRouteBench) for developing, evaluating, and deploying LLM routers. Empirical results show learned routers achieve 14.6% relative improvement over the strongest fixed-model baseline.

0 favorites 0 likes
#benchmark

ArchEGraph: A Large-Scale Graph Dataset for Geometry-Topology-Physics Aligned Building Energy Modeling

arXiv cs.LG · 5h ago Cached

Presents ArchEGraph, a large-scale graph dataset for building energy modeling with aligned geometry, topology, weather, and thermal loads, along with benchmark tasks for graph reconstruction and load prediction.

0 favorites 0 likes
#benchmark

TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade

arXiv cs.CL · 5h ago Cached

Introduces TradeVerse, a benchmark built from WTO meeting minutes that tests LLMs on longitudinal political trade negotiations, including tasks like predicting product categories, identifying responding countries, and generating statements.

0 favorites 0 likes
#benchmark

Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand

arXiv cs.CL · 5h ago Cached

This paper introduces the Cross-Lingual Comprehension Gap (CLCG) metric to measure how LLM response quality degrades when content is presented in non-English languages. Across 18 languages and multiple models, it finds a significant performance drop, especially for low-resource languages, questioning the assumption of English-centric capability transfer.

0 favorites 0 likes
#benchmark

ED-CSP: Crystal Structure Prediction from Electron Diffraction

arXiv cs.LG · 5h ago Cached

ED-CSP is a machine learning model that predicts crystal structures from electron diffraction patterns, achieving improved match rates over the PXRD-based PXRDGen and demonstrating the value of multi-view diffraction data.

0 favorites 0 likes
#benchmark

Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

arXiv cs.AI · 5h ago Cached

This paper investigates whether large language models reproduce critical acclaim versus commercial success hierarchies in film preferences, finding across 20,000 pairwise comparisons that models consistently favor critically acclaimed but commercially obscure films, with this orientation growing with model scale.

0 favorites 0 likes
#benchmark

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

arXiv cs.AI · 5h ago Cached

This paper introduces SEE, a multimodal benchmark of expert-curated questions for scientific discovery in chemistry, biology, and materials science. Evaluation of 19 MLLMs shows the best model reaches only 48.7% accuracy, and even with tool use only 52.7%, revealing that current models lack reliable evidence-bounded scientific reasoning.

0 favorites 0 likes
#benchmark

Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework

arXiv cs.AI · 5h ago Cached

This paper introduces a unified benchmark and fine-grained annotation framework for long-horizon agent trajectory attribution, enabling evaluation of primary attribution localization and attribution-chain recovery across diverse settings. It provides over 1,300 annotated trajectories from existing agent benchmarks and releases a reusable annotation skill for standardizing future trajectories.

0 favorites 0 likes
#benchmark

WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance

arXiv cs.AI · 5h ago Cached

WebRider is a hierarchical framework that formalizes delegated web tasks as intent contracts, preserving persona-conditioned policies through every browsing step. It includes RiderBench, a benchmark of 4,096 live-web contracts, and an 8B action-policy model trained through its guarded interface.

0 favorites 0 likes
#benchmark

TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure

arXiv cs.AI · 5h ago Cached

TRACE is a new multi-layer benchmark for diagnosing drift and failures in human-AI controller coordination, built from ALFRED traces with 1,918 drifted samples annotated across five execution layers. Baseline results show drift identification and attribution well above random baselines across classical, recurrent, and attention-based model families.

0 favorites 0 likes
#benchmark

endless-frontier/BigBang-v1 - qwen 3.5 finetunes

Reddit r/LocalLLaMA · 11h ago

A skeptical take on the BigBang-v1 finetune of Qwen 3.5, noting suspicious benchmark claims and potential test contamination despite Bartowski's GGUF conversion.

0 favorites 0 likes
#benchmark

We got 100% on ARC-AGI-3 ft09 with zero model calls. The failures are more interesting.

Reddit r/artificial · 13h ago

An experimental reasoning system at Orivael scored 100% on ARC-AGI-3 ft09 with zero model calls, revealing that its failures stem from incorrect environment representations rather than planning errors.

0 favorites 0 likes
#benchmark

@witcheer: someone in the community worked on a real benchmark of seven self-hostable memory providers for Hermes Agent, each fed …

X AI KOLs Timeline · 21h ago Cached

A community member benchmarked seven self-hostable memory providers for Hermes Agent, testing each with 71,060 conversation turns and 3,750 questions about changing facts; full results and GitHub repo are in the thread.

0 favorites 0 likes
#benchmark

Updated benchmark: Deepseek V4 Flash on SlopCodeBench (local)

Reddit r/LocalLLaMA · yesterday

A user shares updated benchmark results for DeepSeek V4 Flash on SlopCodeBench using local quants (antirez imatrix quant) with the pi harness, showing improved performance over previous runs but still slower than the hosted API.

0 favorites 0 likes
#benchmark

enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think

Reddit r/LocalLLaMA · yesterday

Enabling PCI-E peer-to-peer (P2P) for consumer Nvidia GPUs with patched drivers and vLLM environment variables yields roughly 25% prefill throughput improvement for free, as demonstrated by benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback