Tag
This paper defines a hierarchy of faithfulness criteria for knowledge base completion models and evaluates current embedding models, finding they are not logically faithful.
This paper introduces NAF-Bench to study how large language models adhere to specified negation semantics, finding that frontier models like o4-mini perform well while open-source models lag, and suggesting improvements via solver delegation or fine-tuning.
This paper evaluates small language models beyond answer accuracy in knowledge graph question answering by isolating graph navigation capabilities, showing significant differences in path fidelity and the need for broader evaluation metrics.
MolDesignBench is a new benchmark for evaluating LLM-based agents in scenario-grounded molecular design, comprising 2K instances with implicit and explicit constraints, revealing that current frontier LLMs achieve low success rates, especially in reasoning about implicit constraints and infeasibility detection.
PRISM-VLM is a multi-axis discriminative benchmark that evaluates compact vision-language models along seven axes, including task quality and behavioral robustness, to provide more reliable separation and insights compared to single-axis benchmarks. It aims to release an open pipeline for the community to improve AI evaluation methods.
PotARCin is a multi-dimensional benchmark that extends ARC to evaluate abstract reasoning skills across five dimensions, showing that standard evaluations may not fully capture AI models' true capabilities.
The paper introduces TACT, a benchmark for evaluating turn-taking in full-duplex spoken dialogue models that uses intent-conditioned continuous scoring to replace binary rules, demonstrating better alignment with human judgments.
This paper proposes a method to estimate the causal effect of benchmark exposure on AI model performance, moving beyond traditional overlap techniques for more robust evaluation.
This paper presents a controlled study on learned context planning for long-context multiple-choice QA, demonstrating that it does not robustly outperform strong retrieval methods like BM25 and hybrid retrieval under various experimental settings.
The article reflects on the rapid progress in AI models from three years ago, comparing early models like Bard that struggled with coding tests to current models like Qwen 27B and Opus 5.5, and speculates on future advancements.
The paper introduces ExplorationBench, a benchmark for evaluating AI systems' exploration abilities in verifiable alien worlds, addressing challenges in scientific discovery by providing executable rules and preventing recall from pre-training data.
Testing Claude Opus 5 twice in one day revealed significant response differences, likely due to background settings rather than model changes, highlighting the need to document all settings when comparing AI tools.
MentalHealthBench is designed to cover the full spectrum of mental health conversations for AI, from everyday support to acute crisis scenarios, as announced by OpenAI.
DrivingBench evaluates whether frontier AI models like GPT-6 Astra can drive a real car by controlling a Toyota Corolla on a predefined course, tracking metrics such as progress, distance, and token cost.
The article explores methods for evaluating AI agents in production to decide whether to retain, improve, or shut them down, citing research on metrics like cost, reliability, human effort, and business outcomes.
This paper examines the limitations of general LLM rankings, including benchmark saturation and commercial incentives, and advocates for task-specific evaluation methods to improve reliability.
CraftBench-UE presents a deterministic evaluation harness for coding agents in Unreal Engine, enabling assessment of gameplay features through build, asset, and runtime checks without relying on LLM judges.
ISA-Bench introduces a benchmark using programming games with constrained instruction sets to evaluate computational reasoning in large language models, revealing insights into model capabilities and a reasoning-execution gap.
The paper details ufakzeka-1, a 151M-parameter Turkish language model built from scratch with a total cost of about $286, describing the training pipeline, evaluation methods, and key findings on small-model training limitations.
This paper proposes ReliMap, a framework for evaluating reliability in LLM-based human behavior simulations across individual and population levels, emphasizing coordinated improvements in model capacity, profile completeness, and coverage.