Tag
This paper introduces SCoR, a hierarchical framework for forecasting relations between scientific concepts, with a benchmark and model that improve research-direction discovery by predicting typed relations.
The paper introduces Whiteboard, the first benchmark for evaluating imagination in large language models by cross-referencing it with hallucination, and reveals a counterintuitive negative correlation between the two across 79 state-of-the-art LLMs.
HARMONY is an open-source framework using hierarchical agentic reasoning with VLMs to reconstruct compositional 3D scenes from single indoor images, competing with GPT-6 Astra in visual quality and outperforming in geometric alignment.
RoboFollow introduces a diagnostic benchmark to expose the illusion of instruction-following in embodied agents by analyzing high scene entropy and perturbations, revealing gaps in current models despite strong initial performance.
This article presents a definitive benchmark designed to evaluate and compare AI models or systems, establishing a standard for future assessments.
Grok 4.7 AI model is now available in Devin Desktop and CLI, with evaluation results showing strong performance on hard backend engineering tasks.
Grok 4.7 has been released and shows improved performance over Grok 4.6 on the Long Horizon Browser Use Benchmark v2, but still lags significantly behind DeepSeek V4.1 Flash and GPT-6 Astra.
Codos launches a virtual Chief AI Officer that addresses context loss in Forward Deployed Engineers by deploying company-wide memory and automation agents, with preliminary benchmark scores showing high performance on EnterpriseRAG-Bench.
StepFun's new Step 5 Preview model is tested as a coding agent, demonstrating competitive performance with models like GLM 5.3 and excelling in long-horizon tasks due to its effective stopping behavior.
Switching the order of question and context in prompts for local Qwen models improved accuracy from 89% to 100% and reduced latency from ~400 ms to ~80 ms on a decision benchmark.
TypeSafe AI launched its Jev model claiming no hallucination and calibrated probabilities, but the article questions the lack of public evidence for calibration while noting rapid developer adoption.
CAISI's assessment finds that Z.ai's GLM-5.3 is the most cyber-capable open-weight AI model to date, but it still trails U.S. frontier models by approximately four months in capability.
Experiments on AI agents in code editing tasks show that agents often over-edit code when bugs are already fixed, with performance varying based on task instructions. A benchmark using real repositories highlights issues with agent behavior in software development.
OpenMAS-GCom introduces a diagnostic benchmark for evaluating graph-enhanced multi-agent systems by using controlled interventions to attribute performance differences to specific organizational components.
The paper demonstrates that test-time communication among AI agents can significantly outperform independent parallel attempts on challenging tasks like ARC-AGI-3, achieving state-of-the-art results in research-oriented domains.
This paper introduces BI-Bench, the first benchmark for evaluating LLMs on end-to-end business intelligence tasks, and BI-Agent, a tool-augmented agent that decomposes workflows and uses post-training to improve accuracy significantly.
This paper introduces a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate, featuring 148 matches with professional adjudication and tasks at match, stage, and speaker levels.
This paper introduces Omni Demand Understanding (ODU), a benchmark to evaluate how well multimodal AI models infer user demands from complex audio-visual interactions, revealing significant performance gaps in current models.
This paper introduces μ²-Bench, a benchmark for evaluating multilingual machine unlearning in large language models, aiming to ensure that undesired information is effectively removed across diverse languages.
MME-Safety is a rigorously verified benchmark for evaluating the safety of Multimodal Large Language Models, featuring a four-dimensional annotation schema and a hierarchical framework to assess risk scenarios, harm severity, and modality-specific stealth levels.