benchmark

Tag

Cards List
#benchmark

SCoR: A Hierarchical Framework for Forecasting Relations Between Scientific Concepts

arXiv cs.CL ↗ · 6d ago Cached

This paper introduces SCoR, a hierarchical framework for forecasting relations between scientific concepts, with a benchmark and model that improve research-direction discovery by predicting typed relations.

0 favorites 0 likes
#benchmark

Is Imagination Derived from Hallucination? A Cross-Taxonomy Evaluation of Imagination and Hallucination in Large Language Models

arXiv cs.CL ↗ · 6d ago Cached

The paper introduces Whiteboard, the first benchmark for evaluating imagination in large language models by cross-referencing it with hallucination, and reveals a counterintuitive negative correlation between the two across 79 state-of-the-art LLMs.

0 favorites 0 likes
#benchmark

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

Hugging Face Daily Papers ↗ · 6d ago Cached

HARMONY is an open-source framework using hierarchical agentic reasoning with VLMs to reconstruct compositional 3D scenes from single indoor images, competing with GPT-6 Astra in visual quality and outperforming in geometric alignment.

0 favorites 0 likes
#benchmark

RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

Hugging Face Daily Papers ↗ · 6d ago Cached

RoboFollow introduces a diagnostic benchmark to expose the illusion of instruction-following in embodied agents by analyzing high scene entropy and perturbations, revealing gaps in current models despite strong initial performance.

0 favorites 0 likes
#benchmark

The final benchmark

Reddit r/singularity ↗ · 6d ago

This article presents a definitive benchmark designed to evaluate and compare AI models or systems, establishing a standard for future assessments.

0 favorites 0 likes
#benchmark

@devindesktop: Grok 4.7 is now available in Devin Desktop and Devin CLI!

X AI KOLs Following ↗ · 6d ago Cached

Grok 4.7 AI model is now available in Devin Desktop and CLI, with evaluation results showing strong performance on hard backend engineering tasks.

0 favorites 0 likes
#benchmark

@browser_use: Grok 4.7 just dropped. Still chasing DeepSeek Long Horizon Browser Use Benchmark v2 > GPT-6 Astra: 80.6 > DeepSeek V4.1…

X AI KOLs Timeline ↗ · 6d ago Cached

Grok 4.7 has been released and shows improved performance over Grok 4.6 on the Long Horizon Browser Use Benchmark v2, but still lags significantly behind DeepSeek V4.1 Flash and GPT-6 Astra.

0 favorites 0 likes
#benchmark

@rohanpaul_ai: Forward Deployed Engineers have a compounding problem: they spend months learning a company’s systems, and hidden depen…

X AI KOLs Timeline ↗ · 2026-09-21 Cached

Codos launches a virtual Chief AI Officer that addresses context loss in Forward Deployed Engineers by deploying company-wide memory and automation agents, with preliminary benchmark scores showing high performance on EnterpriseRAG-Bench.

0 favorites 0 likes
#benchmark

@omarsar0: StepFun’s new Step 5 Preview model is impressive! Had a chance to test it early. I've been testing it as a coding agent…

X AI KOLs Following ↗ · 2026-09-21 Cached

StepFun's new Step 5 Preview model is tested as a coding agent, demonstrating competitive performance with models like GLM 5.3 and excelling in long-horizon tasks due to its effective stopping behavior.

0 favorites 0 likes
#benchmark

Putting the question before the context took my local Qwen from 89% to 100% on a decision benchmark, and from ~400 ms to ~80 ms

Reddit r/LocalLLaMA ↗ · 2026-09-21

Switching the order of question and context in prompts for local Qwen models improved accuracy from 89% to 100% and reduced latency from ~400 ms to ~80 ms on a decision benchmark.

0 favorites 0 likes
#benchmark

TypeSafe's Jev cannot emit an invalid output, but its calibration claim ships with no ECE or reliability curves

Reddit r/ArtificialInteligence ↗ · 2026-09-21

TypeSafe AI launched its Jev model claiming no hallucination and calibrated probabilities, but the article questions the lack of public evidence for calibration while noting rapid developer adoption.

0 favorites 0 likes
#benchmark

CAISI’s Assessment of Z.ai’s GLM-5.3 Cyber Capabilities

Reddit r/ArtificialInteligence ↗ · 2026-09-21 Cached

CAISI's assessment finds that Z.ai's GLM-5.3 is the most cyber-capable open-weight AI model to date, but it still trails U.S. frontier models by approximately four months in capability.

0 favorites 0 likes
#benchmark

Fifteen runs of "make one button blue" on a page built to tempt the agent stayed inside one selector every time; a 200-task stale-bug benchmark says real repos go the other way 35 to 65% of the time

Reddit r/AI_Agents ↗ · 2026-09-21

Experiments on AI agents in code editing tasks show that agents often over-edit code when bugs are already fixed, with performance varying based on task instructions. A benchmark using real repositories highlights issues with agent behavior in software development.

0 favorites 0 likes
#benchmark

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

arXiv cs.LG ↗ · 2026-09-21 Cached

OpenMAS-GCom introduces a diagnostic benchmark for evaluating graph-enhanced multi-agent systems by using controlled interventions to attribute performance differences to specific organizational components.

0 favorites 0 likes
#benchmark

Scaling Discovery through Test-Time Communication

arXiv cs.LG ↗ · 2026-09-21 Cached

The paper demonstrates that test-time communication among AI agents can significantly outperform independent parallel attempts on challenging tasks like ARC-AGI-3, achieving state-of-the-art results in research-oriented domains.

0 favorites 0 likes
#benchmark

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

arXiv cs.LG ↗ · 2026-09-21 Cached

This paper introduces BI-Bench, the first benchmark for evaluating LLMs on end-to-end business intelligence tasks, and BI-Agent, a tool-augmented agent that decomposes workflows and uses post-training to improve accuracy significantly.

0 favorites 0 likes
#benchmark

Chinese Competitive Debating Dataset and Benchmark

arXiv cs.CL ↗ · 2026-09-21 Cached

This paper introduces a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate, featuring 148 matches with professional adjudication and tasks at match, stage, and speaker levels.

0 favorites 0 likes
#benchmark

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

arXiv cs.CL ↗ · 2026-09-21 Cached

This paper introduces Omni Demand Understanding (ODU), a benchmark to evaluate how well multimodal AI models infer user demands from complex audio-visual interactions, revealing significant performance gaps in current models.

0 favorites 0 likes
#benchmark

$\mu^2$-Bench: A Multilingual Machine Unlearning Benchmark

arXiv cs.CL ↗ · 2026-09-21 Cached

This paper introduces μ²-Bench, a benchmark for evaluating multilingual machine unlearning in large language models, aiming to ensure that undesired information is effectively removed across diverse languages.

0 favorites 0 likes
#benchmark

MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs

arXiv cs.CL ↗ · 2026-09-21 Cached

MME-Safety is a rigorously verified benchmark for evaluating the safety of Multimodal Large Language Models, featuring a four-dimensional annotation schema and a hierarchical framework to assess risk scenarios, harm severity, and modality-specific stealth levels.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback