Tag
SWE-Bench Pro V2 is an updated benchmark for evaluating AI agents in software engineering, featuring 642 tasks across 11 repositories with improved evaluation protocols and contamination controls.
Opus 5.5 scores 93.9% on ParseBench for table parsing in PDFs, outperforming Opus 5 and other models, but is expensive compared to LlamaParse.
M5 ultra AI test results show prompt processing speeds up to 4-4.5 times faster and token generation 1.5x faster than M3 ultra, but with doubled power consumption, increased fan noise, and higher temperatures.
LLM Ass Bench is a benchmark or tool for evaluating Large Language Models, with a focus on prompts.
Opus 5.5 outperforms GPT-6 Astra and Sol in benchmark comparisons, but at a higher cost, as shown in the provided image.
This article evaluates the cost-effectiveness and performance of the Opus 5.5 AI model on the Terminal-Bench 4.0 benchmark.
Claude Opus 5.5 has achieved the top position on the Artificial Analysis Intelligence Index, along with a 20% price reduction and enhanced cache hit discounts.
Open Source MiMo v2.6 models are claimed to establish a new Pareto frontier, offering 57x lower cost than Grok 4.7 and 27% cheaper than DeepSeek v4.1 Flash in benchmarks.
LinearSolveBench is a new benchmark for evaluating linear solvers in C, focusing on speed, accuracy, and generality for large sparse systems to promote algorithmic advances.
JevBench v1.3.0 is a reproducible benchmark for Jev-class decision models, evaluating and ranking 52 systems based on intelligence, calibration, speed, and cost.
A tweet compares GPT-6 Astra and Anthropic’s new Mythos model on the Pelican SVG test, with a link to an external source.
Cognition releases Grok 4.7 in Devin, achieving a 59.4% score on FrontierCode 1.1 Extended tasks, demonstrating strong performance in backend engineering.
An analysis of the MiMo-v2.6-Pro AI model's intelligence, performance, and price using Artificial Analysis's benchmarks and indexes.
This study benchmarks coding agents on reproducing Eurostat statistics, finding that semantic validation and a retry budget are crucial for reliability, not just execution diagnostics.
A systematic study benchmarking training-free uncertainty quantification strategies for multimodal Large Language Models, categorizing methods into token-level, verbalized, and semantic approaches and finding optimal strategies depend on response length.
Introduces PII-TRACE, the first benchmark for context-aware PII detection in multi-turn LLM conversations, and PII-Tracer, a compact detector that achieves high entity-level coverage.
This paper introduces SCoR, a hierarchical framework for forecasting relations between scientific concepts, with a benchmark and model that improve research-direction discovery by predicting typed relations.
The paper introduces Whiteboard, the first benchmark for evaluating imagination in large language models by cross-referencing it with hallucination, and reveals a counterintuitive negative correlation between the two across 79 state-of-the-art LLMs.
HARMONY is an open-source framework using hierarchical agentic reasoning with VLMs to reconstruct compositional 3D scenes from single indoor images, competing with GPT-6 Astra in visual quality and outperforming in geometric alignment.
RoboFollow introduces a diagnostic benchmark to expose the illusion of instruction-following in embodied agents by analyzing high scene entropy and perturbations, revealing gaps in current models despite strong initial performance.