benchmark

Tag

Cards List
#benchmark

SWE-Bench Pro V2 (9 minute read)

TLDR AI ↗ · 5d ago Cached

SWE-Bench Pro V2 is an updated benchmark for evaluating AI agents in software engineering, featuring 642 tasks across 11 repositories with improved evaluation protocols and contamination controls.

0 favorites 0 likes
#benchmark

@jerryjliu0: Opus 5.5 is the best frontier model for parsing tables in PDFs. We ran it through ParseBench, and it scored 93.9%, a 7%…

X AI KOLs Timeline ↗ · 5d ago Cached

Opus 5.5 scores 93.9% on ParseBench for table parsing in PDFs, outperforming Opus 5 and other models, but is expensive compared to LlamaParse.

0 favorites 0 likes
#benchmark

M5 ultra AI test results

Reddit r/LocalLLaMA ↗ · 5d ago

M5 ultra AI test results show prompt processing speeds up to 4-4.5 times faster and token generation 1.5x faster than M3 ultra, but with doubled power consumption, increased fan noise, and higher temperatures.

0 favorites 0 likes
#benchmark

LLM Ass Bench

Hacker News Top ↗ · 5d ago Cached

LLM Ass Bench is a benchmark or tool for evaluating Large Language Models, with a focus on prompts.

0 favorites 0 likes
#benchmark

Every Benchmark Comparison for Opus 5.5 and GPT-6 Astra/Sol

Reddit r/singularity ↗ · 5d ago

Opus 5.5 outperforms GPT-6 Astra and Sol in benchmark comparisons, but at a higher cost, as shown in the provided image.

0 favorites 0 likes
#benchmark

Opus 5.5 Cost vs Performance on Terminal-Bench 4.0

Reddit r/singularity ↗ · 6d ago

This article evaluates the cost-effectiveness and performance of the Opus 5.5 AI model on the Terminal-Bench 4.0 benchmark.

0 favorites 0 likes
#benchmark

Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index, along with a 20% price cut and larger cache hit discount

Reddit r/ArtificialInteligence ↗ · 6d ago

Claude Opus 5.5 has achieved the top position on the Artificial Analysis Intelligence Index, along with a 20% price reduction and enhanced cache hit discounts.

0 favorites 0 likes
#benchmark

@browser_use: Open Source MiMo v2.6 models are the new Pareto frontier > MiMo v2.6 Flash: 37.6 > Grok 4.7’s 39.9 > 57x lower recorded…

X AI KOLs Following ↗ · 6d ago Cached

Open Source MiMo v2.6 models are claimed to establish a new Pareto frontier, offering 57x lower cost than Grok 4.7 and 27% cheaper than DeepSeek v4.1 Flash in benchmarks.

0 favorites 0 likes
#benchmark

LinearSolveBench: new benchmark for linear solvers [P]

Reddit r/MachineLearning ↗ · 6d ago

LinearSolveBench is a new benchmark for evaluating linear solvers in C, focusing on speed, accuracy, and generality for large sparse systems to promote algorithmic advances.

0 favorites 0 likes
#benchmark

Show HN: JevBench, a reproducible benchmark for typed decision models

Hacker News Top ↗ · 6d ago Cached

JevBench v1.3.0 is a reproducible benchmark for Jev-class decision models, evaluating and ranking 52 systems based on intelligence, calibration, speed, and cost.

0 favorites 0 likes
#benchmark

GPT-6 Astra vs. Anthropic’s new Mythos do the Pelican SVG test

Reddit r/singularity ↗ · 6d ago

A tweet compares GPT-6 Astra and Anthropic’s new Mythos model on the Pelican SVG test, with a link to an external source.

0 favorites 0 likes
#benchmark

@kentcdodds: this is not the direction you want to see models go generally

X AI KOLs Timeline ↗ · 6d ago Cached

Cognition releases Grok 4.7 in Devin, achieving a 59.4% score on FrontierCode 1.1 Extended tasks, demonstrating strong performance in backend engineering.

0 favorites 0 likes
#benchmark

MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis

Hacker News Top ↗ · 6d ago Cached

An analysis of the MiMo-v2.6-Pro AI model's intelligence, performance, and price using Artificial Analysis's benchmarks and indexes.

0 favorites 0 likes
#benchmark

Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark

arXiv cs.LG ↗ · 6d ago Cached

This study benchmarks coding agents on reproducing Eurostat statistics, finding that semantic validation and a retry budget are crucial for reliability, not just execution diagnostics.

0 favorites 0 likes
#benchmark

Dissecting Training-Free Uncertainty Estimation in Multimodal Large Language Models

arXiv cs.CL ↗ · 6d ago Cached

A systematic study benchmarking training-free uncertainty quantification strategies for multimodal Large Language Models, categorizing methods into token-level, verbalized, and semantic approaches and finding optimal strategies depend on response length.

0 favorites 0 likes
#benchmark

PII-TRACE: A Benchmark for Context-Aware PII Detection in Multi-Turn LLM Conversations

arXiv cs.CL ↗ · 6d ago Cached

Introduces PII-TRACE, the first benchmark for context-aware PII detection in multi-turn LLM conversations, and PII-Tracer, a compact detector that achieves high entity-level coverage.

0 favorites 0 likes
#benchmark

SCoR: A Hierarchical Framework for Forecasting Relations Between Scientific Concepts

arXiv cs.CL ↗ · 6d ago Cached

This paper introduces SCoR, a hierarchical framework for forecasting relations between scientific concepts, with a benchmark and model that improve research-direction discovery by predicting typed relations.

0 favorites 0 likes
#benchmark

Is Imagination Derived from Hallucination? A Cross-Taxonomy Evaluation of Imagination and Hallucination in Large Language Models

arXiv cs.CL ↗ · 6d ago Cached

The paper introduces Whiteboard, the first benchmark for evaluating imagination in large language models by cross-referencing it with hallucination, and reveals a counterintuitive negative correlation between the two across 79 state-of-the-art LLMs.

0 favorites 0 likes
#benchmark

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

Hugging Face Daily Papers ↗ · 6d ago Cached

HARMONY is an open-source framework using hierarchical agentic reasoning with VLMs to reconstruct compositional 3D scenes from single indoor images, competing with GPT-6 Astra in visual quality and outperforming in geometric alignment.

0 favorites 0 likes
#benchmark

RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

Hugging Face Daily Papers ↗ · 6d ago Cached

RoboFollow introduces a diagnostic benchmark to expose the illusion of instruction-following in embodied agents by analyzing high scene entropy and perturbations, revealing gaps in current models despite strong initial performance.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback