ai-benchmarking

Tag

Cards List
#ai-benchmarking

@bnjmn_marie: If you retry a lot, Qwen3.8 27B can reach 92.04% on DeepSWE 1.1, ~18 pts above GPT-6 Astra’s reported ~74%. I ran 20 co…

X AI KOLs Timeline · 2d ago Cached

Qwen3.8 27B achieves 92.04% on DeepSWE 1.1 with retries, outperforming GPT-6 Astra's 74%, demonstrating smaller models' potential with reliability strategies.

0 favorites 0 likes
#ai-benchmarking

Artificial Analysis Capability Indices v1.1 (2 minute read)

TLDR AI · 3d ago Cached

Artificial Analysis has released Intelligence Index v4.2, an interim update with new evaluations like AA-Briefcase and GDP.pdf, increased private test sets to prevent gaming, and key results showing Anthropic and OpenAI leading.

0 favorites 0 likes
#ai-benchmarking

Animated transition from AA Intelligence Index v4.1 to v4.3

Reddit r/LocalLLaMA · 3d ago

This article analyzes the transition of the AA Intelligence Index from v4.1 to v4.3, detailing how changes in benchmark weights affect AI model intelligence scores and cost-effectiveness, with significant improvements for models like GPT-6 Astra.

0 favorites 0 likes
#ai-benchmarking

Are terminal compression tools actually saving us money?

Reddit r/AI_Agents · 6d ago

Research testing terminal compression tools across multiple AI model runs shows that token savings do not lead to significant cost reductions, highlighting that token compression is not equivalent to cost optimization.

0 favorites 0 likes
#ai-benchmarking

Hyper-𝜏-bench: Evaluating agents that build agents (4 minute read)

TLDR AI · 2026-09-09 Cached

Sierra AI open-sources hyper-𝜏-bench, a new benchmark that evaluates AI models' ability to construct customer-service agents, revealing limitations in autonomous builds and improvements with human assistance.

0 favorites 0 likes
#ai-benchmarking

How many agents can 2×4090 actually run at once? Three weeks of llama.cpp concurrency data — soft cap 5 @ 64k, hard cap 9, and why.

Reddit r/LocalLLaMA · 2026-09-08

A benchmarking report on running multiple AI agents concurrently using llama.cpp on 2× RTX 4090 GPUs, revealing performance limits and optimal configurations for Qwen models.

0 favorites 0 likes
#ai-benchmarking

Artificial Analysis updates its Intelligence Index to version 4.3

Reddit r/singularity · 2026-09-07 Cached

Artificial Analysis has updated its Intelligence Index to version 4.3, incorporating new benchmarks like Terminal-Bench v4.0 and AutomationBench-AA to better evaluate AI model performance and cost-efficiency.

0 favorites 0 likes
#ai-benchmarking

The prevalent problem of misleading benchmark reporting (re: Astra)

Reddit r/singularity · 2026-09-03

OpenAI's benchmark reporting for Astra on ARC-AGI-3 is misleading due to using different harnesses, and the performance gap is less dramatic under standard conditions.

0 favorites 0 likes
#ai-benchmarking

@rohanpaul_ai: A strong coding model is not enough if the model managing its work does not know when to redirect, verify, or stop. Loo…

X AI KOLs Following · 2026-09-03 Cached

LoopArena benchmarks models as runtime controllers for coding tasks, revealing that even GPT-5.5 only achieves a 24.69% success rate, emphasizing the need for better control mechanisms in agent systems.

0 favorites 0 likes
#ai-benchmarking

Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

Hacker News Top · 2026-09-02 Cached

FrontierHarness Eval benchmarks nine software engineering harnesses on a single model, showing that cost per successful task can vary by up to 17x depending on the harness used.

0 favorites 0 likes
#ai-benchmarking

Qwen3.8-Flash-Next-NVFP4 vs Qwen3.8-27B-FP Test Results

Reddit r/LocalLLaMA · 2026-08-31

This article presents detailed test results comparing the performance of Qwen3.8-Flash-Next-NVFP4 and Qwen3.8-27B-FP8 AI models across various tasks, highlighting that Flash-Next is faster with fewer failures but struggles with multi-step symbolic work.

0 favorites 0 likes
#ai-benchmarking

How I combined 11 coding benchmarks without averaging incompatible scores

Reddit r/ArtificialInteligence · 2026-08-30

The author describes a method to compare coding AI models by combining 11 benchmarks using percentile ranks instead of averaging raw scores, with a focus on de-duplication and weighted categories for 98 models.

0 favorites 0 likes
#ai-benchmarking

Qwen3.8-27B thinking xhigh Vs. thinking off - Apple M5 Max

Reddit r/LocalLLaMA · 2026-08-29

Benchmarks on a MacBook Pro M5 Max show that disabling thinking mode in Qwen3.8-27B severely degrades output quality, while xhigh thinking mode uses 5.5x more tokens and runs 6x longer.

0 favorites 0 likes
#ai-benchmarking

Run Qwen3.8 27B locally: real numbers from my Mac Studio

Hacker News Top · 2026-08-28 Cached

The article provides real-world performance benchmarks for running the Qwen3.8 27B AI model locally on a Mac Studio, comparing it to its predecessor and discussing hardware requirements and quantization effects.

0 favorites 0 likes
#ai-benchmarking

Technical Comparative Benchmarking Study: Advanced AI Hybrid Methods for Renewable Energy Farm Optimization and Forecasting

arXiv cs.LG · 2026-08-28 Cached

A comparative benchmarking study evaluates various AI methods for renewable energy farm optimization and forecasting, showing that ensemble and hybrid approaches excel in different data scenarios.

0 favorites 0 likes
#ai-benchmarking

How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models

arXiv cs.CL · 2026-08-28 Cached

The paper presents a systematic cross-model evaluation of how large language models interpret verbal probability expressions, finding they track human benchmarks with fidelity but exhibit biases, particularly for negative expressions, with implications for human-AI uncertainty communication.

0 favorites 0 likes
#ai-benchmarking

Can an AI make other AIs better? We benchmarked 5 frontier LLMs at rewriting other agents' harnesses, scored on a test set they never see (HarnessOpt-Bench, arXiv + MIT code)

Reddit r/artificial · 2026-08-27

The article introduces HarnessOpt-Bench, a benchmark for measuring how LLMs can improve other AI agents' harnesses, and presents findings from 5 frontier models, showing that model choice has a greater impact than harness choice.

0 favorites 0 likes
#ai-benchmarking

@rohanpaul_ai: The expensive part of enterprise AI is often using too much intelligence for the task. Glean is moving that decision in…

X AI KOLs Following · 2026-08-26 Cached

Glean introduces runtime intelligence decisions to reduce enterprise AI token costs by 81%, outperforming Claude in benchmarks, and announces new features like Glean Tau and autorouting.

0 favorites 0 likes
#ai-benchmarking

Forget the Pelican, it's Weevil-Time! / Benchmaxxing-Proof SVG and Vision Benchmark

Reddit r/LocalLLaMA · 2026-08-26

The article describes testing the Qwen3.8-27B AI model with different quantizations and settings to recreate images as SVG, aiming to develop a benchmark resistant to benchmaxxing. Preliminary results indicate that high reasoning effort and specific cache configurations optimize performance.

0 favorites 0 likes
#ai-benchmarking

Underrated Muse Glimmer

Reddit r/LocalLLaMA · 2026-08-26

Benchmarking results reveal that Muse Glimmer surprisingly outperforms qwen3.8 in implicit knowledge tests, indicating smaller models can achieve competitive performance with RAG enhancements.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback