ai-benchmarking

Tag

Cards List
#ai-benchmarking

Hyper-𝜏-bench: Evaluating agents that build agents (4 minute read)

TLDR AI β†— Β· 2d ago Cached

Sierra AI open-sources hyper-𝜏-bench, a new benchmark that evaluates AI models' ability to construct customer-service agents, revealing limitations in autonomous builds and improvements with human assistance.

0 favorites 0 likes
#ai-benchmarking

How many agents can 2Γ—4090 actually run at once? Three weeks of llama.cpp concurrency data β€” soft cap 5 @ 64k, hard cap 9, and why.

Reddit r/LocalLLaMA β†— Β· 2d ago

A benchmarking report on running multiple AI agents concurrently using llama.cpp on 2Γ— RTX 4090 GPUs, revealing performance limits and optimal configurations for Qwen models.

0 favorites 0 likes
#ai-benchmarking

Artificial Analysis updates its Intelligence Index to version 4.3

Reddit r/singularity β†— Β· 3d ago Cached

Artificial Analysis has updated its Intelligence Index to version 4.3, incorporating new benchmarks like Terminal-Bench v4.0 and AutomationBench-AA to better evaluate AI model performance and cost-efficiency.

0 favorites 0 likes
#ai-benchmarking

The prevalent problem of misleading benchmark reporting (re: Astra)

Reddit r/singularity β†— Β· 2026-09-03

OpenAI's benchmark reporting for Astra on ARC-AGI-3 is misleading due to using different harnesses, and the performance gap is less dramatic under standard conditions.

0 favorites 0 likes
#ai-benchmarking

@rohanpaul_ai: A strong coding model is not enough if the model managing its work does not know when to redirect, verify, or stop. Loo…

X AI KOLs Following β†— Β· 2026-09-03 Cached

LoopArena benchmarks models as runtime controllers for coding tasks, revealing that even GPT-5.5 only achieves a 24.69% success rate, emphasizing the need for better control mechanisms in agent systems.

0 favorites 0 likes
#ai-benchmarking

Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

Hacker News Top β†— Β· 2026-09-02 Cached

FrontierHarness Eval benchmarks nine software engineering harnesses on a single model, showing that cost per successful task can vary by up to 17x depending on the harness used.

0 favorites 0 likes
#ai-benchmarking

Qwen3.8-Flash-Next-NVFP4 vs Qwen3.8-27B-FP Test Results

Reddit r/LocalLLaMA β†— Β· 2026-08-31

This article presents detailed test results comparing the performance of Qwen3.8-Flash-Next-NVFP4 and Qwen3.8-27B-FP8 AI models across various tasks, highlighting that Flash-Next is faster with fewer failures but struggles with multi-step symbolic work.

0 favorites 0 likes
#ai-benchmarking

How I combined 11 coding benchmarks without averaging incompatible scores

Reddit r/ArtificialInteligence β†— Β· 2026-08-30

The author describes a method to compare coding AI models by combining 11 benchmarks using percentile ranks instead of averaging raw scores, with a focus on de-duplication and weighted categories for 98 models.

0 favorites 0 likes
#ai-benchmarking

Qwen3.8-27B thinking xhigh Vs. thinking off - Apple M5 Max

Reddit r/LocalLLaMA β†— Β· 2026-08-29

Benchmarks on a MacBook Pro M5 Max show that disabling thinking mode in Qwen3.8-27B severely degrades output quality, while xhigh thinking mode uses 5.5x more tokens and runs 6x longer.

0 favorites 0 likes
#ai-benchmarking

Run Qwen3.8 27B locally: real numbers from my Mac Studio

Hacker News Top β†— Β· 2026-08-28 Cached

The article provides real-world performance benchmarks for running the Qwen3.8 27B AI model locally on a Mac Studio, comparing it to its predecessor and discussing hardware requirements and quantization effects.

0 favorites 0 likes
#ai-benchmarking

Technical Comparative Benchmarking Study: Advanced AI Hybrid Methods for Renewable Energy Farm Optimization and Forecasting

arXiv cs.LG β†— Β· 2026-08-28 Cached

A comparative benchmarking study evaluates various AI methods for renewable energy farm optimization and forecasting, showing that ensemble and hybrid approaches excel in different data scenarios.

0 favorites 0 likes
#ai-benchmarking

How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models

arXiv cs.CL β†— Β· 2026-08-28 Cached

The paper presents a systematic cross-model evaluation of how large language models interpret verbal probability expressions, finding they track human benchmarks with fidelity but exhibit biases, particularly for negative expressions, with implications for human-AI uncertainty communication.

0 favorites 0 likes
#ai-benchmarking

Can an AI make other AIs better? We benchmarked 5 frontier LLMs at rewriting other agents' harnesses, scored on a test set they never see (HarnessOpt-Bench, arXiv + MIT code)

Reddit r/artificial β†— Β· 2026-08-27

The article introduces HarnessOpt-Bench, a benchmark for measuring how LLMs can improve other AI agents' harnesses, and presents findings from 5 frontier models, showing that model choice has a greater impact than harness choice.

0 favorites 0 likes
#ai-benchmarking

@rohanpaul_ai: The expensive part of enterprise AI is often using too much intelligence for the task. Glean is moving that decision in…

X AI KOLs Following β†— Β· 2026-08-26 Cached

Glean introduces runtime intelligence decisions to reduce enterprise AI token costs by 81%, outperforming Claude in benchmarks, and announces new features like Glean Tau and autorouting.

0 favorites 0 likes
#ai-benchmarking

Forget the Pelican, it's Weevil-Time! / Benchmaxxing-Proof SVG and Vision Benchmark

Reddit r/LocalLLaMA β†— Β· 2026-08-26

The article describes testing the Qwen3.8-27B AI model with different quantizations and settings to recreate images as SVG, aiming to develop a benchmark resistant to benchmaxxing. Preliminary results indicate that high reasoning effort and specific cache configurations optimize performance.

0 favorites 0 likes
#ai-benchmarking

Underrated Muse Glimmer

Reddit r/LocalLLaMA β†— Β· 2026-08-26

Benchmarking results reveal that Muse Glimmer surprisingly outperforms qwen3.8 in implicit knowledge tests, indicating smaller models can achieve competitive performance with RAG enhancements.

0 favorites 0 likes
#ai-benchmarking

PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation

arXiv cs.LG β†— Β· 2026-08-26 Cached

PhysicsBench introduces a unified benchmark and leaderboard to standardize the evaluation of generative and predictive AI models for engineering design and simulation across multiple tasks and data scales.

0 favorites 0 likes
#ai-benchmarking

Gemini 3.7 Flash is currently 75% off on OpenRouter, beating DeepSeek on price/performance

Reddit r/singularity β†— Β· 2026-08-22

Google has reduced prices for Gemini 3.7 Flash on OpenRouter by 75%, making it cheaper and outperforming DeepSeek models in price/performance based on Artificial Analysis.

0 favorites 0 likes
#ai-benchmarking

I benchmarked Ox Alpha on SWE-bench Verified Mini (50 tasks): 96% resolved. Now I’m skeptical of myself.

Reddit r/LocalLLaMA β†— Β· 2026-08-21

The author benchmarked Ox Alpha on SWE-bench Verified Mini, achieving a 96% resolution rate, but raises concerns about the score's validity due to data contamination, small sample size, and non-comparable baselines.

0 favorites 0 likes
#ai-benchmarking

Qwen 3.8 is off to College. Got a 34 on the ACT

Reddit r/AI_Agents β†— Β· 2026-08-20

An experiment tested the Qwen 3.8 27B AI model on ACT practice exams using vision capabilities, achieving high composite scores of 34-36, showcasing strong performance in standardized testing.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback