ai-benchmarks

Tag

Cards List
#ai-benchmarks

Google Gemini 4 scores same as GPT 6 Astra on Artificial Analysis Benchmark, while costing 40% less.

Reddit r/singularity ↗ · 9h ago

Google's Gemini 4 matches GPT 6 Astra on the Artificial Analysis Benchmark while being 40% cheaper, highlighting its cost-performance advantage.

0 favorites 0 likes
#ai-benchmarks

LLM Pareto frontiers split by benchmark category, Sonnet 5.5 pushes last OpenAI models off the frontier

Reddit r/ArtificialInteligence ↗ · 2d ago Cached

The article discusses AI models ranked on Pareto frontiers across benchmark categories, highlighting how Claude Sonnet 5.5 has improved Anthropic's position and pushed some OpenAI models off the frontiers in terms of quality, price, and speed.

0 favorites 0 likes
#ai-benchmarks

@wenhaocha1: The Pareto frontier is misleading for general users. Price it by subscription instead of API. With cache hits, Opus 5.5…

X AI KOLs Timeline ↗ · 2d ago Cached

The author argues that Pareto frontier benchmarks mislead general users, claiming that when pricing is compared via subscription rather than API rates (with cache hits), Opus 5.5 using Claude Code at 20x is superior to every other model.

0 favorites 0 likes
#ai-benchmarks

HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases

arXiv cs.AI ↗ · 3d ago Cached

HARDEN introduces a constrained evolutionary search to generate more challenging evaluation cases for language models while preserving expected outputs, demonstrating significant accuracy reductions across benchmarks.

0 favorites 0 likes
#ai-benchmarks

Benchy: towards a universal language for task-oriented AI benchmarks

arXiv cs.AI ↗ · 3d ago Cached

Benchy introduces a semantic language and execution engine for benchmarking AI programs, standardizing benchmark definitions and execution through a canonical representation.

0 favorites 0 likes
#ai-benchmarks

WideSWE: Can Coding Agents Coordinate Changes Across Repositories?

Hugging Face Daily Papers ↗ · 4d ago Cached

The paper introduces WideSWE, a benchmark for evaluating coding agents on cross-repository tasks, based on 120 real-world tasks, showing that agent success rates vary and joint execution can improve performance.

0 favorites 0 likes
#ai-benchmarks

@charliermarsh: I fear the gap between benchmarks and capabilities is getting wider

X AI KOLs Timeline ↗ · 4d ago Cached

A Twitter discussion highlights growing concerns that AI benchmarks fail to reflect real-world model capabilities, as they test isolated tasks rather than complex, long-running prompts encountered by users.

0 favorites 0 likes
#ai-benchmarks

Nerf detector?

Reddit r/singularity ↗ · 4d ago

The content discusses the common issue of AI models degrading in performance over time after initial excitement, and questions whether longitudinal benchmark studies have been conducted to track trends across models and labs.

0 favorites 0 likes
#ai-benchmarks

@cline: Gemini 3.8 Flash is free in Cline. It scores 41 on the Artificial Analysis Intelligence Index, the highest of any model…

X AI KOLs Timeline ↗ · 6d ago Cached

Gemini 3.8 Flash is available for free in Cline, scoring 41 on the Artificial Analysis Intelligence Index, the highest in its price tier, and outperforming models like GPT-5.6 Luna and DeepSeek v4.1 Flash.

0 favorites 0 likes
#ai-benchmarks

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

arXiv cs.LG ↗ · 2026-09-24 Cached

This paper examines why tasks in agentic AI benchmarks like Terminal-Bench may fail to be solved, distinguishing between genuine difficulty and issues such as broken oracles or infrastructure failures, and emphasizes the need to validate all-fail tasks for accurate capability claims.

0 favorites 0 likes
#ai-benchmarks

can someone explain why we think a 90%+ bench is considered saturated?

Reddit r/ArtificialInteligence ↗ · 2026-09-23

The article questions why AI benchmarks are considered saturated at 90%+ accuracy and advocates for aiming at 100% or developing new benchmarks.

0 favorites 0 likes
#ai-benchmarks

@browser_use: Browser Use Bench v2 Pareto frontier got completely redrawn today > Claude Opus 5.5: 59.4 > GPT‑6 Sol medium: 66.9 (3.5…

X AI KOLs Following ↗ · 2026-09-23 Cached

The Browser Use Bench v2 benchmark has updated its Pareto frontier, showcasing GPT-6 models from OpenAI outperforming and being more cost-effective than Claude Opus 5.5 from Anthropic.

0 favorites 0 likes
#ai-benchmarks

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Hugging Face Daily Papers ↗ · 2026-09-23 Cached

This paper introduces Just-in-Time Memory (JitMem), a method for LLM agents that defers memory curation to read-time for task-adaptive payloads, demonstrating significant performance improvements over baseline methods in benchmarks like ALFWorld and WebShop.

0 favorites 0 likes
#ai-benchmarks

@anshnanda: The current benchmarks no longer work for the new models. That’s why we are seeing results like this.

X AI KOLs Timeline ↗ · 2026-09-22 Cached

The tweet by @anshnanda points out that existing AI benchmarks are not suitable for evaluating new models, explaining unexpected results in model performance.

0 favorites 0 likes
#ai-benchmarks

Benchmarks Grok 4.7, GPT 6 Astra Fable 4.1 and DeepSeek V4.1 Flash

Reddit r/artificial ↗ · 2026-09-21

Posts comprehensive benchmarks for the latest AI models, including Grok 4.7, GPT 6, Astra Fable 4.1, and DeepSeek V4.1 Flash, to provide unbiased comparisons.

0 favorites 0 likes
#ai-benchmarks

I just realized that Felony Bench is really not a great measure of AI intelligence...

Reddit r/singularity ↗ · 2026-09-19

The article critiques the Felony Bench as an inadequate measure of AI intelligence, noting that only caught AIs are included, and highlights Google's Gemini for its hacking capabilities.

0 favorites 0 likes
#ai-benchmarks

Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs

arXiv cs.AI ↗ · 2026-09-18 Cached

This paper introduces the checkpoint handoff protocol to attribute gains in agentic reinforcement learning by separating 'Reach' (arriving at useful states) and 'Solve' (solving from those states), showing RL improvements stem from both components.

0 favorites 0 likes
#ai-benchmarks

@charliermarsh: Lots of good examples in these reports. For example: in DeepSWE v1.1, the verifier discards the agent's changes to some…

X AI KOLs Following ↗ · 2026-09-18 Cached

Epoch AI introduces Benchmark Reviews to audit AI benchmarks, revealing flaws such as in DeepSWE v1.1 where the verifier discards changes without informing the model, leading to failures.

0 favorites 0 likes
#ai-benchmarks

I could reproduce the AI benchmarks. I still couldn’t verify the business claims.

Reddit r/artificial ↗ · 2026-09-17

A researcher reproduced AI benchmarks but found that verifying business claims was challenging due to metric aggregation and model integration issues, emphasizing the need for careful scrutiny of AI performance claims.

0 favorites 0 likes
#ai-benchmarks

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

arXiv cs.AI ↗ · 2026-09-12 Cached

Benchmark Radar introduces a living database and search engine for AI benchmarks, enabling researchers to discover, retrieve, and analyze evaluation resources across various AI domains.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback