ai-benchmarks

Tag

Cards List
#ai-benchmarks

One AI just scored 1753 on a test where 'human expert' is 1000. Here's why I don't fully trust that number

Reddit r/ArtificialInteligence · 2d ago

The article critiques a viral AI benchmark that claims Grok 4.6 scored 1753 vs 1000 for human experts, highlighting that the test uses preference-based comparisons between AI outputs rather than objective correctness, so polished-looking work may win without being truly better.

0 favorites 0 likes
#ai-benchmarks

@BenjaminDEKR: Lot of people acting like Grok 4.6 just beat Anthropic and OpenAI when really, it didn't. The Grok 4.6 numbers show tha…

X AI KOLs Timeline · 3d ago Cached

Commentary on Grok 4.6 benchmark results, arguing that xAI hasn't beaten Anthropic or OpenAI but remains competitive in the middle-to-upper range.

0 favorites 0 likes
#ai-benchmarks

Goodhart's Law Comes for Every Benchmark You Trust

Hacker News Top · 2026-07-31

Discusses how Goodhart's Law undermines trust in AI benchmarks, as metrics become targets and lose their validity as evaluation measures.

0 favorites 0 likes
#ai-benchmarks

We're measuring AI with benchmarks that barely predict real world performance

Reddit r/ArtificialInteligence · 2026-07-21 Cached

A study by METR found that experienced open-source developers using AI tools (primarily Cursor Pro with Claude 3.5/3.7 Sonnet) took 19% longer to complete real-world issues, contradicting both their own expectations and expert forecasts of 24% speedup.

0 favorites 0 likes
#ai-benchmarks

DeepSeek v4 flash release version appears to have been activated on api. Open weights imminent?

Reddit r/LocalLLaMA · 2026-07-20

DeepSeek v4 flash release version appears to have been activated on API, with open weights imminent, potentially outperforming competitors in its weight class.

0 favorites 0 likes
#ai-benchmarks

KIMI K3 Beats Claude Fable and GPT 5.6 sol in arena.ai!!!

Reddit r/LocalLLaMA · 2026-07-16

KIMI K3 reportedly outperforms Claude Fable and GPT 5.6 on the arena.ai benchmark.

0 favorites 0 likes
#ai-benchmarks

Why aren't any American open-source AI labs even close to Chinese ones on benchmarks yet?

Reddit r/LocalLLaMA · 2026-07-14

The article questions why American open-source AI labs have not achieved top benchmark results like their Chinese counterparts, highlighting a perceived gap in open-source AI development between the two nations.

0 favorites 0 likes
#ai-benchmarks

@gurtej__gill_: ByteDance’s Seed team just dropped their Seed2.0 model card. Its genuinely a fascinating read for anyone tired of watch…

X AI KOLs Timeline · 2026-07-12 Cached

ByteDance's Seed team has released the Seed2.0 model card, detailing a model designed to bridge the gap between lab benchmarks and real-world software engineering. The card highlights deployment tiers, performance comparisons, and honest acknowledgment of gaps versus frontier models.

0 favorites 0 likes
#ai-benchmarks

What is the meaning of AI benchmarks?

Reddit r/artificial · 2026-07-11

A simple explanation about AI benchmarks, what scores mean, and why 100% does not mean AI cannot improve further.

0 favorites 0 likes
#ai-benchmarks

@OpenAI: We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures fr…

X AI KOLs · 2026-07-08 Cached

OpenAI audited the SWE-Bench Pro coding benchmark and found approximately 30% of tasks broken, retracting their previous recommendation for its use as a leading coding evaluation.

0 favorites 0 likes
#ai-benchmarks

A planetary test for local models

Reddit r/LocalLLaMA · 2026-07-05

A planetary-scale test for locally-run AI models is being conducted, likely benchmarking their performance across different environments.

0 favorites 0 likes
#ai-benchmarks

What happens after all AI hit % 100 on benchmarks

Reddit r/AI_Agents · 2026-07-03

The article speculates on what will happen when all AI models achieve 100% on benchmarks, questioning how they will demonstrate superiority.

0 favorites 0 likes
#ai-benchmarks

Design Arena: Gemini Omni Flash is now 1st overall on Video Arena with an Elo of 1404, and a 101 point Elo gap over Seedance 2.0 Mini.

Reddit r/singularity · 2026-07-02

Gemini Omni Flash has reached first place overall on the Video Arena leaderboard with an Elo of 1404, leading Seedance 2.0 Mini by 101 points.

0 favorites 0 likes
#ai-benchmarks

@Meer_AIIT: most "legal ai" just summarizes a pdf and calls it a day. ivo benchmarks does the thing a wrapper can't. it pulls your …

X AI KOLs Timeline · 2026-07-01 Cached

Ivo Benchmarks is a legal AI tool that goes beyond simple PDF summarization by pulling an entire company's negotiation history into contract review. It scores each clause based on past handling, enabling more informed decisions during contract negotiations.

0 favorites 0 likes
#ai-benchmarks

GLM-5.2 Raises the Bar for Open Models (14 minute read)

TLDR AI · 2026-06-23 Cached

GLM-5.2 is a new open-source AI model that sets a high bar for open models, though it still trails proprietary frontier models and lacks some features like vision.

0 favorites 0 likes
#ai-benchmarks

Did we only ever test AI when the user was ready for it

Reddit r/artificial · 2026-06-22

The article questions whether current AI benchmarks are adequate for evaluating AI in real-time, background contexts like voice calls, autonomous driving, and smart glasses, as they assume a prepared user.

0 favorites 0 likes
#ai-benchmarks

humanity's last exam current benchmarks thoughts?

Reddit r/singularity · 2026-06-15

Discussion of recent AI model scores on the 'humanity's last exam' benchmark, noting improvement from GPT-4o's 2.7% in May 2024 to around 45% by June 2026, questioning the exam's difficulty.

0 favorites 0 likes
#ai-benchmarks

Alibaba's Qwen3.7-Max Ran Autonomously for 35 Hours on Unfamiliar Hardware. It Still Kept Getting Better.

Reddit r/ArtificialInteligence · 2026-05-25 Cached

Alibaba's Qwen3.7-Max model autonomously optimized a production kernel on unfamiliar T-Head PPU hardware over 35 hours, making 1,158 tool calls and achieving a 10x speedup, demonstrating sustained autonomous agentic behavior without human guidance.

0 favorites 0 likes
#ai-benchmarks

Who decides what AI tells you? Campbell Brown, once Meta’s news chief, has thoughts

TechCrunch AI · 2026-05-14 Cached

Campbell Brown, former Meta news chief, launches Forum AI to evaluate foundation model accuracy on high-stakes topics like geopolitics and mental health, aiming to improve AI truthfulness through expert-led benchmarks.

0 favorites 0 likes
#ai-benchmarks

@aaron_epstein: New model just released that beats sonnet 4.6, gemini 3 flash, and gpt 5.4 mini on OCR, vision, and STT tasks @interfaz…

X AI KOLs Following · 2026-05-12

A new AI model from interfaze_ai claims to outperform leading models (sonnet 4.6, gemini 3 flash, gpt 5.4 mini) on OCR, vision, and speech-to-text tasks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback