llm-benchmarks

Tag

Cards List
#llm-benchmarks

@theo: Updated run just finished. GPT-6.1 Sol still crushes. Turns out it performs WAY better in Codex than in mini-swe (what …

X AI KOLs Timeline ↗ · yesterday Cached

Theo (t3.gg) reports that GPT-6.1 Sol performs significantly better in Codex than in mini-swe benchmarks used by Artificial Analysis, achieving Terminal Bench 4 scores better than Opus 5.5 at roughly 1/30th the price.

0 favorites 0 likes
#llm-benchmarks

A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance

arXiv cs.AI ↗ · 2026-09-16 Cached

This paper presents a framework for generating context-specific large language model benchmark datasets using expert guidance and synthetic data, improving validity and scalability over existing methods.

0 favorites 0 likes
#llm-benchmarks

Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage

arXiv cs.CL ↗ · 2026-09-14 Cached

This paper identifies surface-level feature leakage in truthfulness benchmarks like TruthfulQA, where models can cheat by exploiting answer form differences, and introduces Audit-Prune to clean benchmarks, ensuring more reliable evaluations.

0 favorites 0 likes
#llm-benchmarks

BenchMIRT: What are LLM benchmarks actually measuring?

Hugging Face Blog ↗ · 2026-09-01 Cached

BenchMIRT introduces a method to audit LLM benchmarks at the individual prompt level using multidimensional item response theory, separating underlying capabilities like safety and general reasoning to reveal what benchmarks actually measure.

0 favorites 0 likes
#llm-benchmarks

All currently popular local models in one table + Opus 4.8 results

Reddit r/LocalLLaMA ↗ · 2026-09-01

This article provides a comparative table of popular local AI models across various benchmarks, including agentic, coding, general, and multimodal tasks, to help users choose models based on hardware specs and use cases.

0 favorites 0 likes
#llm-benchmarks

Qwen-3.8-27B, Nemotron-3.5-Lightning-30B-A3B, Ornith-1.5-35B-A3B, Muse-Glimmer-30B oQ8e comparison

Reddit r/LocalLLaMA ↗ · 2026-08-24

A comparison of multiple AI models including Qwen, Nemotron, Ornith, and Muse-Glimmer on benchmarks, with Ornith performing well and TielCoder showing potential in coding tasks.

0 favorites 0 likes
#llm-benchmarks

Why your local LLM feels dumber than it is

Hacker News Top ↗ · 2026-08-22 Cached

The article explores why locally run large language models might seem less intelligent, addressing potential performance or perception issues.

0 favorites 0 likes
#llm-benchmarks

What Happens When the Cost of Intelligence Drops 100x

Hacker News Top ↗ · 2026-08-21 Cached

The article analyzes the rapid decline in AI intelligence costs, with a 56x drop in six months, and discusses implications for scaling AI applications and future trends.

0 favorites 0 likes
#llm-benchmarks

Qwen 3.8 - 27B is a game changer

Reddit r/LocalLLaMA ↗ · 2026-08-15

The article highlights the exceptional performance of the Qwen 3.8 27B model in cybersecurity tasks, particularly malware analysis, surpassing previous models like Opus. It discusses benchmarks and implications for AI capabilities in exploiting vulnerabilities.

0 favorites 0 likes
#llm-benchmarks

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper introduces BenchDrift, a method for quantifying how LLM benchmark performance changes when problems are rephrased without changing meaning or answer. It shows that rephrasing causes bidirectional correctness flips across models and benchmarks, with stronger models becoming more sensitive to phrasing.

0 favorites 0 likes
#llm-benchmarks

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Hacker News Top ↗ · 2026-08-04 Cached

A systematic study defining and analyzing benchmark saturation across 60 language model benchmarks, finding nearly half exhibit saturation and that expert-curated benchmarks are more resilient, suggesting design choices for durable evaluation.

0 favorites 0 likes
#llm-benchmarks

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

Reddit r/LocalLLaMA ↗ · 2026-07-31

The author shares hands-on comparisons showing Gemma 4 outperforming larger models like Gemini 3.5 Flash and Claude Opus 5 on practical instruction-following, arguing that current LLM benchmarks fail to capture real-world usability.

0 favorites 0 likes
#llm-benchmarks

Quantifying Ranking Uncertainty in LLM Benchmarks

arXiv cs.LG ↗ · 2026-07-21 Cached

This paper analyzes sources of ranking uncertainty in LLM benchmarks like MMLU, proposing modifications to hypothesis tests for constructing rank confidence intervals, and shows that variability across subjects is substantial.

0 favorites 0 likes
#llm-benchmarks

I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM

Reddit r/LocalLLaMA ↗ · 2026-07-20

A user tested quantized 1-bit and 2-bit versions of the 27B-parameter Bonsai model on Terminal-Bench 2.0, achieving results within 8GB VRAM.

0 favorites 0 likes
#llm-benchmarks

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

arXiv cs.AI ↗ · 2026-07-14 Cached

This paper proposes a submodular coreset selection method for LLM benchmarks that selects a subset of prompts without using model evaluation outcomes, achieving score preservation across 35 benchmarks and 18 LLMs.

0 favorites 0 likes
#llm-benchmarks

Is LM Arena over?

Reddit r/LocalLLaMA ↗ · 2026-07-10

A user criticizes LM Arena for not including recent open-source models like Qwen3.6 and Step 3.7 Flash, questioning the platform's relevance for comparing new open models against closed ones.

0 favorites 0 likes
#llm-benchmarks

Every AI Visibility Tool Is Lying to You

Hacker News Top ↗ · 2026-07-03 Cached

This article critically examines the accuracy of AI visibility tools that claim to measure brand presence in generative AI responses, arguing that they provide false precision due to nondeterminism, personalization, and scraping biases. It calls for transparency in methodology and warns against treating opaque dashboards as stable truth.

0 favorites 0 likes
#llm-benchmarks

The gap between open weights LLMs and closed source LLMs

Hacker News Top ↗ · 2026-06-26 Cached

Analyzes the gap between open weights and closed source LLMs using the Artificial Analysis Intelligence Index and other benchmarks, finding that the gap is shrinking on some metrics but stable on others.

0 favorites 0 likes
#llm-benchmarks

@ms_aifrontiers: Most LLM benchmark scores are predictable before you ever run them. New from the MS AI Frontiers team: BenchPress. The …

X AI KOLs Following ↗ · 2026-06-25 Cached

The MS AI Frontiers team introduces BenchPress, a method that uses matrix completion to predict LLM benchmark scores from just five probes, showing the score matrix is effectively rank-2.

0 favorites 0 likes
#llm-benchmarks

Benchmarks from the latest eBay special: W6800 (modded V620)

Reddit r/LocalLLaMA ↗ · 2026-06-17

A user benchmarks a modded AMD V620 GPU flashed with W6800 firmware and a custom blower fan for running LLMs via Vulkan and ROCm backends, comparing performance on Qwen2.5-27B at various quantization levels.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback