llm-benchmarks

Tag

Cards List
#llm-benchmarks

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

Reddit r/LocalLLaMA · 22h ago

The author shares hands-on comparisons showing Gemma 4 outperforming larger models like Gemini 3.5 Flash and Claude Opus 5 on practical instruction-following, arguing that current LLM benchmarks fail to capture real-world usability.

0 favorites 0 likes
#llm-benchmarks

Quantifying Ranking Uncertainty in LLM Benchmarks

arXiv cs.LG · 2026-07-21 Cached

This paper analyzes sources of ranking uncertainty in LLM benchmarks like MMLU, proposing modifications to hypothesis tests for constructing rank confidence intervals, and shows that variability across subjects is substantial.

0 favorites 0 likes
#llm-benchmarks

I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM

Reddit r/LocalLLaMA · 2026-07-20

A user tested quantized 1-bit and 2-bit versions of the 27B-parameter Bonsai model on Terminal-Bench 2.0, achieving results within 8GB VRAM.

0 favorites 0 likes
#llm-benchmarks

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

arXiv cs.AI · 2026-07-14 Cached

This paper proposes a submodular coreset selection method for LLM benchmarks that selects a subset of prompts without using model evaluation outcomes, achieving score preservation across 35 benchmarks and 18 LLMs.

0 favorites 0 likes
#llm-benchmarks

Is LM Arena over?

Reddit r/LocalLLaMA · 2026-07-10

A user criticizes LM Arena for not including recent open-source models like Qwen3.6 and Step 3.7 Flash, questioning the platform's relevance for comparing new open models against closed ones.

0 favorites 0 likes
#llm-benchmarks

Every AI Visibility Tool Is Lying to You

Hacker News Top · 2026-07-03 Cached

This article critically examines the accuracy of AI visibility tools that claim to measure brand presence in generative AI responses, arguing that they provide false precision due to nondeterminism, personalization, and scraping biases. It calls for transparency in methodology and warns against treating opaque dashboards as stable truth.

0 favorites 0 likes
#llm-benchmarks

The gap between open weights LLMs and closed source LLMs

Hacker News Top · 2026-06-26 Cached

Analyzes the gap between open weights and closed source LLMs using the Artificial Analysis Intelligence Index and other benchmarks, finding that the gap is shrinking on some metrics but stable on others.

0 favorites 0 likes
#llm-benchmarks

@ms_aifrontiers: Most LLM benchmark scores are predictable before you ever run them. New from the MS AI Frontiers team: BenchPress. The …

X AI KOLs Following · 2026-06-25 Cached

The MS AI Frontiers team introduces BenchPress, a method that uses matrix completion to predict LLM benchmark scores from just five probes, showing the score matrix is effectively rank-2.

0 favorites 0 likes
#llm-benchmarks

Benchmarks from the latest eBay special: W6800 (modded V620)

Reddit r/LocalLLaMA · 2026-06-17

A user benchmarks a modded AMD V620 GPU flashed with W6800 firmware and a custom blower fan for running LLMs via Vulkan and ROCm backends, comparing performance on Qwen2.5-27B at various quantization levels.

0 favorites 0 likes
#llm-benchmarks

These LLMs are the best at resisting Russian propaganda

Ars Technica · 2026-06-04 Cached

A benchmark study by the Estonian Language Institute evaluates LLMs on their ability to resist Russian propaganda, finding that Nvidia's Nemotron, Alibaba's Qwen, and OpenAI's GPT-5.4 perform well, while Google's Gemini models show notable weaknesses, especially when prompted in Russian.

0 favorites 0 likes
#llm-benchmarks

Auditing LLM Benchmarks with Item Response Theory

arXiv cs.CL · 2026-06-01 Cached

This paper introduces an Item Response Theory-based method to detect mislabeled examples in LLM benchmarks at 95% precision, tracing errors to labeling heuristics and annotation issues.

0 favorites 0 likes
#llm-benchmarks

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks

arXiv cs.AI · 2026-05-26 Cached

This paper identifies systemic measurement bias in production LLM inference benchmarks caused by single-process Python clients using asyncio, and proposes a multi-process evaluation framework and a new metric (NTPOT) to accurately profile serving engines at scale.

0 favorites 0 likes
#llm-benchmarks

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation

arXiv cs.AI · 2026-05-11 Cached

This paper introduces EnvSimBench, a benchmark for evaluating Large Language Models' ability to simulate environments for agent training. It identifies a 'state change cliff' in current LLMs and proposes a constraint-driven pipeline to reduce hallucinations and costs.

0 favorites 0 likes
#llm-benchmarks

@omarsar0: Cool paper from Apple. Most evaluation of tool-calling agents happens after the trajectory is over. By then the wrong c…

X AI KOLs Timeline · 2026-05-10 Cached

This Apple research paper introduces 'Reinforced Agent,' a method that moves evaluation into the execution loop using a specialized reviewer agent to correct tool-calling errors in real-time. It demonstrates significant accuracy improvements on benchmarks like BFCL and τ²-Bench without retraining the base agent.

0 favorites 0 likes
← Back to home

Submit Feedback