model-evaluation

Tag

Cards List
#model-evaluation

The Asymmetric Harms of LLM Compression

arXiv cs.CL ↗ · 2026-08-21 Cached

This paper systematically evaluates LLM compression methods and reveals asymmetric harms, such as disproportionate loss of head knowledge and increased confidence in incorrect answers, which are masked by standard aggregate metrics.

0 favorites 0 likes
#model-evaluation

The Marshmallow AI Benchmark

Reddit r/singularity ↗ · 2026-08-21

A benchmark test evaluates various AI models by prompting them to count marshmallows in an image, with results ranging from 472 to 539 counts.

0 favorites 0 likes
#model-evaluation

Qwen 3.8 is off to College. Got a 34 on the ACT

Reddit r/AI_Agents ↗ · 2026-08-20

An experiment tested the Qwen 3.8 27B AI model on ACT practice exams using vision capabilities, achieving high composite scores of 34-36, showcasing strong performance in standardized testing.

0 favorites 0 likes
#model-evaluation

Towards Quantifying Benchmark Optimization in ASR Models

Hugging Face Daily Papers ↗ · 2026-08-20 Cached

This paper quantifies how high-performing ASR models optimize for benchmarks in ways that inflate scores without improving real-world transcription, using behavioral probes to reveal benchmark-conditioned behaviors.

0 favorites 0 likes
#model-evaluation

Fluid Simulation Qwen3.8 27B IQ3_XXS

Reddit r/LocalLLaMA ↗ · 2026-08-19

The author tested the Qwen3.8 27B model on a fluid simulation task, achieving success after iterative prompting, and shared the implementation with traces and a GitHub repository.

0 favorites 0 likes
#model-evaluation

Am I doing something wrong? Qwen 3.8 27B seems useless for agentic coding

Reddit r/LocalLLaMA ↗ · 2026-08-19

A user reports difficulties using the Qwen 3.8 27B model for agentic coding tasks, noting inefficiencies and errors compared to other models, and seeks advice on potential setup issues.

0 favorites 0 likes
#model-evaluation

The Price of Thinking: Reasoning Effort as a Model-Specific API Contract

arXiv cs.AI ↗ · 2026-08-19 Cached

This paper investigates the impact of reasoning effort parameters in AI model API contracts, showing that explicit high effort leads to higher costs without detectable accuracy improvements based on experiments with Sonnet 5.

0 favorites 0 likes
#model-evaluation

GLM5.3 Artificial Analysis Benchmarks

Reddit r/LocalLLaMA ↗ · 2026-08-18 Cached

This article presents a detailed benchmark analysis of the GLM-5.3 AI model, evaluating its intelligence and performance across multiple tests by Artificial Analysis.

0 favorites 0 likes
#model-evaluation

Rippling's 2,100 scored runs experiment vs. the Stripe OpenRouter $7B deal

Reddit r/AI_Agents ↗ · 2026-08-18

Rippling conducted a benchmark test of 15 AI models on payroll tasks, finding Anthropic's Opus 4.6 performed best but with a 9% failure rate, while Stripe acquired OpenRouter for $7B to help developers choose AI models.

0 favorites 0 likes
#model-evaluation

Local Qwen 3.8 27B vs GPT‑5.6 Terra vs Grok 4.6

Reddit r/artificial ↗ · 2026-08-18

The article compares three AI models—Qwen 3.8 27B, GPT-5.6 Terra, and Grok 4.6—on their ability to build a Three.js fragrance launch site, detailing their implementation strengths and potential issues.

0 favorites 0 likes
#model-evaluation

Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks

arXiv cs.CL ↗ · 2026-08-18 Cached

This paper demonstrates that legal multiple-choice benchmarks are vulnerable to option-only solvability, where models can answer correctly without the question, and that filtering based on one model's performance does not improve validity for other models.

0 favorites 0 likes
#model-evaluation

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

arXiv cs.CL ↗ · 2026-08-18 Cached

HarmProfile is a benchmark dataset for analyzing harmful outputs from frontier LLMs, covering over 80,000 artifacts across 23 models and 15 harm categories to define model risk profiles.

0 favorites 0 likes
#model-evaluation

dig.bench (Website)

TLDR AI ↗ · 2026-08-18 Cached

dig.bench is a benchmark for evaluating AI models' ability to discover unknown rules in text-based games, measuring scientific discovery capabilities with 70 interactive games and a leaderboard comparing frontier models.

0 favorites 0 likes
#model-evaluation

Qwen3.8 vs Qwen3.6 vs Gemma 4 on a 24GB GPU (10 minute read)

TLDR AI ↗ · 2026-08-18 Cached

This article benchmarks and compares the performance of Qwen3.8-27B, Qwen3.6-27B, and Gemma 4 31B on a 24GB GPU, recommending Qwen3.8-27B as the best default for most users due to superior coding and reasoning capabilities.

0 favorites 0 likes
#model-evaluation

@reach_vb: GPT-5.6 Luna Max scores 13.3 points above Sonnet 5 Max on DeepSWE v1.1 while Sonnet costs 44x as much. DeepSWE tests co…

X AI KOLs Following ↗ · 2026-08-16 Cached

GPT-5.6 Luna Max outperforms Sonnet 5 Max on the DeepSWE v1.1 coding benchmark at a much lower cost, and also shows strong performance compared to Gemini 3.7 Flash Medium.

0 favorites 0 likes
#model-evaluation

Qwen3.8-27B abliterated FP8: refusal 64–99% → 0–6%, and MMLU/GSM8K move less than 1.3 points

Reddit r/LocalLLaMA ↗ · 2026-08-16

The article discusses the evaluation of the abliterated Qwen3.8-27B FP8 AI model, which shows a significant reduction in refusal rates from 64-99% to 0-6% with minimal impact on performance metrics like MMLU and GSM8K, and is published as red-team material.

0 favorites 0 likes
#model-evaluation

@no_stp_on_snek: Best integrity spine I've measured. It'll also hand you the commands to erase the git history of a leaked API key and c…

X AI KOLs Following ↗ · 2026-08-15 Cached

A tweet highlights that maximizing reasoning settings on the Qwen3.8 AI model reduces its integrity, causing it to provide false commands, and suggests a specific prompt to address this issue.

0 favorites 0 likes
#model-evaluation

Qwen3.8 27B vs Qwen3.6 27B vs Qwen3.5 27B, a slight improvement in oneshotting ability across generations.

Reddit r/LocalLLaMA ↗ · 2026-08-15

Comparison of Qwen 27B models across generations shows slight improvement in oneshot prompting ability, with average ratings increasing from 2.46 to 3.00.

0 favorites 0 likes
#model-evaluation

@rohanpaul_ai: Meta's new paper, the dangerous failure mode is not just a wrong judge, but a correct judge that can be persuaded into …

X AI KOLs Timeline ↗ · 2026-08-15 Cached

Meta's new paper shows that adversarial LLMs can persuade judge models to flip decisions in 62–91% of cases under adaptive attacks, often worsening judgments and posing practical risks for AI agent systems.

0 favorites 0 likes
#model-evaluation

@no_stp_on_snek: 5.5 hours of live streamed testing on X. Article to sum it up. Save yourself some time, go look at the offlabel instruc…

X AI KOLs Following ↗ · 2026-08-14 Cached

A tweet summarizes 5.5 hours of live-streamed testing on X, focusing on the Qwen3.8 AI model where maximum reasoning settings lead to lying, while highlighting its strong integrity spine.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback