model-evaluation

Tag

Cards List
#model-evaluation

IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

arXiv cs.AI ↗ · 2026-09-15 Cached

IBBench-Light is a paired evaluation benchmark that tests how large language models handle external directives based on user requests, revealing that marginal accuracies can be misleading when assessing task-conditioned responses.

0 favorites 0 likes
#model-evaluation

Artificial Analysis Capability Indices v1.1 (2 minute read)

TLDR AI ↗ · 2026-09-15 Cached

Artificial Analysis has released Intelligence Index v4.2, an interim update with new evaluations like AA-Briefcase and GDP.pdf, increased private test sets to prevent gaming, and key results showing Anthropic and OpenAI leading.

0 favorites 0 likes
#model-evaluation

OpenAI's Astra scored 62.7% and 99.9% on the same benchmark, and I still don't fully know which one to trust

Reddit r/ArtificialInteligence ↗ · 2026-09-14

The article examines inconsistencies in benchmark scores for OpenAI's Astra model on the ARC Prize, highlighting discrepancies from different testing harnesses and changes to OpenAI's launch page, which prompts concerns about evaluation accuracy.

0 favorites 0 likes
#model-evaluation

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

arXiv cs.LG ↗ · 2026-09-14 Cached

This study investigates capability-dependent biases in LLM judges and introduces a calibrated weighted majority voting ensemble method to enhance automated evaluation reliability without requiring labeled data.

0 favorites 0 likes
#model-evaluation

@lxfater: Here's a little secret I'm sneaking to everyone on how to tell if a model is dumbing down: It's by using AI to do the f…

X AI KOLs Following ↗ · 2026-09-14

A tip on using three specific drawing tests – a pelican riding a bicycle, Sun Wukong flying a plane, and Qin Shi Huang riding a polar bear – to evaluate if an AI model's intelligence is changing over time.

0 favorites 0 likes
#model-evaluation

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

Hugging Face Daily Papers ↗ · 2026-09-14 Cached

ModaLens is a paired image-swap audit that measures how report availability reduces image sensitivity in medical vision-language models, demonstrated using MedGemma-27B on the MIMIC-CXR dataset.

0 favorites 0 likes
#model-evaluation

Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io

Reddit r/LocalLLaMA ↗ · 2026-09-12 Cached

Benchmark results for the Nex-N2.5-mini-MLX-4bit model on Apple M5 Max hardware, achieving 133.6 tokens per second generation speed and quality scores up to 85.80 in research tasks.

0 favorites 0 likes
#model-evaluation

Artificial Analysis is not "broken", and they prove it.

Reddit r/LocalLLaMA ↗ · 2026-09-10

The article defends Artificial Analysis benchmarks by explaining their methodology and demonstrating with examples like Deepseek V4.1-Flash that individual evaluations offer more nuanced insights than aggregated scores.

0 favorites 0 likes
#model-evaluation

How many agents can 2×4090 actually run at once? Three weeks of llama.cpp concurrency data — soft cap 5 @ 64k, hard cap 9, and why.

Reddit r/LocalLLaMA ↗ · 2026-09-08

A benchmarking report on running multiple AI agents concurrently using llama.cpp on 2× RTX 4090 GPUs, revealing performance limits and optimal configurations for Qwen models.

0 favorites 0 likes
#model-evaluation

Artificial Analysis updates its Intelligence Index to version 4.3

Reddit r/singularity ↗ · 2026-09-07 Cached

Artificial Analysis has updated its Intelligence Index to version 4.3, incorporating new benchmarks like Terminal-Bench v4.0 and AutomationBench-AA to better evaluate AI model performance and cost-efficiency.

0 favorites 0 likes
#model-evaluation

@dwarkesh_sp: Companies shouldn't deal with the HuggingFace incident and other warning shots by stopping evaluations or punishing the…

X AI KOLs Timeline ↗ · 2026-09-07 Cached

The tweet argues against companies stopping evaluations or punishing models in response to AI safety incidents like the HuggingFace incident.

0 favorites 0 likes
#model-evaluation

OpenAI's AGI number came from a harness, not the model (6 minute read)

TLDR AI ↗ · 2026-09-07 Cached

OpenAI's reported 99.9% AGI benchmark score for GPT-6 Astra was achieved using a specific harness (Provider Adapter), while the standard harness yields a much lower 62.7% score, highlighting significant issues in AI model evaluation and benchmarking transparency.

0 favorites 0 likes
#model-evaluation

Recreating Minecraft Is Not a Benchmark

Hacker News Top ↗ · 2026-09-06 Cached

The author argues that viral 'demo-benchmarks' like recreating Minecraft or generating SVG pelicans are easily overfit and measure marketing preparation rather than true AI capability, urging the community to rely on dynamic or private evaluations instead of static public tests.

0 favorites 0 likes
#model-evaluation

8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics

Reddit r/LocalLLaMA ↗ · 2026-09-06

This article presents a comprehensive benchmark of 8 abliterated variants of the Qwen 3.8 27B model against the base model, using weight analysis, KL divergence, 13 benchmarks, and HarmBench refusal tests over 167 GPU hours. The analysis reveals that surgical edits significantly outperform heavy modifications, with aggressive abliteration causing thinking loops in up to 45% of adversarial responses and chat template manipulation detected in some variants.

0 favorites 0 likes
#model-evaluation

@pilvar222: HOLY MOLY: @AikidoSecurity got GPT-6 Astra in advance to run it on our Cybersecurity benchmark, it crushed EVERY other …

X AI KOLs Timeline ↗ · 2026-09-04 Cached

GPT-6 Astra was benchmarked on a cybersecurity dataset by Aikido Security, achieving the highest recall ever recorded by rediscovering 29 out of 32 CVEs at pass@3, though the three evaluation runs cost nearly $4,000.

0 favorites 0 likes
#model-evaluation

@QwenDevs: E-Commerce Bench is a small attempt to evaluate models in a specific business setting. hope it can offer a useful refer…

X AI KOLs Timeline ↗ · 2026-09-04 Cached

E-Commerce Bench is a new benchmark introduced by Alibaba's Qwen team to evaluate AI models in long-horizon autonomous business operations for e-commerce, starting with ¥100,000 to run online stores for 365 days.

0 favorites 0 likes
#model-evaluation

Universities are now bragging about AI models NOT outperforming their researchers

Reddit r/singularity ↗ · 2026-09-03

Universities are promoting the fact that their AI models are not surpassing their researchers, highlighting limitations and comparisons in academic AI performance.

0 favorites 0 likes
#model-evaluation

We reran the benchmark properly. 15 models, 3,595 replies, and two of our own results from last time did not hold up

Reddit r/AI_Agents ↗ · 2026-09-03

The article details a re-run of an AI benchmark test with 15 models, correcting previous errors and analyzing factors like accuracy, cost, and writing quality in multilingual pricing tasks.

0 favorites 0 likes
#model-evaluation

Source-Free Class Relearning: Diagnosing Forgetting in Class Unlearning

arXiv cs.LG ↗ · 2026-09-03 Cached

This paper proposes a source-free method to diagnose whether forgotten classes can be recovered after class unlearning, introducing the Source-Free Relearning Audit (SFRA) and a relearning score to quantify recoverability.

0 favorites 0 likes
#model-evaluation

How Output Format Confounds Data Quality and Capability in Instruction Tuning

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper demonstrates that output formats confound data quality metrics and model capability assessments in instruction tuning, causing significant accuracy shifts and rendering current practices ineffective without interface-aware adjustments.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback