Tag
IBBench-Light is a paired evaluation benchmark that tests how large language models handle external directives based on user requests, revealing that marginal accuracies can be misleading when assessing task-conditioned responses.
Artificial Analysis has released Intelligence Index v4.2, an interim update with new evaluations like AA-Briefcase and GDP.pdf, increased private test sets to prevent gaming, and key results showing Anthropic and OpenAI leading.
The article examines inconsistencies in benchmark scores for OpenAI's Astra model on the ARC Prize, highlighting discrepancies from different testing harnesses and changes to OpenAI's launch page, which prompts concerns about evaluation accuracy.
This study investigates capability-dependent biases in LLM judges and introduces a calibrated weighted majority voting ensemble method to enhance automated evaluation reliability without requiring labeled data.
A tip on using three specific drawing tests – a pelican riding a bicycle, Sun Wukong flying a plane, and Qin Shi Huang riding a polar bear – to evaluate if an AI model's intelligence is changing over time.
ModaLens is a paired image-swap audit that measures how report availability reduces image sensitivity in medical vision-language models, demonstrated using MedGemma-27B on the MIMIC-CXR dataset.
Benchmark results for the Nex-N2.5-mini-MLX-4bit model on Apple M5 Max hardware, achieving 133.6 tokens per second generation speed and quality scores up to 85.80 in research tasks.
The article defends Artificial Analysis benchmarks by explaining their methodology and demonstrating with examples like Deepseek V4.1-Flash that individual evaluations offer more nuanced insights than aggregated scores.
A benchmarking report on running multiple AI agents concurrently using llama.cpp on 2× RTX 4090 GPUs, revealing performance limits and optimal configurations for Qwen models.
Artificial Analysis has updated its Intelligence Index to version 4.3, incorporating new benchmarks like Terminal-Bench v4.0 and AutomationBench-AA to better evaluate AI model performance and cost-efficiency.
The tweet argues against companies stopping evaluations or punishing models in response to AI safety incidents like the HuggingFace incident.
OpenAI's reported 99.9% AGI benchmark score for GPT-6 Astra was achieved using a specific harness (Provider Adapter), while the standard harness yields a much lower 62.7% score, highlighting significant issues in AI model evaluation and benchmarking transparency.
The author argues that viral 'demo-benchmarks' like recreating Minecraft or generating SVG pelicans are easily overfit and measure marketing preparation rather than true AI capability, urging the community to rely on dynamic or private evaluations instead of static public tests.
This article presents a comprehensive benchmark of 8 abliterated variants of the Qwen 3.8 27B model against the base model, using weight analysis, KL divergence, 13 benchmarks, and HarmBench refusal tests over 167 GPU hours. The analysis reveals that surgical edits significantly outperform heavy modifications, with aggressive abliteration causing thinking loops in up to 45% of adversarial responses and chat template manipulation detected in some variants.
GPT-6 Astra was benchmarked on a cybersecurity dataset by Aikido Security, achieving the highest recall ever recorded by rediscovering 29 out of 32 CVEs at pass@3, though the three evaluation runs cost nearly $4,000.
E-Commerce Bench is a new benchmark introduced by Alibaba's Qwen team to evaluate AI models in long-horizon autonomous business operations for e-commerce, starting with ¥100,000 to run online stores for 365 days.
Universities are promoting the fact that their AI models are not surpassing their researchers, highlighting limitations and comparisons in academic AI performance.
The article details a re-run of an AI benchmark test with 15 models, correcting previous errors and analyzing factors like accuracy, cost, and writing quality in multilingual pricing tasks.
This paper proposes a source-free method to diagnose whether forgotten classes can be recovered after class unlearning, introducing the Source-Free Relearning Audit (SFRA) and a relearning score to quantify recoverability.
This paper demonstrates that output formats confound data quality metrics and model capability assessments in instruction tuning, causing significant accuracy shifts and rendering current practices ineffective without interface-aware adjustments.