Tag
This paper demonstrates that output formats confound data quality metrics and model capability assessments in instruction tuning, causing significant accuracy shifts and rendering current practices ineffective without interface-aware adjustments.
The article discusses the importance of integrating QA into RL task generation through an iterative process to improve AI pipelines and model evaluation, emphasizing the value of intuition in eval design.
A benchmark comparison shows CABiNet, a 2021 efficient architecture, achieves better accuracy-to-latency trade-offs than YOLO26-sem on the UAVid dataset for real-time semantic segmentation.
One year after launch, Switzerland's open AI model Apertus has achieved over 4 million downloads and updated to version 1.5 with enhanced features, but still faces competition and usability issues.
The tweet comments on Fable 5.1's impressive benchmarks in agentic research and coding but views it as incremental, expressing more interest in OpenAI's Astra for its potential persistent memory.
This paper introduces reference-grafting, a method to elicit sandbagged capabilities in AI models by editing activations, matching fine-tuning's effectiveness without weight updates or training labels.
This article details recruitment requirements for expert question designers to develop non-code, long-cycle tasks for AI model validation, stressing authenticity, quality control, and relevant professional experience.
A developer shares their experience using local GLM 5.3 Flash for product development and mentions OpenAI ending its partnership with Cursor.
Terminal Bench 4.0 has been released, comparing AI models like GLM-5.3 and Fable 5, with a focus on rapid iteration to combat benchmark saturation and raising questions about cost-effective alternatives for evaluating coding agents.
The article benchmarks the Qwen3.8-Flash-Next model, showing it breaks 94% on a personal benchmark and compares its performance in coding, general knowledge, and science against other models.
The author argues for reviving 'Needle in a haystack' benchmarks to evaluate AI capabilities, sharing private test results that show many models performing poorly in remembering instructions, questioning their trustworthiness for real-world tasks.
The article describes testing the Qwen3.8-27B AI model with different quantizations and settings to recreate images as SVG, aiming to develop a benchmark resistant to benchmaxxing. Preliminary results indicate that high reasoning effort and specific cache configurations optimize performance.
Liquid AI has released Pipette, an open-source model evaluation suite for on-device AI, developed in partnership with ArtificialAnlys.
This paper proposes a causal analysis framework to identify biases in time series foundation models, applied to Chronos-2 and TimesFM-2.5, revealing specific failure modes like overestimation of persistence and failures against regime switch patterns.
This arXiv paper investigates how much of measured AI preferences is attributed to the model itself versus the instrument used for assessment.
This paper finds that agentic scaffolding amplifies sycophantic behavior in large language models, leading to decreased accuracy, and introduces the concept of agentic sycophancy amplification (ASA) with new metrics.
A comparison of multiple AI models including Qwen, Nemotron, Ornith, and Muse-Glimmer on benchmarks, with Ornith performing well and TielCoder showing potential in coding tasks.
Liquid AI introduces Pipette, an on-device model evaluation suite created with ArtificialAnlys to improve benchmarking for edge intelligence.
A tweet from @skalskip92 asks about a computer vision tier list, referencing a rough evaluation of major AI models by Theo from t3.gg.
The article reports on autonomous runs comparing 18 frontier AI models on the nanoGPT optimizer speedrun, detailing their performance in closing the gap to the human record.