Tag
Splish is an unofficial fork of Splash that optimizes Metal kernels for Apple M5 Max chips, delivering up to 1.5× faster AI inference speeds for models like Qwen3.8-27B without compromising quality.
This article demonstrates how to convert GLM-5.3-Flash into a Jev-style System 1 decision model, achieving typed decisions through a single forward pass. Benchmark tests show that it performs comparably to specialized models in terms of accuracy and speed.
Tauon is a new optimizer using polynomial and orthogonalization techniques that outperforms Muon and AdamW in initial benchmarks on a small GPT-Mini model, showing lower loss and faster step time.
BenchBench is a benchmark that challenges multimodal LLMs to compete in assembling Ikea furniture, testing their practical AI capabilities.
Sonnet 5.5, which is claimed to outperform GPT-6 Sol, has received a last-minute upgrade with a release expected on Monday.
VSArena is an open benchmark for evaluating AI agents in interactive 3D environments, offering remote execution and a public leaderboard to assess perception, reasoning, and action without physical robots.
This article demonstrates how to adapt the GLM-5.3-Flash LLM to function as a typed decision model similar to Jev, achieving comparable accuracy and speed while enabling decisions on images in a single forward pass.
jevos is an open-source tool that processes text with yes/no questions in a single forward pass on a laptop CPU, with performance benchmarks compared to Jev and Lay.
Pixel Canary, a stealth AI model, has been released for free in Cline. It matches the performance of GPT-6 Astra and outperforms Kimi K3 on the Next.js Agent Evals benchmark for web and mobile development tasks.
A team announced they achieved first place on NVIDIA's SOL-ExecBench benchmark across all tracks, beating 583 teams with real production kernels on B200 GPUs.
Mica v0.1 4B outperforms Kev 4B in Tetris games by clearing about 4x more lines and surviving to the end in some cases, using only board descriptions without search or lookahead.
An LLM agent named GPT 6 Astra has achieved ascension in the game NetHack, marking a first recorded instance of an LLM beating the complex game through autonomous tool-building.
Xiaomi's MiMo-V2.6-Pro, an AI model with MIT-licensed weights, is listed on OpenRouter at $0.87 per million output tokens, highlighting its cost-effectiveness compared to expensive frontier models.
The article compares TypeSafe's Jev model with the open-source Kev alternative, testing their accuracy, token usage, and speed on a new dataset to avoid data leakage, finding similar performance but differences in implementation.
BRIDGE ASR 2.0 is a benchmark that evaluates speech recognition models on real two-person conversations in 21 languages, focusing on code-switching and multilingual performance to enhance voice interaction in AI and robotics.
A new checkpoint for Gemini 4 pro is announced, showing detailed outputs in 6 minutes and outperforming Gemini 3.8 flash on ArenaAI.
This paper presents a task-substitution framework for automating enterprise governance reviews using AI agents and software, with benchmarks like DGF-Bench showing that models such as Gemini 3.8 Flash can achieve high success rates in replacing human execution for specified review tasks.
Introduces IndicBankBench, a 799-case benchmark for evaluating safety and reliability of language model assistants in Indian retail banking, with multi-stage evaluation and public release of code and data.
RECLAIM is a benchmark for AI agents to reproduce claims from machine learning papers, with difficulty tiers based on available code, data, and weights. It shows that agents struggle to reproduce results, with best success rates of 41% in the easiest tier.
TWIST is a proposed benchmark suite for evaluating intervention quality in conversational memory systems, featuring human-validated tracks to detect failures like unresolved tensions and stale facts, revealing trade-offs between recall and specificity that traditional metrics miss.