Tag
Locus, an automated AI research system, achieves SOTA on PostTrainBench and post-trains Qwen3 base models that surpass human post-trained models. It also shows strong performance on Kaggle competitions.
Kimi K3 achieves state-of-the-art results on the new pmpp-hard benchmark, scoring 0.71 across 69 GPU kernel tasks and beating other frontier models.
Qwen released Qwen-CUA, a native computer-use agent with a 397B-A17B mixture-of-experts backbone, achieving state-of-the-art results on OSWorld-Verified and ranking #2 on WebArena. A technical report is available on Papers with Code.
Mike Bradley shares benchmark results claiming DeepSeek V4 Flash 0731 is the current state-of-the-art for 190GB VRAM systems, matching or exceeding an Unsloth 3-bit Qwen3.5-397B in quality while running about 3x faster.
Announcement of an open-source SOTA video generation model, MiniMax H3, offering commercial-grade generation, unbeatable cost efficiency, and open weights.
KAT-Coder-V2.5-Dev is an open-weight MoE coding model with 35B total parameters (3B active), achieving state-of-the-art results on agentic coding benchmarks through SFT and RL training.
LLM-as-a-Verifier is a simple, cheap, general-purpose self-improvement technique for agentic tasks, using fine-grained scoring and logprob-based ranking to achieve SOTA on multiple benchmarks like SWE-Bench Verified and Terminal-Bench V2.
Today, we are releasing Le Chaton L∃∀N, aka Leanstral 1.5, which achieves SOTA performance on graduate algebra benchmarks FATE-H and FATE-X and improves the Pareto Frontier on PutnamBench, solving 587/672 problems with a x10 cheaper budget.
Ornith-1.0 has been released on Hugging Face, featuring a collection of models ranging from 9B to 397B parameters, including dense and MoE architectures, claiming state-of-the-art performance on various benchmarks.
GLM-5.2 achieves state-of-the-art results on PostTrainBench, outperforming GPT-5.5 and Opus 4.8.
SAG (SQL-Augmented Generation) is a novel SQL-based retrieval augmented generation method that converts data chunks into events and entities, enabling multi-hop reasoning via SQL join queries. On the MuSiQue dataset, recall increased from 65.13% to 80.04%. It supports second-level online retrieval of approximately 500 million data entries and has been open-sourced.
Harrison Chase announces a post-trained model for detecting issues in production agent traces, claiming SOTA accuracy at 10-100x cheaper rates than frontier models.
Recursive's automated AI research system achieves state-of-the-art results on NanoChat, NanoGPT Speedrun, and GPU kernel benchmarks by automating the research loop without task-specific adaptations, and open-sourcing artifacts for further inspection.
Browser Use Beta achieved state-of-the-art results on a difficult internal web agent benchmark, using Fable for optimization and analysis.
Fable-5/Mythos achieves new SOTA on agentic search but is expensive for self-hosting, while open-weight Harness-1 offers a cost-effective alternative with fewer query restrictions.
Anthropic announces Fable 5, a state-of-the-art model for mechanical engineering tasks capable of generating intricate assemblies and mechanisms from a single prompt.
Argus-Retriever is a new late-interaction visual document retriever that adapts document representation to the query, achieving SOTA performance on ViDoRe benchmarks with a smaller index.
Miles Brundage announces a state-of-the-art (SOTA) score improvement on the Clear AVERI Pronunciation Guide Bench achieved by colleague Carly.
Recommended reading: the MAI-Thinking-1 technical paper, which details almost all the steps to train a SOTA large language model.
NVIDIA releases Cosmos 3 (Mixture-of-Transformers models up to 64B), Nemotron 3 Ultra (550B-A55B LLM), and previews RTX Spark personal superchip at Computex 2026, achieving SOTA on multiple open model leaderboards.