Tag
Harrison Chase announces a post-trained model for detecting issues in production agent traces, claiming SOTA accuracy at 10-100x cheaper rates than frontier models.
Recursive's automated AI research system achieves state-of-the-art results on NanoChat, NanoGPT Speedrun, and GPU kernel benchmarks by automating the research loop without task-specific adaptations, and open-sourcing artifacts for further inspection.
Browser Use Beta achieved state-of-the-art results on a difficult internal web agent benchmark, using Fable for optimization and analysis.
Fable-5/Mythos achieves new SOTA on agentic search but is expensive for self-hosting, while open-weight Harness-1 offers a cost-effective alternative with fewer query restrictions.
Anthropic announces Fable 5, a state-of-the-art model for mechanical engineering tasks capable of generating intricate assemblies and mechanisms from a single prompt.
Argus-Retriever is a new late-interaction visual document retriever that adapts document representation to the query, achieving SOTA performance on ViDoRe benchmarks with a smaller index.
Miles Brundage announces a state-of-the-art (SOTA) score improvement on the Clear AVERI Pronunciation Guide Bench achieved by colleague Carly.
Recommended reading: the MAI-Thinking-1 technical paper, which details almost all the steps to train a SOTA large language model.
NVIDIA releases Cosmos 3 (Mixture-of-Transformers models up to 64B), Nemotron 3 Ultra (550B-A55B LLM), and previews RTX Spark personal superchip at Computex 2026, achieving SOTA on multiple open model leaderboards.
Nick Kang adds a new task to his Twitter benchmark collection; Claude Opus 4.8 and other SOTA models pass, while Sonnet 4.6 and Grok 4.3 fail. Alfin remarks on Opus 4.8's dangerous capabilities.
Extend released Parse 2.0, a state-of-the-art document parsing API that achieves top accuracy on real-world documents, outperforming competitors on the open-source RealDoc-Bench benchmark.
LongCat released an open-source talking-avatar model (likely state-of-the-art) under MIT license, with a Hugging Face demo, enabling various applications like AI tutors, dubbing, and coding agents.
HRM-text is a 1B-parameter hierarchical reasoning language model proposed by Sapient Intelligence. It thinks efficiently through internal latent space, achieving performance surpassing most models of the same size with extremely low training cost.
Ant Group released Ring-2.6-1T, a 1 trillion parameter reasoning model for agent workflows, featuring MIT license, extended context, and Async RL + IcePop training, achieving state-of-the-art results.
NielsRogge announces a revival of PapersWithCode, featuring SOTA per domain, leaderboards, and methods parsed at scale using AI agents.
Poetiq claims new state-of-the-art coding performance using a self-optimizing harness with Gemini 3 Flash, surpassing Opus 4.7.
Poetiq's Meta-System achieved state-of-the-art results on LiveCodeBench Pro by autonomously building a coding harness using standard APIs and Gemini 3.1 Pro, without fine-tuning or special model access.
Two new open-source small language models are being released: one matches state-of-the-art accuracy at up to 93x smaller size, and the other outperforms a recent OpenAI model. The first model drops tomorrow.
A new open-source memory layer called Memvid claims to outperform all existing RAG systems, achieving +35% SOTA on LoCoMo and +76% on multi-hop reasoning, packaged as a single .mv2 file.
Xiaomi launched MiMo-V2.5-Pro, claiming state-of-the-art performance.