Tag
Teacher He Jiyan and 7 PhD students from Zhongguancun College trained a 7B model from scratch in 3 months, achieving state-of-the-art performance in the 7B model category, with insights on efficient training and data quality.
TabPFN-3.5 is released, claiming state-of-the-art performance for tabular data beyond IID small data settings, with features like handling grouped and temporal data, uncertainty calibration, and faster inference.
A new approach called 'Thinking with Looped Flows' trains recurrent reasoning with local denoising objectives, achieving state-of-the-art performance on ARC-AGI benchmarks among looped models.
GPT-6 Astra is announced as the new state-of-the-art in agentic CAD, marking a significant advancement in AI capabilities for mechanical design.
DeepSWE v1.1 shows a 50% improvement on in-house Code Bench and sets new state-of-the-art scores on Terminal-Bench 3.0 and Agents' Last Exam, with emerging cyber capabilities advancing faster than expected.
Runway shares a video highlighting the capabilities of current state-of-the-art image generation models and tooling.
The article discusses Grok 4.6's performance, which is close to Opus 5 in benchmarks at lower cost, and speculates that with more reinforcement learning on browser tasks, Grok 4.7 could become state-of-the-art.
Gemini 3.7 Flash is highlighted as an upgrade over Gemini 3.6, scoring 67% on a browser use dataset compared to 83% SOTA, while Luna offers a cost-effective alternative with a 5x price reduction.
Elon Musk announces that Grok 4.6 reached #1 on Databricks, with Ivan Zhou reporting SOTA performance on OfficeQA Pro V2 using Databricks's Genie harness.
Locus, an automated AI research system, achieves SOTA on PostTrainBench and post-trains Qwen3 base models that surpass human post-trained models. It also shows strong performance on Kaggle competitions.
Kimi K3 achieves state-of-the-art results on the new pmpp-hard benchmark, scoring 0.71 across 69 GPU kernel tasks and beating other frontier models.
Qwen released Qwen-CUA, a native computer-use agent with a 397B-A17B mixture-of-experts backbone, achieving state-of-the-art results on OSWorld-Verified and ranking #2 on WebArena. A technical report is available on Papers with Code.
Mike Bradley shares benchmark results claiming DeepSeek V4 Flash 0731 is the current state-of-the-art for 190GB VRAM systems, matching or exceeding an Unsloth 3-bit Qwen3.5-397B in quality while running about 3x faster.
Announcement of an open-source SOTA video generation model, MiniMax H3, offering commercial-grade generation, unbeatable cost efficiency, and open weights.
KAT-Coder-V2.5-Dev is an open-weight MoE coding model with 35B total parameters (3B active), achieving state-of-the-art results on agentic coding benchmarks through SFT and RL training.
LLM-as-a-Verifier is a simple, cheap, general-purpose self-improvement technique for agentic tasks, using fine-grained scoring and logprob-based ranking to achieve SOTA on multiple benchmarks like SWE-Bench Verified and Terminal-Bench V2.
Today, we are releasing Le Chaton L∃∀N, aka Leanstral 1.5, which achieves SOTA performance on graduate algebra benchmarks FATE-H and FATE-X and improves the Pareto Frontier on PutnamBench, solving 587/672 problems with a x10 cheaper budget.
Ornith-1.0 has been released on Hugging Face, featuring a collection of models ranging from 9B to 397B parameters, including dense and MoE architectures, claiming state-of-the-art performance on various benchmarks.
GLM-5.2 achieves state-of-the-art results on PostTrainBench, outperforming GPT-5.5 and Opus 4.8.
SAG (SQL-Augmented Generation) is a novel SQL-based retrieval augmented generation method that converts data chunks into events and entities, enabling multi-hop reasoning via SQL join queries. On the MuSiQue dataset, recall increased from 65.13% to 80.04%. It supports second-level online retrieval of approximately 500 million data entries and has been open-sourced.