Tag
DeepSeek v4 Pro 0813 outperforms all other models on a cybersecurity vulnerability-finding benchmark, achieving 87.5% CVE rediscovery at pass@3, though with lower precision and run consistency.
This paper introduces Macaron-V1, an open continual learning agent-model family using Mixture-of-LoRA to compose specialist adapters on frozen base models, with recursive self-improvement and model-harness co-design.
A new tool called kimi-k3-in-c runs the 2.78T-parameter Kimi K3 open model on a single CPU with as little as 8.24 GB RAM, streaming experts from disk and achieving deterministic output at 10-32 seconds per token.
NVIDIA released Alpamayo 2 Super, an open reasoning model for autonomous vehicles, now available for commercial use under the OpenMDW-1.1 license, delivering frontier-scale reasoning for robotaxis and AV development.
Arcee AI, Loka, AWS, and Prime Intellect post-trained an open model using reinforcement learning to improve scientific tool use and biological reasoning, achieving notable gains on drug tool and Gene Ontology benchmarks.
Kimi K3 is a 2.8T parameter open model from Moonshot AI, showing strong benchmark performance but likely over-optimized and lagging behind top closed models by months. It is distilled from Claude and its release may precede an IPO.
NVIDIA released Cosmos 3 Edge, a 4-billion-parameter open world model for edge devices that helps robots and vision AI agents understand surroundings, reason in real time, and generate actions. It achieves best-in-class throughput and accuracy among similar-sized models.
Google announced Gemma 4 E2B optimized for the Pixel 10's TPU, enabling on-device multimodal AI capabilities like offline chat, image recognition, and audio transcription.
An open model that predicts a robot's actions from a control signal, raising questions about whether it constitutes a world model or just a video generator.
This paper analyzes open language model adoption, finding that Chinese models led by Qwen now dominate downloads, surpassing US models by March 2026. Qwen's lead comes from a diverse range of model sizes, while DeepSeek leads in very large models, and some US models still show strong momentum.
Introducing Leanstral 1.5, a 119B parameter (6B active) open model for formal proof engineering in Lean 4, achieving 100% on miniF2F, state-of-the-art scores on PutnamBench and FATE benchmarks, and discovering previously unknown bugs in open-source repositories.
GLM 5.2, an open AI model, is now available in the Cursor coding tool via a partnership with Fireworks.
NVIDIA released Nemotron-TwoTower-30B-A3B-Base-BF16, a diffusion-based language model that uses block-wise autoregressive diffusion to generate text by iterative denoising of token blocks, achieving 2.42× the generation throughput of the autoregressive baseline while retaining 98.7% of benchmark quality.
GLM-5.2 is a new open-source AI model that sets a high bar for open models, though it still trails proprietary frontier models and lacks some features like vision.
NVIDIA released the Nemotron 3 open model, offering three sizes: Nano, Super, and Ultra. It optimizes hardware efficiency through architectural innovations such as hybrid Mamba Transformer, latent MoE, and multi-token prediction, and adopts the Open MDW 1.1 open license.
NVIDIA optimizes Google DeepMind's DiffusionGemma, an open model that generates text in parallel 256-token blocks, achieving up to 4x faster performance on local RTX GPUs, DGX Spark, and DGX Station systems.
Google introduces DiffusionGemma, an experimental 26B MoE open model that achieves up to 4x faster text generation on GPUs using text diffusion, targeting speed-critical interactive local workflows.
Nex AGI releases Nex-N2, an open-source agentic model series for coding, tool use, deep research, and long-horizon workflows, with state-of-the-art benchmarks and Apache 2.0 license.
NVIDIA announces Alpamayo 2 Super, a 32B open reasoning model for Level 4 robotaxis, featuring 360-degree perception, meta-actions, and a full stack including AlpaGym simulation and OmniDreams scenario generation.
The article observes that Tencent's Hy3 Preview open model performs surprisingly well in evaluations, narrowing the gap with top closed models, yet remains underdiscussed compared to Western AI labs.