Tag
Inkling, a 975B Mixture-of-Experts model (41B active) with 1M context and Apache-2.0 license, introduces a novel audio front end using a 7.9M parameter lookup table instead of a traditional encoder, achieving strong performance on speech tasks. The model was pretrained on 45 trillion tokens of text, images, audio, and video.
Kimi K3, the latest open-weight model from Moonshot, launches with a new architecture, agent swarm capabilities, and focus on long-horizon agent workflows, positioning it as a major contender against other top models.
NVIDIA AI releases a 75B MoE model (9.3B active) compressed from Nemotron-3-Super-120B using the Iterative Puzzle framework, with 1M token context support.
Google quietly made its newest flash model, Gemini 3.5 Flash, free-tier eligible with no credit card, a 1M context window, and 1,500 requests per day, including native multimodal support.
Mimo 2.5 demonstrates fast performance with large context windows using dual RTX Pro 6000 GPUs.
Zhipu AI (zai_org) has open-sourced GLM 5.2, a flagship model for coding and long-horizon agentic tasks with a usable 1M-token context. Modular is a Day Zero launch partner, offering optimized serving on Modular Cloud.
The user reports that the Gemma 4 12B unified audio model stops attending to speech when the system prompt is large (~21k tokens), and asks for workarounds or explanations, noting the issue persists across vLLM, llama.cpp, and LiteRT-LM backends.
NVIDIA released Nemotron Ultra, a hybrid MoE model with 55B/550B parameters and a 1M context window, supporting MTP speculative decoding and available day-0 in transformers.
MiniMax released M3, an open-weights model combining frontier coding, 1M context, and native multimodality, offering comparable performance to Opus at a fraction of the cost.
StepFun releases Step-3.7-Flash, a new large vision-language MoE model with 198B parameters (11B active), 256K context, and up to 400 tokens/sec inference speed.
OpenAI released GPT Realtime-2 and two accompanying models during Build Hour, enhancing the intelligence and naturalness of voice interaction. It supports 128k context, parallel tool calls, and dynamic voice cloning, demonstrating production-grade applications such as voice-driven shopping assistants and analytics dashboards.