Tag
Alibaba released Qwen3.8-Max, a 2.4 trillion-parameter sparse MoE model with 95B active parameters per token, 1M token context, and strong agentic and benchmark results, including autonomously coding for days, circuit design, and outperforming rivals on Terminal Bench and PaperBench.
ByteDance is in the early stages of training a large language model with up to 10 trillion parameters, signaling a massive scale-up in AI development.
Google's Gemini AI model represents a significant advancement in multimodal AI capabilities.
LG AI Research presents K-EXAONE 2.0, a 750B-parameter MoE foundation model upcycled from K-EXAONE, supporting 256K context and six languages, with self-speculative decoding for efficient inference.
This paper introduces an activation-guided neuron intervention framework using Qwen3-8B to induce Alzheimer's-related computational language phenotypes, demonstrating that amplifying AD-associated neurons produces graded impairments in multiple cognitive domains.
Qwen announces Qwen3.8-Max, a 2.4T parameter model, will be released with open weights next week, along with the Qwen3.8-27B, marking a major push for open-source AI and potentially competitive with top proprietary models.
The paper proposes SentiLLM, a framework that uses semantic-aligned structural abstraction to distill non-verbal modalities into text-like tokens for multimodal sentiment analysis with LLMs. It introduces a dual-stream salience-context calibration mechanism and achieves superior performance on four datasets.
Kimi K3 is a massive 2.9 trillion parameter mixture-of-experts model with 104B active parameters, 1M context length, and native MXFP4 training, now available in GGUF quantizations ranging from 540GB to smaller sizes, though requiring substantial hardware to run.
SKT released A.X K2, a 688B-parameter sparse MoE language model with 33B active parameters, natively trained in FP8 and featuring Think/Non-Think reasoning modes, on Hugging Face.
埃隆·马斯克宣布Grok 4.6将于8月7日左右发布,这是一个1.5T参数的模型,经过改进的SFT和RL;随后几周将发布更好的Grok 4.7(2.1T参数)。
Elon Musk announces that Grok 4.6 will be released around August 7, featuring a 1.5T parameter model with improved SFT and RL, followed by Grok 4.7 with 2.1T parameters.
Moonshot AI releases Kimi K3, a 2.8 trillion parameter open-weight MoE model with native multimodal capabilities and a 1 million token context window, claiming it as the world's first open 3T-class model.
Kimi Moonshot released Kimi K3, a 2.8-trillion-parameter multimodal model with a 1M context window and architectural innovations like Kimi Delta Attention and Attention Residuals, claiming significant efficiency gains and outperforming Claude Opus 4.8 and GPT-5.5 on internal benchmarks.
Moonshot released the weights for their 2.8 trillion parameter Kimi K3 model under a modified license requiring a separate agreement for large commercial Model-as-a-Service businesses.
Kimi K3, a 2.8-trillion-parameter open-source model from Moonshot AI with 1M token context and native vision, is now available on the Telnyx Inference API for tasks like coding, reasoning, and multimodality.
MoonshotAI releases Kimi-K3, a 2.8T-parameter open-weight multimodal agentic model with a 1M-token context window, built on new Kimi Delta Attention and Attention Residuals architecture, achieving significant scaling improvements.
Kimi.ai released the model weights and technical report for Kimi K3, a 2.8T parameter MoE model with native visual understanding and a 1M-token context window, claiming 2.5x intelligence per unit of compute.
Macaron V1 is released in two variants: Venti (748B flagship with 4×1B LoRA specialists) and Tall (50B for local deployment with 4×3.7B LoRA specialists), post-trained on GLM-5.2 and Qwen 3.6 respectively.
Moonshot released the largest open-weight model, Kimi K3 (2.8T parameters), boasting strong benchmarks but requiring massive hardware for self-hosting, making API access the practical choice for most users.
Krasis, a MoE-focused runtime, enables running the 397B-parameter Ornith model on a single RTX PRO 6000 Blackwell 96GB GPU with ~20-24 tok/s decode by dynamically managing expert residency in VRAM.