Tag
Kimi K3 is a 2.8 trillion parameter Mixture-of-Experts model with 104 billion active parameters, native vision, and a 1-million-token context window, achieving frontier-level performance across multiple domains and released as open weights.
Compares two new AI models for agentic workloads: the compact Nanbeige4.2-3B with a looped transformer architecture and the large Mixture-of-Experts Laguna S2.1, both released on Hugging Face.
A leaked report claims that Z.ai's upcoming GLM-5.5 model, with over 1 trillion parameters and 1M-token context, will compete directly with Fable 5 and Mythos, targeting an August release and remaining open-weight.
Kimi Linear 48B A3B appears to be a new large language model with 48 billion parameters, likely from Moonshot AI.
Anthropic releases Claude Opus 5, a new large language model with enhanced capabilities and safety features, as detailed in its system card.
Anthropic publishes the system card for Claude 5 Opus, detailing its capabilities, safety evaluations, and deployment details.
This paper presents a retrieval-augmented, multi-agent LLM framework with human-in-the-loop for detecting cutaneous immune-related adverse events from clinical notes, achieving higher accuracy, improved inter-rater agreement, and halved review time compared to manual review.
Ant Group released Ling-3.0-flash, a 124B MoE model with 5.1B active parameters per token and 256K context expandable to 1M, matching or outperforming their 1T flagship model on most benchmarks.
Arcee AI announces Genesis-Science-1 (GS1), a 1 trillion parameter open-weight model to be released later this year.
Poolside releases Laguna S 2.1, a 118B MoE model with 8B activated parameters per token, optimized for agentic coding. It claims to outperform DeepSeek V4 Pro while being cheaper than DeepSeek V4 Flash, with a 1M context window and open-source license.
The new Laguna S 2.1 model by poolsideai, with 118B total parameters and 8B active, is now free for two weeks on Nous Portal.
poolside released Laguna-S-2.1, a new 120 billion parameter language model that emerges as a strong contender in the LLM landscape.
Using the fastjson 0day to test LLM vulnerability discovery capabilities, found Claude and Doubao performed outstandingly.
Debate-on-Graph (DoG) is a framework that enhances LLM reasoning by leveraging uncertain knowledge graphs (UKGs) with confidence scores, using a heuristic search and multi-agent debate mechanism to produce reliable answers. It achieves state-of-the-art performance on four QA benchmarks.
Moonshot's Kimi K3, a 2.8 trillion parameter open weights model with 896 experts (16 active per token), exemplifies the trend of scaling total parameters while holding active compute constant, and uses attention compression to reduce KV cache size, making frontier inference more accessible but with high storage costs.
Motif Technologies releases an intermediate beta checkpoint of Motif-3, a large-scale Mixture-of-Experts language model with ~314B total parameters (~13B active), 256K context length, and custom architectures like Grouped Differential Latent Attention (GDLA), openly available for non-commercial research.
A 744B parameter mixture-of-experts model boots on a laptop with 25GB RAM by storing expert weights on SSD and only loading the active ~40B parameters per token, enabling local inference despite the model's size.
The paper introduces a knowledge-centric framework for generating ComfyUI workflows by distilling hierarchical knowledge (pseudo-codes, skeletons, strategies) from real workflows and using LLMs to perform reasoning from task descriptions to executable structures, achieving higher node diversity and execution success rates.
Qwen3.8, a 2.4-trillion-parameter language model, is being released as open-weight by Alibaba, with a preview available on Alibaba's Token Plan, Qoder, and QoderWork.
Tested the new Qwen 3.8 model, a large language model with 2.4 trillion parameters.