@AdinaYakup: Keye VL 2.0-30B-A3B New multimodal model from @KwaiKeye 30B/3B active - Apache 2.0 256K context via DeepSeek Sparse Att…

X AI KOLs Following Models

Summary

KwaiKeye releases Keye VL 2.0-30B-A3B, a multimodal model with 30B total / 3B active parameters, 256K context via DeepSeek Sparse Attention, and Apache 2.0 license, claiming it matches Qwen3 VL and Gemini 3 in accuracy.

Keye VL 2.0-30B-A3B 🔥 New multimodal model from @KwaiKeye ✨ 30B/3B active - Apache 2.0 ✨ 256K context via DeepSeek Sparse Attention (probably the first model to ship this in production?👀) ✨ Gets MORE accurate as you feed it more frames ✨ Matchs Qwen3 VL and Gemini 3 https://t.co/B2MO3zMIad
Original Article
View Cached Full Text

Cached at: 06/01/26, 01:23 PM

Keye VL 2.0-30B-A3B 🔥 New multimodal model from @KwaiKeye

✨ 30B/3B active - Apache 2.0 ✨ 256K context via DeepSeek Sparse Attention (probably the first model to ship this in production?👀) ✨ Gets MORE accurate as you feed it more frames ✨ Matchs Qwen3 VL and Gemini 3 https://t.co/B2MO3zMIad

Similar Articles

Kwai-Keye/Keye-VL-2.0-30B-A3B

Hugging Face Models Trending

Kwai-Keye releases Keye-VL-2.0-30B-A3B, a 30B-class vision-language model with advanced video understanding, sparse attention, and agent capabilities, achieving top benchmarks.

Kwai Keye-VL-2.0 Technical Report

Hugging Face Daily Papers

This technical report presents Kwai Keye-VL-2.0, an open-source Mixture-of-Experts multimodal foundation model designed for long-video understanding and agentic intelligence, leveraging DeepSeek Sparse Attention and cross-modal distillation to achieve state-of-the-art performance among similar-scale models.

Qwen3.8-27B

Hacker News Top

Qwen released open weights for Qwen3.8-27B, a native multimodal dense model with 27B parameters that outperforms Qwen3.7-Plus, supports 262K native context extendable to 1M, and is licensed under Apache 2.0.