Keye-VL-2.0-30B-A3B -- Introducing DSA attention into multimodality for the first time
Summary
Kwai releases Keye-VL-2.0-30B-A3B, a 30B-class multimodal base model that introduces DSA attention to multimodality for the first time, targeting long-video understanding and agent capabilities.
Similar Articles
Kwai-Keye/Keye-VL-2.0-30B-A3B
Kwai-Keye releases Keye-VL-2.0-30B-A3B, a 30B-class vision-language model with advanced video understanding, sparse attention, and agent capabilities, achieving top benchmarks.
@AdinaYakup: Keye VL 2.0-30B-A3B New multimodal model from @KwaiKeye 30B/3B active - Apache 2.0 256K context via DeepSeek Sparse Att…
KwaiKeye releases Keye VL 2.0-30B-A3B, a multimodal model with 30B total / 3B active parameters, 256K context via DeepSeek Sparse Attention, and Apache 2.0 license, claiming it matches Qwen3 VL and Gemini 3 in accuracy.
Kwai Keye-VL-2.0 Technical Report
This technical report presents Kwai Keye-VL-2.0, an open-source Mixture-of-Experts multimodal foundation model designed for long-video understanding and agentic intelligence, leveraging DeepSeek Sparse Attention and cross-modal distillation to achieve state-of-the-art performance among similar-scale models.
Introducing DWARF-55M-Base
DWARF-55M-Base is a new language model using a nearly all-sparse attention architecture (DSQG) with a single full causal attention layer, achieving reliable retrieval up to 2048 tokens and extrapolating to 3x that context. It is released as a research prototype for community experimentation.
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Video-DeepResearch (Video-DR) extends multimodal agents from static images to continuous video streams, introducing a decoupled perception-exploration pipeline and a new benchmark Video-DR-Bench. Their Video-DeepResearch-35B-A3B model achieves 64.0% accuracy, surpassing Claude-4.5-Sonnet, GPT-5, and Gemini 2.5 Pro.