Introducing DWARF-55M-Base
Summary
DWARF-55M-Base is a new language model using a nearly all-sparse attention architecture (DSQG) with a single full causal attention layer, achieving reliable retrieval up to 2048 tokens and extrapolating to 3x that context. It is released as a research prototype for community experimentation.
Similar Articles
Subquadratic AI introduces SubQ-1.1-Small, a new model using Smart Sparse Attention
Subquadratic AI introduces SubQ-1.1-Small, a model leveraging Smart Sparse Attention to achieve near-perfect long-context retrieval up to 12M tokens with up to 1,000x attention compute reduction. It balances long-context optimization with strong general reasoning, outperforming baselines on benchmarks like NIAH and RULER.
@omarsar0: Banger paper from Google DeepMind and colleagues. (bookmark it) A model reads its entire KV cache on every generated to…
This paper introduces Declarative Attention, a protocol that allows language models to declare where to attend in their chain-of-thought, reducing attended tokens by 52% on Gemma-4-31B and 31.1% on Qwen-3.6-27B with minimal accuracy drops.
[RELEASE] SupraBrain-50M-v0.1
SupraLabs releases SupraBrain-50M, a hybrid language model combining Gated DeltaNet, Sliding-Window Attention, and Surprise-Gated mechanisms, achieving near-parity with Supra-Base-50M despite training on far fewer tokens.
Keye-VL-2.0-30B-A3B -- Introducing DSA attention into multimodality for the first time
Kwai releases Keye-VL-2.0-30B-A3B, a 30B-class multimodal base model that introduces DSA attention to multimodality for the first time, targeting long-video understanding and agent capabilities.
deepseek-ai/DeepSeek-V4-Flash-DSpark
DeepSeek releases V4 series of Mixture-of-Experts language models (Pro 1.6T/49B activated, Flash 284B/13B activated) supporting one-million-token context with hybrid attention and speculative decoding, claiming best open-source model performance.