delta-attention

Tag

Cards List
#delta-attention

You Could Have Come Up with Kimi Delta Attention

Hacker News Top · 2026-07-28 Cached

This blog post derives Kimi Delta Attention step by step from standard softmax attention through linear attention and DeltaNet variants, explaining the state update equations used by recent Qwen and Kimi models.

0 favorites 0 likes
#delta-attention

@akshay_pachaar: one matrix replaced the KV cache. (the technique is 100% open source) Kimi just dropped K3, an open model at frontier s…

X AI KOLs Following · 2026-07-24 Cached

Kimi released K3, a 2.8T-parameter open model using delta attention to avoid growing KV cache, enabling a 1-million-token context window with linear memory cost.

0 favorites 0 likes
#delta-attention

Semidirect Fourier Delta Attention: Phase-Controlled Delta Memory with Constructive Chunk-WY Kernels

arXiv cs.LG · 2026-07-15 Cached

Introduces Semidirect Fourier Delta Attention (SFDA), a phase-controlled delta-rule layer that extends Kimi Delta Attention with block-rotational Fourier control operators, providing a constructive chunk-WY theorem for efficient chunkwise computation and demonstrating expressivity for cyclic and register memories.

0 favorites 0 likes
#delta-attention

Delta Attention Residuals

Hugging Face Daily Papers · 2026-05-13 Cached

Delta Attention Residuals improve layer-wise routing in transformer models by attending to feature changes (deltas) rather than cumulative hidden states, achieving 1.7–8.2% validation perplexity gains across scales from 220M to 7.6B parameters.

0 favorites 0 likes
← Back to home

Submit Feedback