@ProfTomYeh: Self-Attention vs Cross-Attention interactive diagram. Open https://byhand.ai/self-vs-cross
Summary
An interactive diagram comparing self-attention and cross-attention mechanisms in AI models, published as part of an educational library on attention.
View Cached Full Text
Cached at: 09/20/26, 07:17 AM
Self-Attention vs Cross-Attention interactive diagram. Open https://t.co/RgpT6ger01 https://t.co/FETxf0Zzql
Self Attention vs Cross Attention
Source: https://www.byhand.ai/p/library-models-attention-self-vs-cross
Library›Attention
- QKV Projection
- Attention Computation
- Self Attention
- Cross Attention
- Self Attention vs Cross Attention
- Self Attention (Shared KV)
- Multi-Head Attention
- Fused QKV (Multi-Head)
- Single vs Multi-Head Attention
- Multi-Query Attention
- Grouped-Query Attention
This is a review of the two previous articles, shown side by side.
**Top: self-attention.**Midnight. N neighbors are awake, every one of them asking “whose dog is barking?” Each neighbor queries every other neighbor. The score matrix is square: the queriers and the answerers are the same N people.
**Bottom: cross-attention.**Saturday morning. N neighbors are going on trips, asking M teens “who can sit my dog?” Each neighbor queries every teen, but not each other. The score matrix is rectangular: the queriers and the answerers are two different groups.
Both use the same X for queries. The only difference is where K and V come from. In self-attention, K and V are drawn from X itself. In cross-attention, K and V come from a second sequence E. That single change is what shifts the matrix from N × N to N × M.
Notice what stays the same in both: Q and K must share the same key dimension, because the dot product Kᵀ × Q only works when they land in the same space. V is free to have its own dimension in either case.
Next: 6. Self Attention (Shared KV)
Discussion about this post
Ready for more?
Similar Articles
@ProfTomYeh: Autoencoder by hand interactive diagram. Open https://byhand.ai/autoencoder ~ Prof. Tom Yeh
Prof. Tom Yeh shares an interactive diagram for learning about autoencoders, part of his 'AI by Hand' series focused on multi-layer perceptrons.
@thtrkim: Visual deep dive on FlashAttention by hand (drawn with Excalidraw) https://winterrykim.github.io/blog/2026/training-lm-…
A visual deep dive into FlashAttention, explaining memory optimization and operator fusion for efficient attention computation in language model training.
Kuramoto Attention: Synchronizing Self-Attention on the Torus
Introduces Kuramoto attention, a self-attention layer where hidden states are phase angles on a torus, enabling synchronization through gated cosine similarity and circular mean updates. The layer performs comparably to standard transformers on character-level language modeling.
@_rohit_tiwari_: https://x.com/_rohit_tiwari_/status/2063982924714901858
This article provides a visual guide to the Transformer architecture in Large Language Models, covering self-attention, causal self-attention, masked multi-head attention, and the output layer with step-by-step explanations and examples.
IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training
IMPACT is a scalable framework for training interaction-aware world models by using cross-attention as an internal interaction map to reweight denoising supervision, improving performance without external representations or inference-time changes.