@ProfTomYeh: Self-Attention vs Cross-Attention interactive diagram. Open https://byhand.ai/self-vs-cross

X AI KOLs Timeline Tools

Summary

An interactive diagram comparing self-attention and cross-attention mechanisms in AI models, published as part of an educational library on attention.

Self-Attention vs Cross-Attention interactive diagram. Open https://t.co/RgpT6ger01 https://t.co/FETxf0Zzql
Original Article
View Cached Full Text

Cached at: 09/20/26, 07:17 AM

Self-Attention vs Cross-Attention interactive diagram. Open https://t.co/RgpT6ger01 https://t.co/FETxf0Zzql


Self Attention vs Cross Attention

Source: https://www.byhand.ai/p/library-models-attention-self-vs-cross

Library›Attention

  1. QKV Projection
  2. Attention Computation
  3. Self Attention
  4. Cross Attention
  5. Self Attention vs Cross Attention
  6. Self Attention (Shared KV)
  7. Multi-Head Attention
  8. Fused QKV (Multi-Head)
  9. Single vs Multi-Head Attention
  10. Multi-Query Attention
  11. Grouped-Query Attention

This is a review of the two previous articles, shown side by side.

**Top: self-attention.**Midnight. N neighbors are awake, every one of them asking “whose dog is barking?” Each neighbor queries every other neighbor. The score matrix is square: the queriers and the answerers are the same N people.

**Bottom: cross-attention.**Saturday morning. N neighbors are going on trips, asking M teens “who can sit my dog?” Each neighbor queries every teen, but not each other. The score matrix is rectangular: the queriers and the answerers are two different groups.

Both use the same X for queries. The only difference is where K and V come from. In self-attention, K and V are drawn from X itself. In cross-attention, K and V come from a second sequence E. That single change is what shifts the matrix from N × N to N × M.

Notice what stays the same in both: Q and K must share the same key dimension, because the dot product Kᵀ × Q only works when they land in the same space. V is free to have its own dimension in either case.

Next: 6. Self Attention (Shared KV)

Discussion about this post

Ready for more?

Similar Articles

Kuramoto Attention: Synchronizing Self-Attention on the Torus

arXiv cs.LG

Introduces Kuramoto attention, a self-attention layer where hidden states are phase angles on a torus, enabling synchronization through gated cosine similarity and circular mean updates. The layer performs comparably to standard transformers on character-level language modeling.