attention-compression

Tag

Cards List
#attention-compression

Sparse By Design (5 minute read)

TLDR AI · 2d ago Cached

Moonshot's Kimi K3, a 2.8 trillion parameter open weights model with 896 experts (16 active per token), exemplifies the trend of scaling total parameters while holding active compute constant, and uses attention compression to reduce KV cache size, making frontier inference more accessible but with high storage costs.

0 favorites 0 likes
← Back to home

Submit Feedback