attention-modification

Tag

Cards List
#attention-modification

Rebuilding Gemma 4 31b... better... As 26b...

Reddit r/LocalLLaMA · 2026-07-02

A developer is rebuilding Gemma 4 31b into a smaller 26b model by removing weak SWA layers, adding attention-based residual networks (from Moonshot), and using topK logit targets for retraining, aiming for better long context and performance.

0 favorites 0 likes
← Back to home

Submit Feedback