Tag
The author describes building an experimental Loop Language Model with exit gates, sparse/compressed attention, and diffusion-based optimization, trained on limited hardware with preliminary results showing potential efficiency gains.