mla

Tag

Cards List
#mla

@AI_Whisper_X: Reposting Su Jianlin's review of the K3 architecture. In one sentence, K3 = KDA + MLA + Stable LatentMoE + AttnRes. The whole design isn't about showing off; the core is making trade-offs among model performance, computational efficiency, and training stability. Here's a brief explanation: KD…

X AI KOLs Timeline · 2026-08-04 Cached

Su Jianlin reviews the K3 architecture, focusing on the combination of KDA + MLA + Stable LatentMoE + AttnRes. He explains the design trade-offs, MoE stability improvements, why MLA was kept, and the relationship between DSV4 and MLA.

0 favorites 0 likes
#mla

Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding

arXiv cs.LG · 2026-07-31 Cached

This paper proposes functional reconstruction for converting MHA/GQA checkpoints into MLA draft models for speculative decoding, directly optimizing attention modules to preserve token acceptance. It reports consistent improvements across 192 configurations involving Llama/Qwen models and multiple conversion methods.

0 favorites 0 likes
#mla

@classiclarryd: Question for the LLM Research Community: Is anyone aware of fully reproducible experimental results showing that MLA be…

X AI KOLs Following · 2026-07-20 Cached

A researcher questions the reproducibility of MLA outperforming GQA under same KV cache, sharing early small-scale ablation results and plans for scaling experiments to decide on architecture for next large-scale run.

0 favorites 0 likes
#mla

@IndieDevHailey: 22-Year-Old Reverses Anthropic's Black Box in Two Days! Claude Mythos Architecture Fully Open-Sourced Anthropic Unveils Most Dangerous AI—Claude Mythos, Capable of Autonomously Uncovering 20-Year-Old Zero-Day Vulnerabilities in OS and Browsers, Only Available in Project Glasswi…

X AI KOLs Timeline · 2026-06-28 Cached

22-year-old developer Kye Gomez reversed Anthropic's Claude Mythos black box architecture in just two days and open-sourced the OpenMythos project, using Recurrent-Depth Transformer and other techniques, achieving performance equivalent to a 1.3B model with 770M parameters.

0 favorites 0 likes
#mla

@rasbt: Always back to the basics: LatentMoE was probably inspired by MLA, which was inspired by LoRA, which was inspired by SV…

X AI KOLs Timeline · 2026-06-09 Cached

Sebastian Raschka points out the chain of inspiration from LatentMoE back to eigendecomposition through MLA, LoRA, and SVD.

0 favorites 0 likes
← Back to home

Submit Feedback