hidden-decoding

Tag

Cards List
#hidden-decoding

@Fenng: This paper from WeChat's WeLM team reveals at least two model scales: 80B and 617B. Each scale includes a standard version and an HD4 version, namely WeLM-80B, WeLM-HD4-80B, WeLM-617B, and WeLM-HD4-617B. Among them, WeLM-617…

X AI KOLs Timeline · 2026-07-12 Cached

The WeChat WeLM team published a paper introducing the Hidden Decoding method, which extends computation through hidden flows without increasing the Transformer backbone parameters, training the WeLM-HD4-80B and WeLM-HD4-617B MoE models, surpassing autoregressive baselines on multiple benchmarks.

0 favorites 0 likes
#hidden-decoding

Hidden Decoding at Scale: Latent Computation Scaling for Large Language Models

arXiv cs.CL · 2026-07-10 Cached

This paper introduces Hidden Decoding, a sequence-length scaling method for LLMs that adds internal computation per token by expanding each token into multiple streams with independent embeddings, using Stream-Factorized Attention to keep costs low. Experiments on models up to 617B parameters show consistent improvements over baselines, demonstrating a practical fixed-backbone scaling path.

0 favorites 0 likes
← Back to home

Submit Feedback