Tag
The WeChat WeLM team published a paper introducing the Hidden Decoding method, which extends computation through hidden flows without increasing the Transformer backbone parameters, training the WeLM-HD4-80B and WeLM-HD4-617B MoE models, surpassing autoregressive baselines on multiple benchmarks.
Fenng shares a self-media comparison between the fourth-generation WeLM-80B (80B total params, 3B activated, 3.75% activation rate) and DeepSeek-V4-Flash (284B total, 13B activated, 4.6% activation rate), with a humorous comment.
Fenng explained the positioning difference between Tencent's WeLM (closed-source large model) and Hunyuan (open-source large model), pointing out that the closed-source model will not disclose technical details or participate in evaluations, and suggested understanding it through product experience.
The WeLM large model developed by the WeChat team is considered to have entered the top tier of domestically produced large models.
WeChat launches AI Agent product 'Xiao Wei', main model uses WeLM, some answers fallback to DeepSeek, grayscale testing has begun.