@huangyihe: When joining Tencent as a consultant, Yao Shunyu had only one task: to investigate the reasons behind the long-term lag of the Hunyuan large model. The finding was: "In simple terms, almost every link is leaking." Hunyuan excessively pursued leaderboard results, contaminating the training data with test-prep materials. The model became good at exams but performed poorly in real-world scenarios. At that time, the accuracy acceptance threshold for data annotation was set at 95%, but in reality it remained at 60%-70% for a long time. Teams produced large amounts of unusable data just to meet deadlines, and the algorithm team tacitly allowed this.

X AI KOLs Timeline News

Summary

Tencent consultant Yao Shunyu investigated the lagging performance of the Hunyuan large model, finding systemic issues such as data contamination from chasing leaderboard scores and poor annotation quality.

When joining Tencent as a consultant, Yao Shunyu had only one task: to investigate the reasons behind the long-term lag of the Hunyuan large model. The finding was: "In simple terms, almost every link is leaking." Hunyuan excessively pursued leaderboard results, contaminating the training data with test-prep materials. The model became good at exams but performed poorly in real-world scenarios. At that time, the accuracy acceptance threshold for data annotation was set at 95%, but in reality it remained at 60%-70% for a long time. Teams produced large amounts of unusable data just to meet deadlines, and the algorithm team tacitly allowed this. From "LatePost".
Original Article
View Cached Full Text

Cached at: 07/14/26, 10:22 AM

When Yao Shunyu joined Tencent as a consultant, his only task was to diagnose the reasons behind the long‑term underperformance of Hunyuan, Tencent’s large‑scale AI model. His conclusion: “In simple terms, almost every link in the chain was leaking.” Hunyuan had been overly optimized for leaderboard performance. The training data was contaminated with leaderboard‑oriented content, making the model highly skilled at passing exams but poor in real‑world scenarios. The acceptance threshold for data annotation accuracy was set at 95%, but the actual accuracy had remained at only 60%–70% for a long time. To meet deadlines, the team produced large amounts of unusable data, and the algorithm team tacitly allowed this to happen.

From LatePost

Similar Articles

@latepostnews: 29-Year-Old Yao Shunyu Takes Over: 300 Days of Reforming Tencent Hunyuan - In 2024, Tencent's high-level recruitment team met Yao Shunyu at a top academic conference. At that time, the young man born in 1997 was still a researcher at OpenAI, and he was introduced to Tencent President Liu Chiping. A year later, he returned to China and became the head of Tencent's large language model...

X AI KOLs Timeline

Under Yao Shunyu's leadership, Tencent's Hunyuan large language model undergoes deep reforms: simplifying hierarchy, focusing on data quality, abandoning benchmark chasing, with a goal of entering the domestic first tier by 2027. The article details the changes Yao Shunyu drove within 300 days after parachuting into Tencent from OpenAI, including replacing key responsible persons, strengthening infrastructure, and promoting model-product co-design.

@MaxForAI: After reading the remarks of former Qwen researcher Hongyi Yuan @yiguyuan20, I finally understand why the foundational models from major Chinese tech companies can't make real progress... So pointing out mistakes or telling the truth requires seniority?

X AI KOLs Timeline

Remarks by former Qwen researcher Hongyi Yuan reveal a culture of seniority in foundational model R&D at major Chinese tech companies, which stifles newcomers from raising objections or innovating, hindering model breakthroughs.

@Sxy_Cherotich: Recently I've been talking with quite a few model researchers, and a consensus conclusion is: the importance of data is once again highlighted. A while ago I got to know ex-Kimi's @FanqingMengAI, who is doing a startup in the data direction, and invited him to record a podcast. The biggest non-consensus from our conversation is Fanqing's view on the difference between domestic and foreign models...

X AI KOLs Timeline

A podcast about AI model competition, discussing the importance of data, distillation and pre-training innovation, and an interview with Evolvent AI co-founder Meng Fanqing, covering topics such as synthetic data, RSI, and differences in domestic models.

@AYi_AInotes: A counter-intuitive judgment: 80% of Agent production crashes have nothing to do with model IQ — they're all from context overflow, tool misconfiguration, sub-agent runaway. The real watershed in 2026 is Harness and Loop, not the model. Bro, @wizardly_ai's engineering note...

X AI KOLs Timeline

This article points out that 80% of AI Agent production crashes are not due to model intelligence, but are caused by context overflow, tool misconfiguration, and sub-agent runaway. The author emphasizes that the watershed in 2026 lies in Harness (office systems, security) and Loop (automatic cycling mechanism), not the model itself.

@jakevin7: Current models all seem to have some serious issues—they've caught comment-mania syndrome.... This leads to big problems: lots of outdated information gets injected into the code as context, and this is done autonomously by agents, making it uncontrollable.

X AI KOLs Following

The author complains that current AI models generally have a tendency to over-comment, leading to outdated information being injected into code context, and this is done autonomously by agents, posing a risk of losing control.