@LiorOnAI: Hy3 spent less time chasing another benchmark point and more time fixing the things that make agents quietly fail. Tool…

X AI KOLs Following Models

Summary

Tencent released Hy3, a 295B MoE model focused on practical reliability for agentic tasks, with open-source Apache 2.0 license and a free API for two weeks.

Hy3 spent less time chasing another benchmark point and more time fixing the things that make agents quietly fail. Tool-call recovery. Output formats. Multi-turn constraint tracking. Hallucinations. Token efficiency. Those don't usually move leaderboard positions much. They decide whether a workflow finishes without a human stepping in. Tencent says it evaluated the model with 270 domain experts doing work from their own jobs instead of synthetic benchmark tasks. That's a direction I'd like to see more labs take, even if every company will naturally design those evaluations differently. The more models converge on reasoning ability, the more the competition shifts toward failure rates across thousands of small interactions. Nobody notices the model that solves one more math problem. Everyone notices the model that silently drops a tool call on the 47th step. That's also why token efficiency matters more than it did a year ago. A model that finishes the same workflow with half the tokens changes deployment economics without changing the user's experience much.
Original Article
View Cached Full Text

Cached at: 07/07/26, 06:15 AM

Hy3 spent less time chasing another benchmark point and more time fixing the things that make agents quietly fail.

Tool-call recovery. Output formats. Multi-turn constraint tracking. Hallucinations. Token efficiency.

Those don’t usually move leaderboard positions much. They decide whether a workflow finishes without a human stepping in.

Tencent says it evaluated the model with 270 domain experts doing work from their own jobs instead of synthetic benchmark tasks. That’s a direction I’d like to see more labs take, even if every company will naturally design those evaluations differently.

The more models converge on reasoning ability, the more the competition shifts toward failure rates across thousands of small interactions.

Nobody notices the model that solves one more math problem.

Everyone notices the model that silently drops a tool call on the 47th step.

That’s also why token efficiency matters more than it did a year ago. A model that finishes the same workflow with half the tokens changes deployment economics without changing the user’s experience much.

Tencent Hy (@TencentHunyuan): 🚀Hy3 is here.

295B MoE. Best in its size class. Rivals trillion-scale flagships. Reliable and affordable for most agentic usecases. Apache 2.0. Friendly for commercial use. FREE API for 2 weeks → https://t.co/EyURKwTdgi

🤗 https://t.co/twqJpqb2SL 📖

Similar Articles

Hy3 (1 minute read)

TLDR AI

Tencent released Hy3, a 295B-parameter MoE model with 21B active parameters, outperforming similar-sized models and rivaling larger open-source models. It is Apache 2.0 licensed, available on Hugging Face and free on OpenRouter until July 21st.

HY-3 PREVIEW

Reddit r/LocalLLaMA

Tencent releases Hy3-preview, a 295B-parameter MoE model with 21B active parameters that excels in STEM reasoning, instruction following, coding and agent tasks.