Xing4.0-29B-A4B is a next-generation MoE large language model developed by China Telecom, featuring 29B total parameters with 4B active per token, native support for 256K context length, and optimization for Ascend NPU with agent-oriented architecture for complex engineering tasks.
I find another new model at HF: https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B "Xing4.0-29B-A4B is a next-generation large language model in the Xing series (formerly TeleChat), developed by China Telecom Artificial Intelligence Technology Co., Ltd. With 29B total parameters and only 4B activated per token, it natively supports a 256K context length, extensible to 512K. It is the first model of this scale trained entirely on the Ascend NPU platform with the MindSpore framework, and deeply optimized for complex engineering tasks. For more information, please refer to our GitHub repository. Highlights Agent-Oriented Architecture: Built on the mHC + MLA + MTP architecture, supporting multi-step planning, tool calling, and complex reasoning chain execution, ensuring task coherence and execution stability under long contexts. Deep Co-optimization with Ascend NPU: Adapted for Ascend 910C clusters using MindSpore/MindFormers, including feature adaptation for mHC and fused operator development, enabling stable and efficient training on the Ascend platform. Significant Training Efficiency Gains: Through multi-level co-optimization — including fine-grained MoE communication optimization, selective recomputation, DVM automatic graph-operator fusion, and Ascend C mHC fused operators — overall training throughput was improved by approximately 96% over out-of-the-box performance. Full Open-Source Ecosystem Compatibility: Supports LLaMA-Factory and MindFormers for fine-tuning; SGLang, vLLM, and KTransformers for inference and deployment; with targeted adaptation and format alignment for agent frameworks such as OpenCode, Claude Code, OpenClaw, and Hermes, enabling seamless integration into existing workflows. Easy Adaptation for Domain-Specific Scenarios: The model is well-suited for downstream task fine-tuning, allowing lightweight customization on proprietary data for vertical domains such as intent classification, table understanding, contract auditing, and knowledge-based QA, enabling rapid domain capability development and deployment at low cost." Parameters 29B (4B active) Number of Layers 40 Hidden Size 3584 Dense Intermediate Size 9216 Expert Intermediate Size 1024 Attention Type MLA Number of Routed Experts 64 Active Experts per Token 4 Number of Shared Experts 1 Context Length 256K (extensible to 512K) Benchmark Benchmark Xing4.0-29B-A4B Gemma4-26B-A4B Qwen3.6-35B-A3B IFBench 69.67 72.67 65.50 AIME2026 90.00 88.30 92.70 AA.LCR 61.00 66.00 62.00 Tau3-Bench 64.63 58.90 67.20 Claw-Eval 76.55 71.49 74.54 SWE-bench Verified 75.00 53.00 76.00 Terminal-Bench 2.1 57.50 30.00 51.50 SWE-bench Multilingual 66.00 51.00 67.20 DeepresearchBII 60.80 39.30 59.70
Xing4.0-29B-A4B is an open-source 29B-parameter large language model optimized for agent tasks and Ascend NPU, featuring a MoE architecture and achieving high training efficiency with competitive benchmark results.
Xiaomi releases MiMo-V2.5-Pro, an open-source MoE language model with 1.02T total parameters and 1M token context, optimized for complex agentic and software engineering tasks.
LongCat-2.0 is a large-scale MoE language model with 1.6 trillion total parameters and ~48B activated per token, trained on AI ASIC superpods with 1M-context data. It achieves strong performance on coding and agentic tasks.
LG AI Research introduces K-EXAONE 2.0, a frontier-scale multilingual MoE model with 750B total parameters (37B active), upcycled and trained for advanced reasoning, agentic workflows, and long-context understanding, released under Apache 2.0.
AntLingAGI releases Ling-3.0-flash, a 124B-parameter MoE model with 5.1B active parameters, now free for a week on Nous Portal. It matches or beats their 1T flagship on many benchmarks, designed for agent workloads like coding and tool use.