SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Hugging Face Daily Papers Papers

Summary

This paper presents SLAI T-Rex, a full-parameter post-training optimization framework for trillion-parameter MoE models on Ascend NPU SuperPOD, achieving 34.22% MFU and outperforming GPT-5.4-Mini on Operations Research tasks by 3.98 percentage points.

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.
Original Article
View Cached Full Text

Cached at: 07/24/26, 05:09 AM

Paper page - SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Source: https://huggingface.co/papers/2607.20145 Published on Jul 22

#1 Paper of the day Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Full-parameterpost-trainingoftrillion-parameter-scaleMoEmodelsintroducessubstantialsystem-levelchallengesforlarge-scaledistributedtraining,includingseverememorypressure,non-overlappedcommunicationoverhead,andinefficientkernelexecution.Whilemostlarge-scaleLLMtrainingsystemsarebuiltaroundGPU-basedclusters,thisreportpresentsanend-to-endoptimizationpracticeontheAscendNPUSuperPOD.UsingtheDeepSeek-V4modelfamilyasthetargetworkload,wedevelopahierarchicaloptimizationframeworkspanningmodel-levelparallelism,computation-communicationorchestration,andlow-levelkernelexecution.Theresultingsystemachieves34.22%ModelFLOPsUtilization(MFU)witha2.93ximprovementovertheopen-sourcebaselinerecipewhilemaintainingtrainingstability.Buildingonthisoptimizedinfrastructure,wefurtherestablishaCPTandSFTworkflowforcomplexOperationsResearch(OR)tasks.WerefertotheintegratedframeworkasSLAIT-Rex.UsingDeepSeek-V4-Flash,wedevelopOR-orientedCPTandSFTdatapipelinesthatcombinecollecteddomainresourceswithsolver-verifiedsyntheticoptimizationdocuments.Theresultingdatasetcontains10Khigh-qualitySFTsamplesspanningfourtaskcategoriesandthreeproblemrepresentations.Thespecializedmodelachievesthehighestaveragezero-shotPass@1scoreamongtheevaluatedmodels,reaching71.81%andoutperformingGPT-5.4-MiniandthebaseDeepSeek-V4-Flashmodelby3.98and11.27percentagepoints,respectively.Overall,thisworkdemonstratesafull-stackpathwayfromefficienttrillion-parametermodelpost-trainingonAscendinfratodomain-specializedFlashmodelsforsolver-groundedmathematicalmodeling,advancingfrontier-modelsystemsforcomplexreasoning.

View arXiv pageView PDFGitHub19Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.20145 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.20145 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.20145 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

DeepSeek V4 paper full version is out, FP4 QAT details and stability tricks [D]

Reddit r/MachineLearning

DeepSeek released the full V4 paper detailing FP4 quantization-aware training, MoE training stability tricks (anticipatory routing and SwiGLU clamping), and a generative reward model for RLHF, achieving dramatic efficiency gains—V4-Flash uses only 10% of V3.2's FLOPs and 7% of its KV cache at 1M context length.

Pushing the Limits of Serving DeepSeek-V4-Pro (28 minute read)

TLDR AI

The article presents a methodology for optimizing the serving of DeepSeek-V4-Pro, a 1.6-trillion-parameter MoE model, on H20 GPUs, achieving significant performance improvements through scenario-specific configurations and optimizations.