SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Summary
This paper presents SLAI T-Rex, a full-parameter post-training optimization framework for trillion-parameter MoE models on Ascend NPU SuperPOD, achieving 34.22% MFU and outperforming GPT-5.4-Mini on Operations Research tasks by 3.98 percentage points.
View Cached Full Text
Cached at: 07/24/26, 05:09 AM
Paper page - SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Source: https://huggingface.co/papers/2607.20145 Published on Jul 22
#1 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Full-parameterpost-trainingoftrillion-parameter-scaleMoEmodelsintroducessubstantialsystem-levelchallengesforlarge-scaledistributedtraining,includingseverememorypressure,non-overlappedcommunicationoverhead,andinefficientkernelexecution.Whilemostlarge-scaleLLMtrainingsystemsarebuiltaroundGPU-basedclusters,thisreportpresentsanend-to-endoptimizationpracticeontheAscendNPUSuperPOD.UsingtheDeepSeek-V4modelfamilyasthetargetworkload,wedevelopahierarchicaloptimizationframeworkspanningmodel-levelparallelism,computation-communicationorchestration,andlow-levelkernelexecution.Theresultingsystemachieves34.22%ModelFLOPsUtilization(MFU)witha2.93ximprovementovertheopen-sourcebaselinerecipewhilemaintainingtrainingstability.Buildingonthisoptimizedinfrastructure,wefurtherestablishaCPTandSFTworkflowforcomplexOperationsResearch(OR)tasks.WerefertotheintegratedframeworkasSLAIT-Rex.UsingDeepSeek-V4-Flash,wedevelopOR-orientedCPTandSFTdatapipelinesthatcombinecollecteddomainresourceswithsolver-verifiedsyntheticoptimizationdocuments.Theresultingdatasetcontains10Khigh-qualitySFTsamplesspanningfourtaskcategoriesandthreeproblemrepresentations.Thespecializedmodelachievesthehighestaveragezero-shotPass@1scoreamongtheevaluatedmodels,reaching71.81%andoutperformingGPT-5.4-MiniandthebaseDeepSeek-V4-Flashmodelby3.98and11.27percentagepoints,respectively.Overall,thisworkdemonstratesafull-stackpathwayfromefficienttrillion-parametermodelpost-trainingonAscendinfratodomain-specializedFlashmodelsforsolver-groundedmathematicalmodeling,advancingfrontier-modelsystemsforcomplexreasoning.
View arXiv pageView PDFGitHub19Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.20145 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.20145 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.20145 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
DeepSeek V4 paper full version is out, FP4 QAT details and stability tricks [D]
DeepSeek released the full V4 paper detailing FP4 quantization-aware training, MoE training stability tricks (anticipatory routing and SwiGLU clamping), and a generative reward model for RLHF, achieving dramatic efficiency gains—V4-Flash uses only 10% of V3.2's FLOPs and 7% of its KV cache at 1M context length.
Developing open source LLM from ground up from pretrain - rlhf(PPO/GRPO)
A developer shares progress on training a 7B parameter open source LLM from scratch using a DeepSeek architecture optimized for low VRAM, with the goal of democratizing AI development and eventually surpassing large proprietary models.
Deepseek training 2T and plans 8T model
DeepSeek is training a 2 trillion-parameter AI model and plans to eventually build an 8 trillion-parameter model.
Pushing the Limits of Serving DeepSeek-V4-Pro (28 minute read)
The article presents a methodology for optimizing the serving of DeepSeek-V4-Pro, a 1.6-trillion-parameter MoE model, on H20 GPUs, achieving significant performance improvements through scenario-specific configurations and optimizations.
DeepSeek V4.1 Flash is getting surprisingly close to GPT-5.6 Sol territory, while being absurdly cheap
DeepSeek has released V4.1 Flash, a 552B MoE model with efficient active parameters, achieving performance close to GPT-5.6 Sol at a much lower cost.