SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
SLAI
Reports full-parameter post-training of trillion-parameter DeepSeek-V4 mixture-of-experts models on Huawei Ascend NPU clusters rather than GPUs, reaching 34.22% Model FLOPs Utilization (2.93x over the open-source baseline recipe) via hierarchical parallelism and kernel-level optimization, and using the resulting infrastructure to train an operations-research-specialized model reaching 71.81% zero-shot Pass@1.
Why it matters
Demonstrates that trillion-parameter-scale LLM post-training is now practical outside the Nvidia GPU ecosystem, a significant infrastructure milestone for non-Western AI compute independence.
Importance: 3/5
Notable research release: trillion-parameter-scale post-training infrastructure milestone.