SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 4 days ago • 56
Janus: Disaggregating Attention and Experts for Scalable MoE Inference Paper • 2512.13525 • Published Dec 15, 2025 • 6
Janus: Disaggregating Attention and Experts for Scalable MoE Inference Paper • 2512.13525 • Published Dec 15, 2025 • 6