Chinese AI ecosystem coalesces around MoE with wide expert parallelism. Hardware developers now design 'supernode' systems for large domains.
Key facts
- Chinese ecosystem standardizes on MoE architecture.
- Wide expert parallelism enables operation on weaker NPUs.
- Hardware developers now designing 'supernode' systems.
- Huawei, Enflame, Biren building high-radix topologies.
- MoE sparsity may narrow compute gap with US.
The Chinese AI ecosystem has coalesced around mixture-of-experts (MoE) architectures with wide expert parallelism, according to @teortaxestex. This design choice is driven by the need to survive on weaker domestic NPUs, which lack the raw compute of Nvidia's H100 or B200. "It's not just Huawei," the source notes—hardware developers are now designing systems with large domains, with everyone needing a 'supernode' now.
The shift reflects a structural constraint: Chinese AI labs cannot access cutting-edge Western accelerators due to export controls. MoE, which activates only a subset of parameters per token, reduces compute demands per forward pass. Wide expert parallelism spreads these experts across many NPUs, trading inter-chip communication for per-chip memory and compute savings. This allows models like DeepSeek's V2 and Qwen's MoE variants to scale to hundreds of billions of parameters on Huawei Ascend 910B or Cambricon MLU370 clusters.
The 'supernode' trend extends beyond chip design to system architecture. Huawei's CloudEngine switches and Ascend cluster topologies now support up to 2,000+ NPU interconnects per domain, while startups like Enflame and Biren are building similar high-radix topologies. This mirrors Nvidia's DGX SuperPOD approach but tailored for Chinese supply chains and lower per-chip performance.
Implications for global AI competition
China's MoE standardization creates a differentiated path from the US, where dense models (GPT-4, Gemini 1.5 Pro) dominate. MoE's sparsity advantage could narrow the compute gap—a 100B-parameter MoE model on 2,000 NPUs may match a 400B dense model on 1,000 H100s for certain tasks, especially inference. However, training efficiency remains a concern: MoE's communication overhead during training can offset NPU gains. Chinese labs are reportedly developing custom all-to-all communication primitives to address this.
The hardware shift also pressures Nvidia's China-specific A800/H800 products, which already face reduced bandwidth. If Chinese 'supernodes' achieve competitive throughput via parallelism, demand for Nvidia's sanctioned chips could decline further.
Key Takeaways
- China's AI ecosystem standardizes on MoE with wide expert parallelism to survive on weaker NPUs.
- Hardware makers now design 'supernode' systems.
What to watch
Watch for benchmark results from DeepSeek or Qwen showing MoE training throughput on 2,000+ NPU clusters vs. Nvidia H100 equivalents. Also monitor export control updates—any tightening could accelerate Chinese supernode adoption.
[Updated 22 Jul via scmp_tech]
US Treasury Secretary Scott Bessent flagged potential sanctions against Chinese AI models using stolen US intellectual property, stating on Fox Business that 'if we see overseas models are stealing from our great companies, we have the ability to sanction them' [per SCMP]. This threat targets low-cost systems like Moonshot AI's Kimi K3, which revived debate over China's architectural innovation closing the compute gap.







