Huawei's Ascend SuperPOD delivers roughly 1.3x to 1.7x lower decode throughput per GPU than Nvidia's GB300, according to analyst @zephyr_z9. The gap narrows from a 4:1 training FLOPS ratio due to memory bandwidth constraints and sharding strategies.
Key facts
- Training FLOPS ratio: 4:1 (Huawei vs Nvidia)
- Memory bandwidth ratio: 2:1 (8TB/s vs 4TB/s)
- Decode throughput gap: 1.3x to 1.7x
- Huawei's scale-up world size enables aggressive sharding
A detailed technical analysis by @zephyr_z9 on X reveals that the performance gap between Huawei's Ascend SuperPOD and Nvidia's GB300 is highly workload-dependent. The widely cited 4:1 exchange ratio for training FLOPS does not hold for inference.
Memory Bandwidth Bottleneck
On a memory bandwidth basis, the ratio falls to 2:1, with Huawei's SuperPOD delivering approximately 8 TB/s versus Nvidia's 4 TB/s [according to @zephyr_z9]. This bandwidth advantage partially compensates for the raw compute deficit during autoregressive decoding, where memory bandwidth is the primary constraint.
Sharding Strategy Advantage
Huawei's massive scale-up world size enables aggressive sharding strategies that Nvidia cannot match with its smaller node configurations. This architectural flexibility allows the Ascend SuperPOD to distribute memory-bound operations more efficiently, narrowing the decode throughput gap.
Real-World Decode Performance
The analyst estimates the decode throughput per GPU difference between the two systems will be in the 1.3x-1.7x range [per @zephyr_z9]. This is significantly closer than the 4x training FLOPS gap would suggest, making the Ascend SuperPOD a more competitive inference platform than its training numbers imply. The exact figure depends on model architecture, batch size, and sharding configuration.
Implications for Inference Workloads
For deployments dominated by inference (e.g., chatbots, code generation, real-time translation), the Ascend SuperPOD's decode throughput gap of 1.3-1.7x versus GB300 means it can serve a large fraction of workloads at competitive latency, especially when combined with aggressive sharding. The 4x training gap remains a significant disadvantage for model development, but inference-heavy deployments may find the SuperPOD viable.
What to watch

Watch for benchmark results from Huawei's Ascend SuperPOD at scale, particularly SWE-Bench or MT-Bench inference latency numbers. Also monitor whether Nvidia responds with larger-scale-up configurations in the Vera Rubin architecture to close the sharding advantage.
[Updated 24 Jul via gn_gpu_cluster]
A SemiAnalysis report on Vera Rubin NVL72 indicates a 10x inference improvement over Grace Blackwell, which could widen the decode throughput gap with Huawei's Ascend SuperPOD beyond the current 1.3-1.7x estimate [per SemiAnalysis]. The new architecture's larger scale-up domain may also counter Huawei's sharding advantage, shifting the competitive landscape for inference workloads.








