Moonshot AI released Kimi K3, a 1.56 trillion parameter model at 1561 GB. The mixture-of-experts architecture requires at least 2x NVIDIA B200 nodes for inference.
Key facts
- 1.56 trillion total parameters.
- 1561 GB model weight.
- Requires at least 2x NVIDIA B200 nodes.
- Moonshot AI did not disclose training data or benchmarks.
- Largest openly described Chinese AI model by parameter count.
Moonshot AI has released Kimi K3, a 1.56 trillion parameter mixture-of-experts model weighing 1561 GB. According to @_akhaliq, the model requires at least 2x NVIDIA B200 nodes to run inference, making it one of the largest openly described architectures to date.
At 1.56 trillion parameters, Kimi K3 is among the largest openly described dense or MoE architectures. The 1561 GB footprint means even with 8-bit quantization the model would still exceed 780 GB, firmly placing it in the multi-node inference tier. For comparison, OpenAI's GPT-4 is estimated at 1.7 trillion parameters, while Meta's Llama 3.1 405B requires only 800 GB at FP16 — roughly half Kimi K3's weight. DeepSeek's V3, another MoE model, clocks in at 671 billion total parameters with 37 billion activated per token.
Moonshot AI did not disclose training compute, data mix, or benchmark scores. The model's architecture — number of experts, active parameters per token, context window — remains unspecified. This lack of technical detail makes it difficult to assess whether the parameter count reflects genuine capacity or MoE overhead, where total parameters far exceed active parameters.
Hardware implications
The 2x B200 requirement signals both ambition and cost. Each B200 delivers 4.5 TB/s HBM3e bandwidth and 2.8 petaFLOPS of FP8 compute. Two nodes in NVLink-connected configuration provide roughly 1.4 TB of aggregate HBM — barely sufficient for the 1.5 TB model weight plus KV cache overhead. Inference latency and throughput figures are absent, but early adopters should expect significant engineering effort to achieve usable performance.
This release comes amid a broader trend of Chinese AI labs pushing frontier-scale models. DeepSeek's V3 and R1 series, ByteDance's Doubao, and Alibaba's Qwen 2.5 have all demonstrated competitive performance with fewer parameters. Kimi K3's massive footprint may reflect a deliberate strategy to maximize raw capacity, but without benchmark transparency, the trade-off between cost and capability remains opaque.
Key Takeaways
- Moonshot AI released Kimi K3, a 1.56T parameter MoE model at 1561 GB, requiring 2x B200 nodes.
- No benchmarks disclosed.
What to watch

Watch for Moonshot AI to release benchmark scores on standard evals like MMLU, HumanEval, and SWE-Bench. If Kimi K3 matches or exceeds GPT-4-class models despite the MoE overhead, it validates the scale-at-all-costs approach. If not, the 2x B200 requirement becomes a hardware tax with no corresponding accuracy gain.








