Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

Listen to today's AI briefing

Daily podcast — 5 min, AI-narrated summary of top stories

Two large NVIDIA B200 server racks with blue LED lights in a data center, cables connected to a Moonshot AI Kimi K3…

Moonshot AI Releases 1.56T-Parameter Kimi K3, Requires 2x B200 Nodes

Moonshot AI released Kimi K3, a 1.56T parameter MoE model at 1561 GB, requiring 2x B200 nodes. No benchmarks disclosed.

·10h ago·3 min read··9 views·AI-Generated·Report error
Share:
What is Moonshot AI's Kimi K3 model and what hardware does it require?

Moonshot AI released Kimi K3, a 1.56 trillion parameter mixture-of-experts model at 1561 GB, requiring at least 2x NVIDIA B200 nodes for inference.

TL;DR

Kimi K3 is a 1.56 trillion parameter MoE model. · Requires at least 2x B200 nodes to run. · Weighs 1561 GB, pushing inference hardware limits.

Moonshot AI released Kimi K3, a 1.56 trillion parameter model at 1561 GB. The mixture-of-experts architecture requires at least 2x NVIDIA B200 nodes for inference.

Key facts

  • 1.56 trillion total parameters.
  • 1561 GB model weight.
  • Requires at least 2x NVIDIA B200 nodes.
  • Moonshot AI did not disclose training data or benchmarks.
  • Largest openly described Chinese AI model by parameter count.

Moonshot AI has released Kimi K3, a 1.56 trillion parameter mixture-of-experts model weighing 1561 GB. According to @_akhaliq, the model requires at least 2x NVIDIA B200 nodes to run inference, making it one of the largest openly described architectures to date.

At 1.56 trillion parameters, Kimi K3 is among the largest openly described dense or MoE architectures. The 1561 GB footprint means even with 8-bit quantization the model would still exceed 780 GB, firmly placing it in the multi-node inference tier. For comparison, OpenAI's GPT-4 is estimated at 1.7 trillion parameters, while Meta's Llama 3.1 405B requires only 800 GB at FP16 — roughly half Kimi K3's weight. DeepSeek's V3, another MoE model, clocks in at 671 billion total parameters with 37 billion activated per token.

Moonshot AI did not disclose training compute, data mix, or benchmark scores. The model's architecture — number of experts, active parameters per token, context window — remains unspecified. This lack of technical detail makes it difficult to assess whether the parameter count reflects genuine capacity or MoE overhead, where total parameters far exceed active parameters.

Hardware implications

The 2x B200 requirement signals both ambition and cost. Each B200 delivers 4.5 TB/s HBM3e bandwidth and 2.8 petaFLOPS of FP8 compute. Two nodes in NVLink-connected configuration provide roughly 1.4 TB of aggregate HBM — barely sufficient for the 1.5 TB model weight plus KV cache overhead. Inference latency and throughput figures are absent, but early adopters should expect significant engineering effort to achieve usable performance.

This release comes amid a broader trend of Chinese AI labs pushing frontier-scale models. DeepSeek's V3 and R1 series, ByteDance's Doubao, and Alibaba's Qwen 2.5 have all demonstrated competitive performance with fewer parameters. Kimi K3's massive footprint may reflect a deliberate strategy to maximize raw capacity, but without benchmark transparency, the trade-off between cost and capability remains opaque.

Key Takeaways

  • Moonshot AI released Kimi K3, a 1.56T parameter MoE model at 1561 GB, requiring 2x B200 nodes.
  • No benchmarks disclosed.

What to watch

Kimi K3: Moonshot AI's 2.8 Trillion Parameter Open-Source ...

Watch for Moonshot AI to release benchmark scores on standard evals like MMLU, HumanEval, and SWE-Bench. If Kimi K3 matches or exceeds GPT-4-class models despite the MoE overhead, it validates the scale-at-all-costs approach. If not, the 2x B200 requirement becomes a hardware tax with no corresponding accuracy gain.

Source: gentic.news · · author= · citation.json

AI-assisted reporting. Generated by gentic.news from multiple verified sources, fact-checked against the Living Graph of 4,300+ entities. Edited by Ala SMITH.

Following this story?

Get a weekly digest with AI predictions, trends, and analysis — free.

AI Analysis

Kimi K3's 1.56 trillion parameter count places it in the same weight class as GPT-4, but the comparison is misleading without knowing the active parameter count per token. MoE models like Mixtral 8x7B achieve strong results with only 12.9B active parameters out of 46.7B total. If Kimi K3 activates only 10-20% of its parameters per token, the effective compute per forward pass could be comparable to models one-tenth its size. The real question is whether the remaining parameters contribute meaningful capacity or merely inflate the hardware bill. The 1561 GB footprint is striking. Even with FP8 quantization, the model would require ~780 GB of HBM, still exceeding a single B200's 192 GB. This forces multi-node inference with NVLink interconnect, adding latency and complexity. For comparison, Meta's Llama 3.1 405B runs on a single H100 node with 8 GPUs. Kimi K3 demands twice the hardware for an unknown quality delta. Moonshot AI's silence on benchmarks is the most telling signal. In the current competitive landscape, labs typically release at least MMLU or HumanEval scores alongside model weights. The omission suggests either the model is still being evaluated, or the results do not justify the hardware investment. Until benchmarks emerge, the release reads more as a capability statement than a practical tool.
Compare side-by-side
Moonshot AI vs OpenAI

Mentioned in this article

Enjoyed this article?
Share:

AI Toolslive

Five one-click lenses on this article. Cached for 24h.

Pick a tool above to generate an instant lens on this article.

Related Articles

From the lab

The framework underneath this story

Every article on this site sits on top of one engine and one framework — both built by the lab.

More in Products & Launches

View all