SambaNova's SN50 MVP demo of MiniMax M2.7 is limited to batch size 2. SemiAnalysis notes the software stack is not yet ready for larger batches or more frontier models.
Key facts
- SN50 MVP demo of MiniMax M2.7 model
- Software stack limited to batch size 2
- SemiAnalysis notes stack not yet for frontier models
- SN50 peak theoretical performance: 2 PFLOPS FP8
- MiniMax M2.7 is a 2.7B parameter dense model
SambaNova has achieved a milestone with its SN50 MVP, demoing MiniMax's M2.7 model According to @SemiAnalysis_. However, the celebration is tempered by a significant constraint: the software stack only functions at batch size 2. SemiAnalysis did not disclose whether this limitation is hardware or software-bound, but it suggests the SN50's compiler or runtime is still in early optimization stages.
The M2.7 model, part of MiniMax's family, is a dense 2.7B parameter model. Running it at batch size 2 severely limits throughput, likely to a few hundred tokens per second, making it unsuitable for production inference workloads. SemiAnalysis explicitly notes the stack must improve to handle 'more frontier models,' implying the SN50's current software is model-specific and not generalizable.
The Software Stack Gap

SambaNova's SN50 is a dataflow architecture, requiring custom compiler mappings for each model. Unlike Nvidia's CUDA ecosystem, which benefits from decades of optimization for transformer architectures, SambaNova's software team is essentially building from scratch. The batch size 2 cap suggests the compiler cannot yet exploit parallelism across larger batches, which is critical for cost-effective inference.
SemiAnalysis, a respected hardware analysis firm, praises the team's effort but frames the limitation as a known challenge: 'Excited for when the software stack will work above batch size 2.' This implies the hardware is capable, but the software is the bottleneck. The SN50's peak theoretical performance is 2 PFLOPS at FP8, but achieving that requires software maturity.
Competitive Context
Groq's LPU, another custom inference accelerator, supports batch sizes up to 64 for Llama 3 70B with its compiler. Cerebras's CS-3 supports variable batch sizes for GPT-3 class models. SambaNova's batch size 2 is a stark contrast, underscoring the software gap. The company has not publicly disclosed a roadmap for larger batch support or additional model support.
SambaNova previously announced a partnership with Argonne National Laboratory for scientific computing, but inference at scale remains elusive. The M2.7 demo is a proof of concept, not a production-ready deployment.
What to watch
Watch for SambaNova's next software release, likely by Q3 2026, targeting batch size 8+ support. If the compiler cannot scale, expect hardware redesigns or a pivot to specialized verticals like single-stream inference.








