Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

groq

30 articles about groq in AI news

NVIDIA Groq 3 LPX Hits 3,400 Tokens/s on Gemma 4 31B

NVIDIA's Groq 3 LPX claims 3,400 tokens/s on Gemma 4 31B, with Nebius first to deploy. The 4x responsiveness claim targets agent latency.

99% relevant

Nvidia to Acquire Groq; Leftover Biz Buys Blackwell Clusters

Nvidia reportedly acquiring Groq; leftover Groq rents B300/GB300/Rubin GPUs, many Blackwell-only, per SemiAnalysis.

85% relevant

TILERT Boosts Blackwell Decode 1.9x, Pressures Groq/Cerebras

TileRT AI's TILERT claims 1.9x decode interactivity on NVIDIA Blackwell at same cost, pressuring Groq/Cerebras/SambaNova. Claim unverified, but if real, erodes custom-chip value.

85% relevant

Jensen Huang Announces $20B Groq Integration, OpenClaw OS, and $50T+ Physical AI Market Vision on All-In Podcast

NVIDIA CEO Jensen Huang announced a ~$20B Groq integration ending GPU inference monopoly, launched OpenClaw OS for AI agents, and identified physical AI as a $50-70T market. He criticized Anthropic's 'doomer hype' and predicted NVIDIA's path to $1T+ revenue.

99% relevant

Groq's LPU Inference Engine Demonstrates 500+ Token/s Performance on Llama 3.1 70B

Groq's Language Processing Unit (LPU) inference engine achieves over 500 tokens/second on Meta's Llama 3.1 70B model, demonstrating significant performance gains for large language model inference.

85% relevant

Nvidia's Groq Ramps Up AI Chip Production with Samsung in Major Partnership Expansion

Nvidia's recent acquisition Groq has significantly expanded its partnership with Samsung, increasing chip orders from 9,000 to 30,000 wafers. This massive production boost signals accelerated development of Groq's specialized AI inference processors amid growing market demand.

85% relevant

Nvidia's Strategic Shift: Merging Groq Hardware in New AI Chip Targeting OpenAI

Nvidia is reportedly developing a new AI chip that combines its GPU technology with hardware from Groq, with OpenAI potentially becoming a major customer. This move signals Nvidia's recognition of specialized AI hardware beyond traditional GPUs.

95% relevant

SemiAnalysis: Can TileRT Software Match Cerebras on NVIDIA GPUs?

SemiAnalysis is testing TileRT InferenceX, software claiming batch-1 ultra-high interactivity on NVIDIA GPUs, targeting Cerebras, Groq LPU, and SambaNova. No benchmarks disclosed yet.

82% relevant

Building Intelligent Feedback Systems

A technical guide on building a customer review triage system using LangGraph, LangChain, Groq, and Pydantic. It explains how agentic workflows enable conditional routing based on sentiment analysis.

92% relevant

Inference shift opens door for AI chip startups to challenge Nvidia

Inference shift from training to serving creates opportunities for AI chip startups. Nvidia's $20B Groq acquihire validates disaggregated compute strategies.

96% relevant

Nvidia in Talks to Invest in Perplexity at $30B+ Valuation

Nvidia is negotiating a Perplexity investment at $30B+ valuation as Perplexity's ARR tripled to $750M, extending Nvidia's invest-to-sell-chips strategy.

99% relevant

Nvidia Pays $6B for Poolside's Model Factory, 109 Staff

Nvidia is paying $6B to license Poolside's Model Factory and hire 109 staff, plus a $1B investment at $12B pre-money. The deal avoids acquisition optics while deepening Nvidia's model-building push.

100% relevant

Nvidia in Talks to Acquire AI Chip Startup Rebellions

Nvidia is in talks to acquire Korean AI chip startup Rebellions, a deal that would be its largest AI chip acquisition. Terms undisclosed; the move targets inference accelerators amid rising competition.

85% relevant

OpenAI's Ultrafast Mode Hits 750 Tokens/s on GPT-5.6 Sol

OpenAI launched Ultrafast mode for GPT-5.6 Sol at 14x speed and 750 tokens/s, powered by Cerebras. Preview limited to select customers, targeting latency-sensitive enterprise workflows.

100% relevant

SambaNova SN50 MVP Runs MiniMax M2.7, But Batch Size Limit Looms

SambaNova's SN50 MVP runs MiniMax M2.7 but is stuck at batch size 2, highlighting software maturity issues for frontier models.

72% relevant

GPT-5.6 Sol on Cerebras Hits 750 Token/s

GPT-5.6 Sol on Cerebras claimed at 750 token/s, but no official data or model release exists. Unverified claim needs vendor confirmation.

97% relevant

SambaNova Hits 850 t/s on MiniMax M2.7 via Hybrid H200-RDU Inferencing

SambaNovaAI achieved 850 t/s on MiniMax M2.7 by pairing H200 GPUs for prefill with SN50 RDUs for decode at RAISE Paris.

86% relevant

Etched Hits $5B Valuation, $1B in Orders for AI Inference Chip

Etched hits $5B valuation with $1B in orders for TSMC-made inference chips, raising $500M from top investors. The startup targets Nvidia's dominance.

100% relevant

FreeLLMAPI Aggregates 1.7B Free Tokens/Month Across 11 Providers

FreeLLMAPI aggregates 11 free LLM providers into one endpoint, offering 1.7B tokens/month with automatic fallover. Reduces friction for side projects but faces provider tolerance risks.

75% relevant

Jim Keller: Tenstorrent IPO Looms as BlackHole Chip Scales

Jim Keller confirmed Tenstorrent's IPO plans as BlackHole chip scales for AI inference, competing with Nvidia. No revenue disclosed.

98% relevant

Tensordyne Claims 10x Efficiency Gain with Napier Architecture

Tensordyne claims 10x efficiency over Nvidia in inference with Napier gen, but lacks data or verification.

85% relevant

Qualcomm Launches AI Data Center Program With Hyperscaler Customer

Qualcomm launched an AI data center program with a major hyperscaler customer, targeting inference workloads. Financial terms and partner identity undisclosed.

85% relevant

Nvidia Buys Kumo AI for $400M to Predict from Business Data

Nvidia acquired Kumo AI for $400M+ to bring foundation model predictions to enterprise relational data, filling a gap left by LLMs.

88% relevant

Prism v1.8 Adds CLI, MCP Server, and SDKs — Here's How to Use Them with

Prism v1.8's MCP server gives Claude Code direct control over caches, budgets, and routing. Install it in 2 minutes and ditch the dashboard for terminal-based AI infrastructure management.

73% relevant

Karpathy: Neural nets will become the host, CPUs the co-processor

Karpathy predicts neural networks will become the host OS, with CPUs as co-processors, rendering most classical app interfaces obsolete.

85% relevant

Cerebras Challenges Nvidia Inference Monopoly with Wafer-Scale Edge

Cerebras is challenging Nvidia's inference dominance with wafer-scale chips, as inference workloads surpass training in AI compute spend.

70% relevant

Perplexity Claims 3x Blackwell Inference Throughput for 70B Models

Perplexity AI claims 3x inference throughput for 70B models on Nvidia Blackwell GPUs via FP4 and custom scheduling. The gain exceeds Nvidia's own 2x marketing claim.

85% relevant

Google Opens TPU Sales to Select Customers, Raises Capex Forecast

Google sells TPUs to select customers, raising capex forecast for Q1 FY2026, monetizing in-house chips beyond Cloud.

100% relevant

AI Inference Costs Drop 5-10x Yearly: @kimmonismus Challenges Forbes

@kimmonismus claims AI inference costs drop 5-10x yearly, challenging Forbes' static compute cost narrative. This deflation rate implies rapid TCO reduction for enterprise deployments.

75% relevant

Retail traffic from LLMs surged 393% year-on-year, reports CX Network

According to CX Network, retail traffic originating from large language model interfaces increased 393% year-on-year, highlighting the growing role of conversational AI as a customer acquisition channel for retailers.

86% relevant