groq
30 articles about groq in AI news
NVIDIA Groq 3 LPX Hits 3,400 Tokens/s on Gemma 4 31B
NVIDIA's Groq 3 LPX claims 3,400 tokens/s on Gemma 4 31B, with Nebius first to deploy. The 4x responsiveness claim targets agent latency.
Nvidia to Acquire Groq; Leftover Biz Buys Blackwell Clusters
Nvidia reportedly acquiring Groq; leftover Groq rents B300/GB300/Rubin GPUs, many Blackwell-only, per SemiAnalysis.
TILERT Boosts Blackwell Decode 1.9x, Pressures Groq/Cerebras
TileRT AI's TILERT claims 1.9x decode interactivity on NVIDIA Blackwell at same cost, pressuring Groq/Cerebras/SambaNova. Claim unverified, but if real, erodes custom-chip value.
Jensen Huang Announces $20B Groq Integration, OpenClaw OS, and $50T+ Physical AI Market Vision on All-In Podcast
NVIDIA CEO Jensen Huang announced a ~$20B Groq integration ending GPU inference monopoly, launched OpenClaw OS for AI agents, and identified physical AI as a $50-70T market. He criticized Anthropic's 'doomer hype' and predicted NVIDIA's path to $1T+ revenue.
Groq's LPU Inference Engine Demonstrates 500+ Token/s Performance on Llama 3.1 70B
Groq's Language Processing Unit (LPU) inference engine achieves over 500 tokens/second on Meta's Llama 3.1 70B model, demonstrating significant performance gains for large language model inference.
Nvidia's Groq Ramps Up AI Chip Production with Samsung in Major Partnership Expansion
Nvidia's recent acquisition Groq has significantly expanded its partnership with Samsung, increasing chip orders from 9,000 to 30,000 wafers. This massive production boost signals accelerated development of Groq's specialized AI inference processors amid growing market demand.
Nvidia's Strategic Shift: Merging Groq Hardware in New AI Chip Targeting OpenAI
Nvidia is reportedly developing a new AI chip that combines its GPU technology with hardware from Groq, with OpenAI potentially becoming a major customer. This move signals Nvidia's recognition of specialized AI hardware beyond traditional GPUs.
SemiAnalysis: Can TileRT Software Match Cerebras on NVIDIA GPUs?
SemiAnalysis is testing TileRT InferenceX, software claiming batch-1 ultra-high interactivity on NVIDIA GPUs, targeting Cerebras, Groq LPU, and SambaNova. No benchmarks disclosed yet.
Building Intelligent Feedback Systems
A technical guide on building a customer review triage system using LangGraph, LangChain, Groq, and Pydantic. It explains how agentic workflows enable conditional routing based on sentiment analysis.
Inference shift opens door for AI chip startups to challenge Nvidia
Inference shift from training to serving creates opportunities for AI chip startups. Nvidia's $20B Groq acquihire validates disaggregated compute strategies.
Nvidia in Talks to Invest in Perplexity at $30B+ Valuation
Nvidia is negotiating a Perplexity investment at $30B+ valuation as Perplexity's ARR tripled to $750M, extending Nvidia's invest-to-sell-chips strategy.
Nvidia Pays $6B for Poolside's Model Factory, 109 Staff
Nvidia is paying $6B to license Poolside's Model Factory and hire 109 staff, plus a $1B investment at $12B pre-money. The deal avoids acquisition optics while deepening Nvidia's model-building push.
Nvidia in Talks to Acquire AI Chip Startup Rebellions
Nvidia is in talks to acquire Korean AI chip startup Rebellions, a deal that would be its largest AI chip acquisition. Terms undisclosed; the move targets inference accelerators amid rising competition.
OpenAI's Ultrafast Mode Hits 750 Tokens/s on GPT-5.6 Sol
OpenAI launched Ultrafast mode for GPT-5.6 Sol at 14x speed and 750 tokens/s, powered by Cerebras. Preview limited to select customers, targeting latency-sensitive enterprise workflows.
SambaNova SN50 MVP Runs MiniMax M2.7, But Batch Size Limit Looms
SambaNova's SN50 MVP runs MiniMax M2.7 but is stuck at batch size 2, highlighting software maturity issues for frontier models.
GPT-5.6 Sol on Cerebras Hits 750 Token/s
GPT-5.6 Sol on Cerebras claimed at 750 token/s, but no official data or model release exists. Unverified claim needs vendor confirmation.
SambaNova Hits 850 t/s on MiniMax M2.7 via Hybrid H200-RDU Inferencing
SambaNovaAI achieved 850 t/s on MiniMax M2.7 by pairing H200 GPUs for prefill with SN50 RDUs for decode at RAISE Paris.
Etched Hits $5B Valuation, $1B in Orders for AI Inference Chip
Etched hits $5B valuation with $1B in orders for TSMC-made inference chips, raising $500M from top investors. The startup targets Nvidia's dominance.
FreeLLMAPI Aggregates 1.7B Free Tokens/Month Across 11 Providers
FreeLLMAPI aggregates 11 free LLM providers into one endpoint, offering 1.7B tokens/month with automatic fallover. Reduces friction for side projects but faces provider tolerance risks.
Jim Keller: Tenstorrent IPO Looms as BlackHole Chip Scales
Jim Keller confirmed Tenstorrent's IPO plans as BlackHole chip scales for AI inference, competing with Nvidia. No revenue disclosed.
Tensordyne Claims 10x Efficiency Gain with Napier Architecture
Tensordyne claims 10x efficiency over Nvidia in inference with Napier gen, but lacks data or verification.
Qualcomm Launches AI Data Center Program With Hyperscaler Customer
Qualcomm launched an AI data center program with a major hyperscaler customer, targeting inference workloads. Financial terms and partner identity undisclosed.
Nvidia Buys Kumo AI for $400M to Predict from Business Data
Nvidia acquired Kumo AI for $400M+ to bring foundation model predictions to enterprise relational data, filling a gap left by LLMs.
Prism v1.8 Adds CLI, MCP Server, and SDKs — Here's How to Use Them with
Prism v1.8's MCP server gives Claude Code direct control over caches, budgets, and routing. Install it in 2 minutes and ditch the dashboard for terminal-based AI infrastructure management.
Karpathy: Neural nets will become the host, CPUs the co-processor
Karpathy predicts neural networks will become the host OS, with CPUs as co-processors, rendering most classical app interfaces obsolete.
Cerebras Challenges Nvidia Inference Monopoly with Wafer-Scale Edge
Cerebras is challenging Nvidia's inference dominance with wafer-scale chips, as inference workloads surpass training in AI compute spend.
Perplexity Claims 3x Blackwell Inference Throughput for 70B Models
Perplexity AI claims 3x inference throughput for 70B models on Nvidia Blackwell GPUs via FP4 and custom scheduling. The gain exceeds Nvidia's own 2x marketing claim.
Google Opens TPU Sales to Select Customers, Raises Capex Forecast
Google sells TPUs to select customers, raising capex forecast for Q1 FY2026, monetizing in-house chips beyond Cloud.
AI Inference Costs Drop 5-10x Yearly: @kimmonismus Challenges Forbes
@kimmonismus claims AI inference costs drop 5-10x yearly, challenging Forbes' static compute cost narrative. This deflation rate implies rapid TCO reduction for enterprise deployments.
Retail traffic from LLMs surged 393% year-on-year, reports CX Network
According to CX Network, retail traffic originating from large language model interfaces increased 393% year-on-year, highlighting the growing role of conversational AI as a customer acquisition channel for retailers.