gpu
30 articles about gpu in AI news
Anthropic RSI Claim Under Fire After GPU Math Disputed
Analyst accuses Anthropic of overstating RSI progress, citing GPU abundance and vague '<2X' metric. Credibility gap highlighted.
OpenAI Loses GPT-3 GPU Builder Scott Gray; 13 Leaders Out in 2026
Scott Gray, OpenAI's founding GPU engineer, left in 2026—the 13th senior departure. The exodus spans every function, signaling systemic churn beyond the C-suite.
SemiAnalysis: Can TileRT Software Match Cerebras on NVIDIA GPUs?
SemiAnalysis is testing TileRT InferenceX, software claiming batch-1 ultra-high interactivity on NVIDIA GPUs, targeting Cerebras, Groq LPU, and SambaNova. No benchmarks disclosed yet.
Nvidia Vera Rubin Shifts AI Strategy Beyond Raw GPU Speed
Nvidia's Vera Rubin architecture pivots from raw GPU FLOPS to system-level AI infrastructure, targeting memory bandwidth and interconnect bottlenecks that constrain large-scale model training.
The 'pytorch_no_powerplant_blowup' env var keeping frontier GPUs alive
A tweet reveals a PyTorch env var that throttles GPU power to prevent grid failures, exposing the real-world constraints of frontier AI infrastructure.
AMD to Supply Anthropic with 2GW of MI450 GPUs, Invest Up to $5B
AMD to supply Anthropic with 2GW MI450 GPUs in H1 2027 and invest up to $5B, challenging Nvidia's AI hardware dominance.
Moonshot AI Pauses K3 Subscriptions as Demand Exceeds GPU Capacity
Moonshot AI paused Kimi K3 subscriptions due to GPU capacity limits. The open-weight release by July 27 aims to offload compute demand.
CEO of Top AI Company Admits Begging for GPUs: Shortage Is Structural
The CEO of the most valuable AI company admitted begging for GPUs, signaling a structural shortage. The confession, reported by @TheGeorgePu, contradicts vendor narratives of ample supply.
LongStraw Reaches 2.1M Tokens on 8 H20 GPUs via Branch Replay
LongStraw reaches 2.1M token positions for RL post-training on 8 H20 GPUs by replaying short response branches, cutting compute 8-16x vs prior art.
Colibri Runs 744B-Parameter Model on 25GB RAM, No GPU
Colibri claims to run a 744B-parameter model on 25GB RAM without GPU, but lacks evidence. If true, it could democratize large-model inference.
Crusoe Launches Serverless Fine-Tuning, Targets AI Lifecycle Beyond GPUs
Crusoe launched serverless fine-tuning and inference, targeting enterprise AI teams. IDC says GPU access is no longer the differentiator; portability is now a procurement requirement.
95% of Announced Nvidia Blackwell GPUs Yet to Deploy
95% of announced Nvidia Blackwell GPUs remain undeployed per Air Street Capital, signaling a gap between orders and infrastructure.
DeepSeek, Zhipu AI Build Custom Inference Chips to Cut GPU Dependency
DeepSeek and Zhipu AI are developing custom inference chips to cut GPU costs. China's domestic chip budget share hit 46% in July 2026.
Biren Raises $893M to Ramp GPU Production, Challenge Nvidia in China
Biren raises $893M at a discount to fund GPU production and challenge Nvidia in China's AI chip market.
Nvidia Renting Back GPU Capacity from Neoclouds Signals Demand Softening
Nvidia renting back GPU capacity from neoclouds signals demand softening. Analyst @edzitron claims the market cannot absorb current supply.
NHN Cloud Tops Korean TOP500 with FactoryX GPU Clusters
NHN Cloud tops Korean TOP500 with FactoryX GPU clusters delivering 1.2 exaflops, marking first domestic cloud provider to lead the list.
Colossus 2: xAI's Memphis Cluster Hits 300,000 GPUs
xAI's Colossus 2 hits 300,000 GPUs, targeting 1M by year-end. Training Grok-3, the $6B cluster challenges OpenAI and Google.
AMD's Lemonade v10.8 Adds MCP Support, Letting Claude Desktop and Cursor Route Tasks to Local AMD GPUs
AMD-backed Lemonade v10.8, released June 17, now exposes a Model Context Protocol server, letting Claude Desktop, Cursor, and GitHub Copilot route inference tasks to local AMD Ryzen AI NPUs, Radeon GPUs, or plain CPUs — no cloud API required. The update also adds Moonshine speech-to-text, expanded R
TensorWave Raises $350M Series B for AMD-Powered GPU Clusters
TensorWave raised $350M Series B for AMD-powered GPU clusters in North America, challenging Nvidia's dominance.
mlx-vlm v0.6.2 Adds Gemma 4 QAT Support for Local GPUs
mlx-vlm v0.6.2 adds launch-day support for Google DeepMind's Gemma 4 QAT checkpoints, enabling local inference on consumer GPUs and edge devices with video input for the 12B model.
Cerebras Hits 981 Tokens/sec on 1T-Parameter Kimi K2.6, Claims 6.7× GPU Cloud Speedup
Cerebras reported 981 tokens/sec on the 1T-parameter Kimi K2.6 model, a 6.7× speedup over the next GPU cloud, validated by an independent third party.
Nvidia Networking Revenue Hits $14.8B, Up 199% as AI Spending Shifts Beyond GPUs
Nvidia's Q1 FY2027 networking revenue surged 199% to $14.8B, signaling AI infrastructure spending is moving beyond GPUs into full-system networking. New reporting splits into Hyperscale and ACIE segments reflect a broadening customer base beyond hyperscalers.
train-llm-from-scratch: 1B-Parameter LLM on a Single GPU
train-llm-from-scratch trains billion-parameter LLMs on a single GPU, cutting costs from $10M+ to consumer hardware.
vLLM Optimizations Cut Voice AI Latency by 40% on 6-GPU Cluster
vLLM optimizations on a 6-GPU cluster reduced voice AI latency by 40% for a Qwen-based system, enabling 500 concurrent sessions per node without hardware upgrades.
CoreWeave, Nebius Earnings Show AI Race Shifts From GPUs to Power
CoreWeave and Nebius Q1 earnings show AI infrastructure race shifting from GPU supply to power and scale, with combined capex guidance exceeding $55B.
Cerebras IPO Challenges GPU Scaling Orthodoxy
Cerebras filed for IPO on April 21, betting wafer-scale chips can disrupt Nvidia's GPU cluster model for AI workloads.
MLX CUDA Backend Passes All Tests, Closing Apple GPU Gap
MLX CUDA backend passes all tests, enabling NVIDIA GPU support. Milestone bridges Apple Silicon and CUDA ecosystems for ML workloads.
NHN Deploys 7,656-GPU AI Cluster in Seoul
NHN launched a 7,656-GPU cluster in Seoul, South Korea, for domestic enterprise AI workloads. The cluster targets inference and training, competing with Naver and Kakao.
VS Code Now Connects Directly to Google Colab With Free T4 GPU
Google Colab integrates with VS Code, offering a free T4 GPU inside the editor, bypassing cloud GPU providers.
Detecting AI Images: Metadata Exposes Generators, No GPU Needed
AI image detection via metadata analysis exposes generators like Google's Gemini and Meta's Llama without GPU clusters, highlighting a simple but effective method.