cuda
30 articles about cuda in AI news
Alibaba Open-Sources SAIL Stack to Break Nvidia CUDA Lock-In
Alibaba T-Head open-sourced SAIL stack for Zhenwu chips at WAIC, targeting Nvidia CUDA dominance with 7-day migration claim.
Reverse-engineering Nvidia's cuda-checkpoint reveals 70x cold-start speedup path
Reverse-engineering Nvidia's cuda-checkpoint reveals PCIe bandwidth underutilization. The tool enables up to 70x faster cold starts for GPU servers, critical for AI inference scaling.
NanoEuler: GPT-2-Scale 116M Model Built in Pure C/CUDA From Scratch
NanoEuler is a 116M-parameter GPT-2-scale model built in pure C/CUDA from scratch. It provides a complete educational training pipeline for understanding LLMs at the lowest level.
MLX CUDA Backend Passes All Tests, Closing Apple GPU Gap
MLX CUDA backend passes all tests, enabling NVIDIA GPU support. Milestone bridges Apple Silicon and CUDA ecosystems for ML workloads.
OpenAI Codex Now Translates C++, CUDA, and Python to Swift and Python for CoreML Model Conversion
OpenAI's Codex AI code generator is now being used to automatically rewrite C++, CUDA, and Python code into Swift and Python specifically for CoreML model conversion, a previously manual and error-prone process for Apple ecosystem deployment.
ByteDance's CUDA Agent: The AI System Outperforming Human Experts in GPU Code Generation
ByteDance has unveiled CUDA Agent, a large-scale reinforcement learning system that generates high-performance CUDA kernels. The system achieves state-of-the-art results, outperforming torch.compile by up to 100% and beating leading AI models like Claude Opus 4.5 and Gemini 3 Pro by approximately 40% on the most challenging tasks.
WSL 3 Preview: Cut Claude Code's Local Inference Latency on Windows
WSL 3 preview delivers near-native GPU/NPU for Claude Code + Ollama on Copilot+ laptops, but WSL 2 still handles NVIDIA CUDA fine for desktop users.
LlamaFactory Enables No-Code Fine-Tuning for 100+ LLMs Including Llama 4, Qwen, and DeepSeek
The LlamaFactory project eliminates traditional fine-tuning complexity with a drag-and-click interface, supporting over 100 models. This reduces setup from hours of boilerplate code and CUDA debugging to a visual workflow.
Nvidia's Open-Source Gambit: NeMoClaw Aims to Tame Enterprise AI Agents
Nvidia is preparing to launch NeMoClaw, an open-source platform designed for building secure, autonomous AI agents for enterprise workflows. Breaking from its proprietary CUDA tradition, the move targets software ecosystem dominance regardless of hardware.
Google's TPUv8i Starts Software Bring-Up on g3 Codebase
SemiAnalysis reports Google's TPUv8i entered software bring-up on g3 and public stacks, signaling accelerated TPU software externalization. No specs disclosed.
SemiAnalysis: Can TileRT Software Match Cerebras on NVIDIA GPUs?
SemiAnalysis is testing TileRT InferenceX, software claiming batch-1 ultra-high interactivity on NVIDIA GPUs, targeting Cerebras, Groq LPU, and SambaNova. No benchmarks disclosed yet.
Google Open-Sources TPU Raiden, Its NIXL Equivalent for KV Cache
Google open-sourced TPU Raiden, its KV cache transfer library equivalent to NVIDIA NIXL, signaling deeper externalization of its TPU stack.
NVIDIA Releases Nemotron VoiceChat, First Open Full-Duplex Speech Model
NVIDIA released Nemotron VoiceChat, claiming the first open full-duplex speech model with tool calling and barge-in. The move targets real-time voice agents, challenging proprietary APIs.
SambaNova SN50 MVP Runs MiniMax M2.7, But Batch Size Limit Looms
SambaNova's SN50 MVP runs MiniMax M2.7 but is stuck at batch size 2, highlighting software maturity issues for frontier models.
M4 Max Mac Studio Tops GB10 in Local AI Decode Throughput
M4 Max Mac Studio beats GB10 and Strix Halo in local AI decode throughput but memory bandwidth caps large model performance. Tom's Hardware tested llama.cpp across three platforms.
Nvidia Ships Hundreds of Thousands of Grace Standalone Servers
Nvidia shipped hundreds of thousands of Grace standalone servers. The CPU pivot targets agentic AI workloads shifting hardware balance.
Zhipu AI Builds 1GW China-Only Data Center, Acquires Compiler Startup
Zhipu AI builds 1GW all-domestic chip data center, acquires compiler startup, explores custom AI chip development to decouple from Nvidia.
Microsoft to Deploy AMD Helios Rack-Scale AI at Scale on Azure
Microsoft will deploy AMD's Helios rack-scale AI accelerator at scale on Azure, powered by MI455X GPUs and Epyc Venice CPUs. The move diversifies Azure's AI silicon beyond Nvidia.
How This Solo Builder Ships Features While Sleeping with a 5-Machine Local
Alex Finn's build-and-review loop with Claude Code and local models like OpenClaw automates feature shipping on 5 machines. Key takeaway: set up Tailscale and allocate tasks by model strength.
Nvidia, Hugging Face Open-Source Robot Models to Democratize Physical AI
Nvidia and Hugging Face open-sourced robot models to democratize physical AI, providing pre-trained models and simulation tools on the Hugging Face hub.
PKU Chip Hits 2.12ms Brain Latency, 478x A100 Speedup
PKU chip achieves 2.12ms step latency with 478x speedup over Nvidia A100 for brain modeling using phase-change memristors.
Biren Raises $893M to Ramp GPU Production, Challenge Nvidia in China
Biren raises $893M at a discount to fund GPU production and challenge Nvidia in China's AI chip market.
Jim Keller: Tenstorrent IPO Looms as BlackHole Chip Scales
Jim Keller confirmed Tenstorrent's IPO plans as BlackHole chip scales for AI inference, competing with Nvidia. No revenue disclosed.
OpenAI-Broadcom Chip Hints at Token Price Collapse
OpenAI and Broadcom are co-developing a custom AI inference chip that could cut token prices by an order of magnitude, per @mweinbach. The chip targets inference workloads, not training, and aims to reduce dependency on Nvidia.
NVIDIA Vera Rubin: One Rack Matches TOP500, 35 EU Labs Deploy
NVIDIA's Vera Rubin NVL72 delivers TOP500-class performance in a single rack, with 35 European labs deploying the system for AI and HPC.
How Simon Willison Ported a 0.2B Image Model to the Browser with Claude
Simon Willison used Claude Code to port a 0.2B image inpainting model to WebGPU, running it as a parallel side project while his main agent worked on Datasette. The technique? Research with Claude.ai, then hand off to Claude Code with research.md.
Qualcomm in Talks to Acquire Modular for $4B, Landing Lattner
Qualcomm nears $4B acquisition of Modular, Chris Lattner's AI infra startup. Deal targets inference software for edge and data center AI chips.
NVFP4 GEMM on RTX Pro Blackwell: SM12x Breaks from B200 Programming Model
NVIDIA's SM12x architecture drops tcgen05.mma for mma.sync, breaking B200 kernel compatibility. SM8x kernels port easily; developers must maintain separate codebases.
Intel Targets Nvidia, AMD with New AI Chip Launch by End 2026
Intel plans to launch a new AI data center chip by end of 2026, targeting Nvidia and AMD in the AI infrastructure market.
AWS Beats Cloud Rivals to NVIDIA Blackwell with EC2 G7 — 4.6x AI Inference Gain Over G6
AWS launched EC2 G7 instances on June 19, 2026, becoming the first major cloud to offer NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs. The instances claim 4.6x AI inference performance over G6, backed by 700 Gbps EFA networking and 32 GB GDDR7 per GPU. The move arrives the same week AWS confirme