Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

cuda

30 articles about cuda in AI news

Alibaba Open-Sources SAIL Stack to Break Nvidia CUDA Lock-In

Alibaba T-Head open-sourced SAIL stack for Zhenwu chips at WAIC, targeting Nvidia CUDA dominance with 7-day migration claim.

100% relevant

Reverse-engineering Nvidia's cuda-checkpoint reveals 70x cold-start speedup path

Reverse-engineering Nvidia's cuda-checkpoint reveals PCIe bandwidth underutilization. The tool enables up to 70x faster cold starts for GPU servers, critical for AI inference scaling.

86% relevant

NanoEuler: GPT-2-Scale 116M Model Built in Pure C/CUDA From Scratch

NanoEuler is a 116M-parameter GPT-2-scale model built in pure C/CUDA from scratch. It provides a complete educational training pipeline for understanding LLMs at the lowest level.

75% relevant

MLX CUDA Backend Passes All Tests, Closing Apple GPU Gap

MLX CUDA backend passes all tests, enabling NVIDIA GPU support. Milestone bridges Apple Silicon and CUDA ecosystems for ML workloads.

77% relevant

OpenAI Codex Now Translates C++, CUDA, and Python to Swift and Python for CoreML Model Conversion

OpenAI's Codex AI code generator is now being used to automatically rewrite C++, CUDA, and Python code into Swift and Python specifically for CoreML model conversion, a previously manual and error-prone process for Apple ecosystem deployment.

89% relevant

ByteDance's CUDA Agent: The AI System Outperforming Human Experts in GPU Code Generation

ByteDance has unveiled CUDA Agent, a large-scale reinforcement learning system that generates high-performance CUDA kernels. The system achieves state-of-the-art results, outperforming torch.compile by up to 100% and beating leading AI models like Claude Opus 4.5 and Gemini 3 Pro by approximately 40% on the most challenging tasks.

95% relevant

WSL 3 Preview: Cut Claude Code's Local Inference Latency on Windows

WSL 3 preview delivers near-native GPU/NPU for Claude Code + Ollama on Copilot+ laptops, but WSL 2 still handles NVIDIA CUDA fine for desktop users.

98% relevant

LlamaFactory Enables No-Code Fine-Tuning for 100+ LLMs Including Llama 4, Qwen, and DeepSeek

The LlamaFactory project eliminates traditional fine-tuning complexity with a drag-and-click interface, supporting over 100 models. This reduces setup from hours of boilerplate code and CUDA debugging to a visual workflow.

87% relevant

Nvidia's Open-Source Gambit: NeMoClaw Aims to Tame Enterprise AI Agents

Nvidia is preparing to launch NeMoClaw, an open-source platform designed for building secure, autonomous AI agents for enterprise workflows. Breaking from its proprietary CUDA tradition, the move targets software ecosystem dominance regardless of hardware.

97% relevant

Google's TPUv8i Starts Software Bring-Up on g3 Codebase

SemiAnalysis reports Google's TPUv8i entered software bring-up on g3 and public stacks, signaling accelerated TPU software externalization. No specs disclosed.

95% relevant

SemiAnalysis: Can TileRT Software Match Cerebras on NVIDIA GPUs?

SemiAnalysis is testing TileRT InferenceX, software claiming batch-1 ultra-high interactivity on NVIDIA GPUs, targeting Cerebras, Groq LPU, and SambaNova. No benchmarks disclosed yet.

82% relevant

Google Open-Sources TPU Raiden, Its NIXL Equivalent for KV Cache

Google open-sourced TPU Raiden, its KV cache transfer library equivalent to NVIDIA NIXL, signaling deeper externalization of its TPU stack.

92% relevant

NVIDIA Releases Nemotron VoiceChat, First Open Full-Duplex Speech Model

NVIDIA released Nemotron VoiceChat, claiming the first open full-duplex speech model with tool calling and barge-in. The move targets real-time voice agents, challenging proprietary APIs.

100% relevant

SambaNova SN50 MVP Runs MiniMax M2.7, But Batch Size Limit Looms

SambaNova's SN50 MVP runs MiniMax M2.7 but is stuck at batch size 2, highlighting software maturity issues for frontier models.

72% relevant

M4 Max Mac Studio Tops GB10 in Local AI Decode Throughput

M4 Max Mac Studio beats GB10 and Strix Halo in local AI decode throughput but memory bandwidth caps large model performance. Tom's Hardware tested llama.cpp across three platforms.

85% relevant

Nvidia Ships Hundreds of Thousands of Grace Standalone Servers

Nvidia shipped hundreds of thousands of Grace standalone servers. The CPU pivot targets agentic AI workloads shifting hardware balance.

100% relevant

Zhipu AI Builds 1GW China-Only Data Center, Acquires Compiler Startup

Zhipu AI builds 1GW all-domestic chip data center, acquires compiler startup, explores custom AI chip development to decouple from Nvidia.

100% relevant

Microsoft to Deploy AMD Helios Rack-Scale AI at Scale on Azure

Microsoft will deploy AMD's Helios rack-scale AI accelerator at scale on Azure, powered by MI455X GPUs and Epyc Venice CPUs. The move diversifies Azure's AI silicon beyond Nvidia.

100% relevant

How This Solo Builder Ships Features While Sleeping with a 5-Machine Local

Alex Finn's build-and-review loop with Claude Code and local models like OpenClaw automates feature shipping on 5 machines. Key takeaway: set up Tailscale and allocate tasks by model strength.

57% relevant

Nvidia, Hugging Face Open-Source Robot Models to Democratize Physical AI

Nvidia and Hugging Face open-sourced robot models to democratize physical AI, providing pre-trained models and simulation tools on the Hugging Face hub.

98% relevant

PKU Chip Hits 2.12ms Brain Latency, 478x A100 Speedup

PKU chip achieves 2.12ms step latency with 478x speedup over Nvidia A100 for brain modeling using phase-change memristors.

100% relevant

Biren Raises $893M to Ramp GPU Production, Challenge Nvidia in China

Biren raises $893M at a discount to fund GPU production and challenge Nvidia in China's AI chip market.

100% relevant

Jim Keller: Tenstorrent IPO Looms as BlackHole Chip Scales

Jim Keller confirmed Tenstorrent's IPO plans as BlackHole chip scales for AI inference, competing with Nvidia. No revenue disclosed.

98% relevant

OpenAI-Broadcom Chip Hints at Token Price Collapse

OpenAI and Broadcom are co-developing a custom AI inference chip that could cut token prices by an order of magnitude, per @mweinbach. The chip targets inference workloads, not training, and aims to reduce dependency on Nvidia.

75% relevant

NVIDIA Vera Rubin: One Rack Matches TOP500, 35 EU Labs Deploy

NVIDIA's Vera Rubin NVL72 delivers TOP500-class performance in a single rack, with 35 European labs deploying the system for AI and HPC.

95% relevant

How Simon Willison Ported a 0.2B Image Model to the Browser with Claude

Simon Willison used Claude Code to port a 0.2B image inpainting model to WebGPU, running it as a parallel side project while his main agent worked on Datasette. The technique? Research with Claude.ai, then hand off to Claude Code with research.md.

70% relevant

Qualcomm in Talks to Acquire Modular for $4B, Landing Lattner

Qualcomm nears $4B acquisition of Modular, Chris Lattner's AI infra startup. Deal targets inference software for edge and data center AI chips.

82% relevant

NVFP4 GEMM on RTX Pro Blackwell: SM12x Breaks from B200 Programming Model

NVIDIA's SM12x architecture drops tcgen05.mma for mma.sync, breaking B200 kernel compatibility. SM8x kernels port easily; developers must maintain separate codebases.

86% relevant

Intel Targets Nvidia, AMD with New AI Chip Launch by End 2026

Intel plans to launch a new AI data center chip by end of 2026, targeting Nvidia and AMD in the AI infrastructure market.

72% relevant

AWS Beats Cloud Rivals to NVIDIA Blackwell with EC2 G7 — 4.6x AI Inference Gain Over G6

AWS launched EC2 G7 instances on June 19, 2026, becoming the first major cloud to offer NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs. The instances claim 4.6x AI inference performance over G6, backed by 700 Gbps EFA networking and 32 GB GDDR7 per GPU. The move arrives the same week AWS confirme

85% relevant