latency
30 articles about latency in AI news
AWS Personalize cuts recommendation latency to under 10 seconds in new
AWS published a real-time e-commerce personalization blueprint that cuts recommendation latency to under 10 seconds using AWS Personalize, Kinesis, and Lambda. The architecture bridges batch and streaming modes for luxury retailers.
MCP Cuts Token Costs 75% But Adds 30x Latency vs REST APIs
MCP cuts token costs by 75% but adds 30x latency versus REST. The protocol, backed by Anthropic and OpenAI, trades speed for dynamic tool discovery.
PKU Chip Hits 2.12ms Brain Latency, 478x A100 Speedup
PKU chip achieves 2.12ms step latency with 478x speedup over Nvidia A100 for brain modeling using phase-change memristors.
Wan-Streamer v0.1 Cuts Audio-Visual Interaction Latency to 200ms in Single
Wan-Streamer v0.1 achieves 200ms model-side latency in a single Transformer for full-duplex audio-visual interaction, eliminating cascaded modules. The paper lacks parameter count and benchmark comparisons, limiting reproducibility.
llada.cpp Cuts LLaDA-8B Latency 17-42x on Mobile NPU
llada.cpp, the first NPU-aware dLLM inference framework, cuts LLaDA-8B latency 17-42x on smartphones, enabling real-time on-device generation.
Miso One: 8B Open-Source TTS Hits 110ms Latency, Real Emotion
Miso One, an 8B open-source TTS model, achieves 110ms latency with emotional range. Weights are fully open-source for self-hosting, but no benchmark data is provided.
vLLM Optimizations Cut Voice AI Latency by 40% on 6-GPU Cluster
vLLM optimizations on a 6-GPU cluster reduced voice AI latency by 40% for a Qwen-based system, enabling 500 concurrent sessions per node without hardware upgrades.
Codex Update Cuts GUI Workflow Latency 42%
Codex app update cuts GUI workflow latency 42%, enabling near-human-speed interface operation for autonomous app building and debugging.
Google Virgo Fabric: 100K-Accelerator AI Network Cuts Latency
Google unveiled Virgo, a data center fabric for AI clusters of 100,000+ accelerators, using a flatter two-layer topology to reduce latency and improve bisection bandwidth for synchronized training workloads.
OpenCLAW-P2P v6.0 Cuts Paper Lookup Latency to <50ms
OpenCLAW-P2P v6.0 introduces a multi-layer persistence architecture and live reference verification, reducing paper retrieval latency from >3s to <50ms and operating with 14 autonomous agents that scored 50+ papers.
EventChat Study: LLM-Driven Conversational Recommenders Show Promise but Face Cost & Latency Hurdles for SMEs
A new study details the real-world implementation and user evaluation of an LLM-driven conversational recommender system (CRS) for an SME. Results show 85.5% recommendation accuracy but highlight critical business viability challenges: a median cost of $0.04 per interaction and 5.7s latency.
OmniForcing Enables Real-Time Joint Audio-Visual Generation at 25 FPS with 0.7s Latency
Researchers introduced OmniForcing, a method that distills a bidirectional LTX-2 model into a causal streaming generator for joint audio-visual synthesis. It achieves ~25 FPS with 0.7s latency, a 35× speedup over offline diffusion models while maintaining multi-modal fidelity.
Quantized Inference Breakthrough for Next-Gen Recommender Systems: OneRec-V2 Achieves 49% Latency Reduction with FP8
New research shows FP8 quantization can dramatically speed up modern generative recommender systems like OneRec-V2, achieving 49% lower latency and 92% higher throughput with no quality loss. This breakthrough bridges the gap between LLM optimization techniques and industrial recommendation workloads.
WSL 3 Preview: Cut Claude Code's Local Inference Latency on Windows
WSL 3 preview delivers near-native GPU/NPU for Claude Code + Ollama on Copilot+ laptops, but WSL 2 still handles NVIDIA CUDA fine for desktop users.
Shopify's Gisting: Compressing LLM Agent Context to Boost Throughput and
Shopify Engineering unveiled Gisting, a context compression method for LLM agents that boosts throughput and cuts costs. The technique addresses the rising token expenses and latency in long-running agentic workflows.
Kimi K3 in Claude Code: Slow Start, Then It Just Works — What That Tells You
Kimi K3 runs in Claude Code via OpenRouter with a slow first impression but becomes usable. Its Code Arena #1 frontend ranking makes it a solid choice for UI tasks, but Claude Opus 4.8 remains better for latency-sensitive work.
Orbital AI Data Centers: Compute's Next Frontier?
DCD floats orbital AI data centers as next compute frontier. Physics and economics hurdles—power, cooling, latency, launch costs—remain unsolved. No concrete plans announced.
OpenAI's Ultrafast Mode Hits 750 Tokens/s on GPT-5.6 Sol
OpenAI launched Ultrafast mode for GPT-5.6 Sol at 14x speed and 750 tokens/s, powered by Cerebras. Preview limited to select customers, targeting latency-sensitive enterprise workflows.
SemiAnalysis: Linux Runs Claude Shell Commands 3x Faster Than Windows
SemiAnalysis analyzed 2.91M Claude tool calls, finding Linux executes shell commands 3x faster than Windows, reshaping agentic AI latency priorities.
Anthropic Engineer Shows Claude Managed Agents for Server-Side AI
Anthropic engineer demoed Claude Managed Agents, a server-side harness with 90% lower P95 latency and an SRE agent that traced a P99 spike to a commit.
Claude Haiku vs Gemini Flash vs GPT-5.4 Mini: The 5x Speed Gap Explained
Claude Code's /model haiku command routes simple subtasks to Claude Haiku, which benchmarks 5x faster than Gemini Flash and GPT-5.4 Mini, cutting session latency by up to 40%.
OpenAI Details GPT-Live's Continuous Audio Architecture
OpenAI disclosed GPT-Live's continuous audio architecture via @rohanpaul_ai. The design shifts from turn-based voice processing to streaming, targeting latency, though specifics remain undisclosed.
DeepSeek DSpark: Speculative Decoding Unifies Parallel Gen, Adaptive Verification
DeepSeek released DSpark, a speculative decoding framework unifying parallel generation with adaptive verification. No benchmarks disclosed yet; the approach targets inference latency and throughput.
PrismML Shrinks Qwen 3.6 to iPhone 17 Pro, Apple Eyes Deal
PrismML compressed Alibaba's 36B-parameter Qwen 3.6 to run on an iPhone 17 Pro, drawing Apple's interest for on-device AI without cloud latency.
GitHub's Former CEO Launches Distributed Git Network for AI Coding Agents
Claude Code users should monitor Nat Friedman's distributed Git network for faster agentic coding workflows. The new network optimizes Git for AI agents, potentially reducing clone/push latency.
Epoch AI's CursorBench Benchmarks AI Code Editing at Scale
Epoch AI launched CursorBench, a 500-task benchmark for AI code editors. It reveals a 15% accuracy gap vs. humans and 3x latency variance.
VIAVI Ships First Ultra Ethernet Validation Tool for AI Data Centers
VIAVI launched the first Ultra Ethernet validation tool for AI data centers, supporting 800GE/1.6TE links. The tool enables certification of low-latency, lossless transport critical for distributed AI training.
OpenAI 'Bidi' Voice Mode Demo Leaks: Real-Time Interruption
Leaked demo shows OpenAI 'bidi' voice mode handling interruptions with sub-320ms latency. No official release date or pricing announced.
Gemini 3.5 Live Translate Debuts as Real-Time Audio Model
Google DeepMind released Gemini 3.5 Live Translate, an audio model for real-time translation, but disclosed no pricing, latency, or language pair details.
Google Releases Magenta RealTime 2 for Open-Weight Music Generation
Google released Magenta RealTime 2 on Hugging Face, the only open-weights model for real-time continuous music generation on device with ~200ms latency.