Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

quantization

30 articles about quantization in AI news

KV Cache Quantization Silently Breaks Safety Alignment, Paper Shows

KV cache quantization silently breaks LLM safety alignment, with Mistral-7B losing 15.2% refusals at 1.03x perplexity. PCR diagnostic recovers up to 97% alignment in 35 GPU-minutes.

79% relevant

Product Quantization: The Hidden Engine Behind Scalable Vector Search

The article explains Product Quantization (PQ), a method for compressing high-dimensional vectors to enable fast and memory-efficient similarity search. This is a foundational technology for scalable AI applications like semantic search and recommendation engines.

88% relevant

TTQ: A New Framework for On-the-Fly Quantization of LLMs at Inference Time

Researchers propose TTQ, a test-time quantization method that compresses large language models dynamically during inference. It uses efficient online calibration to adapt to any prompt, aiming to solve domain-shift issues and accelerate inference without retraining.

70% relevant

Efficient Fine-Tuning of Vision-Language Models with LoRA & Quantization

A technical guide details methods for fine-tuning large VLMs like GPT-4V and LLaVA using Low-Rank Adaptation (LoRA) and quantization. This reduces computational cost and memory footprint, making custom VLM training more accessible.

80% relevant

The Quantization Paradox: How Compressing Multimodal AI Impacts Reliability

New research reveals that compressing multimodal AI models through quantization significantly reduces their reliability, making them more likely to produce confidently wrong answers. The study identifies methods to mitigate these effects while maintaining efficiency gains.

70% relevant

Why Your GPU's Memory Ceiling Predicts Your Cloud Inference Costs

Towards AI explains how GPU VRAM constraints, not compute, dictate LLM inference costs locally and in the cloud. The article details memory math, 2026 hardware shortages, and pricing trends, urging teams to manage context and quantization.

78% relevant

Massive Activations Found in Hybrid Linear Attention LLMs

New paper finds massive activations in hybrid linear attention LLMs, forming pre-attention spikes and inter-spike plateaus. Code and checkpoints released on Hugging Face; implications for quantization and deployment.

85% relevant

Paper Details Full-Stack MFM Acceleration: Quant, Spec Decode, HW Co-Design

A research paper details a full-stack approach for accelerating multimodal foundation models, combining hierarchy-aware mixed-precision quantization, structural pruning, speculative decoding, model cascading, and a specialized hardware accelerator. Demonstrated on medical and code generation tasks.

72% relevant

Fine-Tuning an LLM on a 4GB GPU: A Practical Guide for Resource-Constrained Engineers

A Medium article provides a practical, constraint-driven guide for fine-tuning LLMs on a 4GB GPU, covering model selection, quantization, and parameter-efficient methods. This makes bespoke AI model development more accessible without high-end cloud infrastructure.

100% relevant

TurboQuant Ported to Apple MLX, Claims 75% Memory Reduction with Minimal Performance Loss

Developer Prince Canuma has successfully ported the TurboQuant quantization method to Apple's MLX framework, reporting a 75% reduction in memory usage with nearly no performance degradation for on-device AI models.

85% relevant

Google's TurboQuant Cuts LLM KV Cache Memory by 6x, Enables 3-Bit Storage Without Accuracy Loss

Google released TurboQuant, a novel two-stage quantization algorithm that compresses the KV cache in long-context LLMs. It reduces memory by 6x, achieves 3-bit storage with no accuracy drop, and speeds up attention scoring by up to 8x on H100 GPUs.

95% relevant

Flash-KMeans Achieves 200x Speedup Over FAISS by Targeting GPU Memory Bottlenecks

Flash-KMeans is an IO-aware GPU implementation of exact k-means that runs 30x faster than cuML and 200x faster than FAISS. At million-scale datasets, it completes iterations in milliseconds, enabling dynamic re-indexing and real-time quantization.

95% relevant

Quantized Inference Breakthrough for Next-Gen Recommender Systems: OneRec-V2 Achieves 49% Latency Reduction with FP8

New research shows FP8 quantization can dramatically speed up modern generative recommender systems like OneRec-V2, achieving 49% lower latency and 92% higher throughput with no quality loss. This breakthrough bridges the gap between LLM optimization techniques and industrial recommendation workloads.

97% relevant

LeCun's Team Uncovers Hidden Transformer Flaws: How Architectural Artifacts Sabotage AI Efficiency

NYU researchers led by Yann LeCun reveal that Transformer language models contain systematic artifacts—massive activations and attention sinks—that degrade efficiency. These phenomena, stemming from architectural choices rather than fundamental properties, directly impact quantization, pruning, and memory management.

95% relevant

LittleBit-2: How Geometric Alignment Unlocks Ultra-Efficient AI Below 1-Bit

Researchers have developed LittleBit-2, a framework that achieves state-of-the-art performance in sub-1-bit LLM compression by solving latent geometry misalignment. The method uses internal latent rotation and joint iterative quantization to align model parameters with binary representations without inference overhead.

75% relevant

AutoQRA: The Breakthrough That Makes AI Fine-Tuning 4x More Efficient

Researchers have developed AutoQRA, a novel framework that jointly optimizes quantization precision and LoRA adapters for large language models. This breakthrough enables near-full-precision performance with dramatically reduced memory requirements, potentially revolutionizing how organizations fine-tune AI models on limited hardware.

75% relevant

Gigabyte Ships GB300 Desktop Superchip PC for 400 Concurrent Users

Gigabyte unveiled a desktop PC with Nvidia's GB300 Grace Blackwell Ultra Superchip, claiming 400 concurrent AI users. The data-center-in-a-box targets on-prem enterprise inference.

85% relevant

How Generative Recommenders Are Redefining RecSys at Scale | NVIDIA

NVIDIA's technical blog details how generative recommenders are redefining RecSys at scale. It highlights a shift from two-stage pipelines to sequence-to-sequence transformers for large-scale platforms.

99% relevant

Ornith-1.5 Open-Source LLM Family: 9B Dense, 35B MoE

Ornith-1.5 open-source LLM family announced with 9B Dense, 35B MoE, and 39B variants. No benchmarks or technical details disclosed, limiting immediate evaluation.

87% relevant

FreeToken Runs 284B MoE Locally on a Gaming Desktop

FreeToken claims 284B MoE serving on a gaming desktop via unified PC inference. No benchmarks disclosed; feasibility depends on MoE sparsity and offload strategy.

87% relevant

Tencent's EVIE Hits ViDoRe SOTA with 128D Embeddings

Tencent's EVIE model achieves SOTA on ViDoRe with 128D embeddings, a major efficiency gain. The release lacks technical details, but the compact vector size signals a shift toward cost-efficient retrieval.

85% relevant

Qwen 3.7 27B Draws Best Local-Model Pelican, Simon Willison Says

Qwen 3.7 27B, as a 17GB GGUF, drew Simon Willison's best local pelican-bicycle image. Signals on-device generation crossed a quality threshold.

95% relevant

OpenAI's Ultrafast Mode Hits 750 Tokens/s on GPT-5.6 Sol

OpenAI launched Ultrafast mode for GPT-5.6 Sol at 14x speed and 750 tokens/s, powered by Cerebras. Preview limited to select customers, targeting latency-sensitive enterprise workflows.

100% relevant

Google's 7-Year-Old TPUs Still Run at 100% Utilization, Says Vahdat

Google's 7-8 year old TPUs still run at 100% utilization per Amin Vahdat, citing Jevons Paradox. Efficiency gains drive demand, keeping all chips busy.

87% relevant

Meta Drops Muse Glimmer 30B Under Apache 2.0, First OSI License

Meta released Muse Glimmer 30B under Apache 2.0, its first OSI-licensed open-weight model, marking a strategic shift toward true open-source AI.

100% relevant

Anthropic Builds In-House Chip Team for Claude Co-Design

Anthropic is building an in-house chip team to co-design hardware with Claude, aiming for faster, cheaper inference. The company maintains a multi-chip strategy with AWS, Google, Nvidia, and AMD.

94% relevant

1-Bit Kimi K3 Quant Cuts 2.8T Model to 590GB, Keeps 78.7% Quality

Atomic Labs' 1-bit quant shrinks 2.8T Kimi K3 to 590GB (-62%), keeping 78.7% quality and 1M context. Runs on 4x B200s.

100% relevant

M4 Max Mac Studio Tops GB10 in Local AI Decode Throughput

M4 Max Mac Studio beats GB10 and Strix Halo in local AI decode throughput but memory bandwidth caps large model performance. Tom's Hardware tested llama.cpp across three platforms.

85% relevant

Moonshot AI Releases 1.56T-Parameter Kimi K3, Requires 2x B200 Nodes

Moonshot AI released Kimi K3, a 1.56T parameter MoE model at 1561 GB, requiring 2x B200 nodes. No benchmarks disclosed.

100% relevant

Alibaba Qwen3.8: 2.4T Parameter Open-Weight Model Incoming

Alibaba's Qwen3.8, a 2.4T parameter open-weight model, was announced. It would be the largest open-weight model ever, but lacks benchmark details.

100% relevant