
dMoE Cuts Active Experts from 69.5 to 14.6, Retains 99.11% Performance
dMoE reduces active experts from 69.5 to 14.6 in diffusion LLMs, retaining 99.11% performance while cutting memory 80% and speeding inference 1.66×.
Every story, newest first — research, funding, product launches, policy, and analysis.