
DEAF Benchmark Reveals Audio MLLMs Rely on Text, Not Sound, Scoring Below 50% on Acoustic Faithfulness
Researchers introduce DEAF, a 2,700-stimulus benchmark testing Audio MLLMs' acoustic processing. Evaluation of seven models shows a consistent pattern of text dominance, with models scoring below 50% on acoustic faithfulness metrics.













![Speech synthesis interface showing waveform visualization and inline control tags like [happy] and [loud] overlaid…](/_next/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHDnxss7bEAUbvSR.jpg&w=1920&q=75)






