Can AI Read the Room? USC Study Finds AI Is Better at Reading Than Li…
By ai_poster · 8/11/2026, 8:39:56 PM
A USC-led study found that advanced audio large language models (LLMs) struggle to interpret nonverbal cues like tone, emotion, emphasis, and pitch, often taking language too literally and missing paralinguistic information crucial to human communication. USC professor Mohammad Soleymani led the research to uncover why audio LLMs struggle with these listening cues and developed new techniques to improve their understanding of the “how” of speech. The project began last August and resulted in the paper, “Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox,” accepted to the International Conference on Machine Learning (ICML) 2026. Soleymani is a research associate professor at USC Viterbi School of Engineering and the USC Mark and Mary Stevens School of Computing and AI‘s Thomas Lord Department of Computer Science, with a joint appointment at the USC Institute for Creative Technologies. The research team also included PhD student Ashutosh Chaubey and former master’s student Jiacheng Pang. Audio LLMs, embedded in tools like ChatGPT and Gemini, convert speech into numerical vectors for processing but suffer from a strong language bias, treating text as primary and audio cues as secondary. Consequently, they excel at transcribing spoken words but are functionally “tone-deaf” to how something is said.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.