ArXiv

Real-Time Voice AI Hears but Does Not Listen

Authors
Martijn Bartelds, Federico Bianchi, James Zou
Categories
cs.CL, eess.AS
arXiv
https://arxiv.org/abs/2606.26083v1
PDF
https://arxiv.org/pdf/2606.26083v1

Brief

The paper evaluates four leading realtime voice AI systems on tasks where both lexical content and vocal delivery convey critical information. Across three consequential scenarios the systems routinely base actions on transcripts (words) rather than tone, even though three of four can explicitly label distress, fear, or sarcasm when asked. Prompting helps only inconsistently, highlighting an "emotional intelligence gap" and urging caution for deployment where vocal affect matters.

Why it matters

Martijn Bartelds, Federico Bianchi, and James Zou (arXiv:2606.26083v1, published 2026-06-24) evaluated four production realtime voice AIs—OpenAI GPT Realtime 2, Google Gemini 3.1 Flash Live, and Alibaba Qwen3.5 Omni Plus and Omni Flash—on three high-stakes scenarios (crying callers denying distress, frightened voices authorizing wire transfers, clearly sarcastic agreement); all four systems acted on words rather than vocal delivery.

Key details

  • When directly queried, three of the four systems reliably identified distress, fear, or sarcasm but then ignored those cues when making decisions; attempts to prompt systems to attend to vocal delivery yielded only partial and inconsistent improvements, a failure the authors call the 'emotional intelligence gap' of voice AI.
Source evidence

Abstract

Speech conveys information through both words and vocal delivery. We evaluate four leading production realtime voice systems-OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live, and Alibaba's Qwen3.5 Omni Plus and Omni Flash-on tasks where the words and the delivery patterns both convey meaningful information. Across three consequential scenarios, all four systems act on the words rather than the voice. They end calls with crying callers who insist nothing is wrong, approve wire transfers authorized in frightened voices, and enroll callers whose agreement is clearly sarcastic. Surprisingly, this is often not a failure of perception. When asked directly, three of the four systems reliably identify the distress, fear, or sarcasm they later ignore when making decisions. We observe a similar pattern when these realtime voice systems estimate accent and age, as their responses frequently follow the biases of the words rather than the acoustic properties of the speaker. We term this disconnect between perception and action the emotional intelligence gap of voice AI. Prompting systems to explicitly attend to vocal delivery improves performance only partially and inconsistently. Our findings show that current realtime voice AI systems often behave as if speech had been reduced to a transcript, suggesting that they should be used with caution in settings where the tone and emotion of delivery convey important information.