When a voice AI is involved in a several-hour-long, multi-round real conversation, can it actually remember “who said this sentence, what tone was used, and what sounds are present in the background”? A joint team from the University of Melbourne and the University of New South Wales believes that current evaluations simply cannot answer this question, so they introduced a benchmark called VoxMem specifically designed to address this weakness.
They first exposed two major flaws in the old evaluation: one is that it only focuses on the text content produced by transcription, ignoring the hidden identity information, emotional cues, and ambient sounds in the voice; the other is that it cannot isolate the “how long the history is” variable, so the model performs poorly. It’s unclear whether the problem lies in poor memory or being overwhelmed by the length. VoxMem’s approach is to establish a two-dimensional classification method of “acoustic evidence × memory operation”, breaking down the examination dimensions orthogonally. It covers 15 combinations, and all questions remain unchanged at different historical lengths from 8K to 64K—essentially just turning the “dialogue length” knob while locking everything else. This allows the changes in memory performance to be cleanly attributed to the length.
The volume of this benchmark is substantial: a total of 3196 testing instances and 34743 speech conversations, amounting to approximately 177 hours of audio, enough to simulate a long multi-round conversation. However, the actual results for the 15 mainstream audio large models are quite disappointing. At a context length of 32K, the overall accuracy of all models failed to exceed 40%. What is even more intriguing is the selective weakness in different areas: the model’s semantic memory of “what was said” is clearly stronger than its memory of the speaker’s identity, paralinguistic cues (tone, emotion), and environmental sounds—the three types of natural sound information. In particular, the accuracy in tracking tone fluctuations and background sound changes is extremely low, and it is basically a futile effort. Moreover, as the length of the history increases, the accuracy of remembering all kinds of information gradually decreases.
The team made a judgment based on this: the voice memory bottleneck of current audio large models mainly lies in the stages of “identification, positioning, and binding information together,” rather than the reasoning process itself. In other words, the model isn’t unable to think, but it simply fails to grasp the underlying facts such as “who said what, at what time, and in what way.” Therefore, higher-level reasoning cannot take place. The VoxMem benchmark is like replacing a thicker thermometer when measuring the temperature of a speech AI, forcing the industry to acknowledge a fact that has been obscured by text transcription for a long time: being able to understand words does not mean remembering people, tone, and the whole world.