Today for AI
HOT RADAR
llmHEAT 9.3°

AI CLUSTERED EVENT · 10/9/2026

VoxMem Benchmark Reveals Audio LLMs Struggle to Retain Speaker Identity and Tone Across Multi-Session Dialogues

1 reports archived1 independent sourcesupdated 10/9/2026, 11:55:00
Synthesis & Latest Updates
1 Sources Cross-Validated

The University of Melbourne and UNSW jointly introduced VoxMem, a multi-session speech memory benchmark designed to address the neglect of acoustic features (speaker identity, tone) and single-segment limitations in existing evaluations. Covering approximately 177 hours of audio across 15 combinations via an 'acoustic evidence × memory operation' framework, tests on 15 leading audio LLMs revealed that none exceeded 40% overall accuracy at 32K token context lengths. Notably, average accuracy for tracking tone and background sound changes was critically low at 3.4% and 1.2%, highlighting severe deficiencies in retaining non-textual acoustic information over long contexts.

LATEST/The University of Melbourne and UNSW jointly introduced VoxMem, a multi-session speech memory benchmark designed to address the neglect of acoustic features (speaker identity, tone) and single-segment limitations in existing evaluations. Covering approximately 177 hours of audio across 15 combinations via an 'acoustic evidence × memory operation' framework, tests on 15 leading audio LLMs revealed that none exceeded 40% overall accuracy at 32K token context lengths. Notably, average accuracy for tracking tone and background sound changes was critically low at 3.4% and 1.2%, highlighting severe deficiencies in retaining non-textual acoustic information over long contexts.

HEAT TRENDHourly heat curve

2 fully observed hours
Fewer than 3 fully observed hours; trend not plotted yet.

TIMELINECoverage timeline

Total 1 reports · Latest first
  1. 新智元 (微信公众号)T2·68 pts
    • VoxMem is the first benchmark to integrate non-textual acoustic cues like speaker identity, tone, and background sounds into multi-session long-term memory evaluation.
    • Experimental data shows that top-tier audio LLMs all have overall memory accuracy below 40% at 32K context length, indicating long-range acoustic memory remains a bottleneck.
    • Models retain explicit textual information far better than implicit acoustic cues, with only 3.4% accuracy in tracking tone changes and 1.2% for background sound changes.
VoxMem Benchmark Reveals Audio LLMs Struggle to Retain Speaker Identity and Tone Across Multi-Session Dialogues | AI Clustered Intelligence | Today for AI