新智元 (微信公众号)T2·68 pts
- VoxMem is the first benchmark to integrate non-textual acoustic cues like speaker identity, tone, and background sounds into multi-session long-term memory evaluation.
- Experimental data shows that top-tier audio LLMs all have overall memory accuracy below 40% at 32K context length, indicating long-range acoustic memory remains a bottleneck.
- Models retain explicit textual information far better than implicit acoustic cues, with only 3.4% accuracy in tracking tone changes and 1.2% for background sound changes.