UNSW
AI overview
Sign in and the AI will write an overview from our coverage.
Headlines · 1
- VoxMem benchmark: audio LLMs score under 40% on multi-session speech memory
A team from the University of Melbourne and UNSW proposes VoxMem, a multi-session speech memory benchmark built on an "acoustic evidence × memory operation" framework covering 15 combinations, with 799 questions, 3,196 evaluation instances and 34,743 spoken sessions (about 177 hours). Each question is held constant across 8K–64K token history lengths. Evaluating 15 audio LLMs, the study reports no model topping 40% overall accuracy at 32K (best: Qwen3.8-Omni-Flash at 38.5%), while average accuracy for tracking tone changes and background-sound changes was just 3.4% and 1.2%. Error analysis attributes most failures to locating and retaining acoustic evidence rather than to reasoning.
新智元 · 🔥 7
Experience and discussion from the community
Share my UNSW experienceAsk about UNSW
Nobody has shared their experience with UNSW yet.