Create

Sign in to ReadmeX

or

How we handle your data: Privacy Policy

ReadmeX
ReadmeX

A clearer picture, in a conversation.

Catch up on what matters, then ask a little deeper.

Your community and people briefings stay personal to you.

VoxMem

AI overview

Sign in and the AI will write an overview from our coverage.

Headlines · 2

  1. VoxMem benchmark: audio LLMs all fall below 40% accuracy at 32K context

    A joint team from the University of Melbourne and UNSW released VoxMem, an audio-LLM memory benchmark built on an "acoustic evidence × memory operation" taxonomy covering 15 combinations, 3,196 evaluation instances, 34,743 speech sessions and roughly 177 hours of audio. Questions are held constant across 8K–64K history lengths so only conversation length varies. Across 15 mainstream audio models, overall accuracy stayed below 40% at 32K context; memory for speaker identity, paralinguistic cues such as tone and emotion, and environmental sound lagged clearly behind semantic recall of what was said, and all categories degraded as history grew.

    AIbase AI新闻 · 🔥 8
  2. VoxMem benchmark: audio LLMs score under 40% on multi-session speech memory

    A team from the University of Melbourne and UNSW proposes VoxMem, a multi-session speech memory benchmark built on an "acoustic evidence × memory operation" framework covering 15 combinations, with 799 questions, 3,196 evaluation instances and 34,743 spoken sessions (about 177 hours). Each question is held constant across 8K–64K token history lengths. Evaluating 15 audio LLMs, the study reports no model topping 40% overall accuracy at 32K (best: Qwen3.8-Omni-Flash at 38.5%), while average accuracy for tracking tone changes and background-sound changes was just 3.4% and 1.2%. Error analysis attributes most failures to locating and retaining acoustic evidence rather than to reasoning.

    新智元 · 🔥 7

Experience and discussion from the community

Share my VoxMem experienceAsk about VoxMem

Nobody has shared their experience with VoxMem yet.