Benchmark · Spoken Conversational Memory

VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

Do audio language models remember not only what was said, but who said it, how it was said, and what was audible?

Yang Xiao1 Vidhyasaharan Sethu2 Eun-Jung Holden1 Ting Dang1
1University of Melbourne 2University of New South Wales
AIMS Lab
3,196evaluation instances
799questions × 4 lengths
34,743spoken sessions
177 hof audio
8K–64Kcontext budgets
15LALMs evaluated
01

Overview

Every question is posed over a multi-session spoken history. The answer may depend on what a user said, on which speaker said it, on how it was said, or on what could be heard in the background.

TL;DR

Across 15 LALMs, no model exceeds 40% overall accuracy at 32K. Models retain what was said far better than who said it, how, or what was audible, and the gap grows with history length.

Four representative VoxMem items: (a) Information Extraction × Speaker, (b) Multi-Session Reasoning × Environmental, (c) Temporal Evolution Tracking × Paralinguistic, (d) Answer Refusal × Semantic, each with a timeline of evidence, haystack, and filler sessions and a gold answer.
Representative VoxMem examples. (a) IE × Speaker: recall a fact from the querying speaker; (b) MSR × Environmental: link sessions by a shared background sound; (c) TET × Paralinguistic: track how a speaker's vocal state changes; (d) AR × Semantic: abstain when evidence is missing.

Abstract

Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single continuous recording. Existing benchmarks fall short on all three dimensions: they focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem. We argue that principled memory evaluation requires jointly characterizing the acoustic evidence to be retained and the operations applied to it, and introduce a taxonomy along these two axes. Building on this taxonomy, we present VoxMem: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal tracking, and answer refusal), grounded in multi-session histories and stratified across context budgets from 8K to 64K tokens. Evaluating 15 LALMs, no model exceeds 40% at 32K. Models retain what was said far better than who said it, how, or what was audible, a gap that widens for complex operations, grows with history length, and manifests as qualitatively distinct failure modes across evidence types. VoxMem aims to provide a foundation to measure and drive progress on the full scope of spoken conversational memory.

02

Two axes of spoken memory

Acoustic evidence specifies what must be recovered from the history; memory operation specifies how it must be used. Their valid combinations form 15 cells. Select a cell to see details.

Acoustic evidence: what the answer depends on

Speech Semantics
what the user said; the transcript-sufficient reference
Speaker Identity
who was speaking
Paralinguistic Cues
how it was said: vocal delivery, such as emotion
Environmental Sound
what could be heard around the user

Memory operation: what the question asks for

IE · Information Extraction
recover one fact from one session
MSR · Multi-Session Reasoning
combine evidence across sessions
TET · Temporal Evolution Tracking
track how something changed over time
AR · Answer Refusal
recognise that the history does not contain the answer
03

Leaderboard

Accuracy (%) over all 799 questions at each reference history budget, broken down by answer-critical evidence type. Answer Refusal items are included, with appropriate abstention counted as correct. Click a column header to sort.

# Model Type

04

Key findings

Each question and its required evidence are held fixed across context budgets, so changes in accuracy across lengths reflect the history, not the question.

1

Current LALMs remain far from reliable spoken conversational memory.

At the 32K budget no evaluated model exceeds 40% overall accuracy; the strongest reaches 38.5%. The five proprietary models average 33.0% and the ten open-weight models 21.9%.

2

Non-lexical acoustic information is substantially less accessible than speech semantics.

At 32K, proprietary models average 55.6% on Speech Semantics, but 32.7% on Speaker Identity, 20.0% on Paralinguistic Cues and 21.9% on Environmental Sound. Open-weight models show the same separation (31.6%, 26.7%, 14.5% and 15.5%).

3

Memory difficulty depends on the interaction between the operation and the evidence type.

Multi-session reasoning is relatively robust (41.8% on semantics, 43.1% on speaker identity). Temporal tracking works for semantics (44.5%) but nearly collapses for paralinguistic (3.4%) and environmental (1.2%) evidence: maintaining states carried by vocal delivery or ambient sound is the bottleneck, not temporal reasoning alone. Explore the cells in the matrix above.

4

Answer refusal follows a different profile from answering.

Models abstain more successfully on paralinguistic and environmental questions (34.3% and 34.2%) than on semantics and speaker identity (19.9% and 16.2%), the reverse of answerable questions. This suggests that high refusal accuracy on acoustic questions reflects difficulty using acoustic evidence rather than genuine awareness that it is missing.

5

Access to the same evidence declines as history grows, at different rates across evidence types.

From 8K to 32K, mean accuracy drops from 40.1% to 33.0% for proprietary models and from 26.6% to 21.9% for open-weight models, and the decline continues to 64K for all four evidence types. At 64K, models retain 70.3% of their 8K accuracy on speech semantics and 69.8% on paralinguistic cues, against 66.5% for speaker identity and 63.2% for environmental sound.

Line plots of accuracy by evidence type from 8K to 64K: raw accuracy on the left and retention relative to each type's 8K baseline on the right.
Accuracy scaling by evidence type (8K–64K). Left: raw accuracy. Right: retention relative to each type's own 8K baseline.
6

Error profiles differ qualitatively across both evidence types and memory operations.

Speaker errors are most often binding failures (48%): the cue is recovered but attached to the wrong person. Paralinguistic errors are dominated by localization failures (63%): the answer-critical vocal cue is often not retrieved at all. Environmental errors split between localization (41%) and unsupported answers (39%).

  • Evidence recognition / localization
  • Binding / association
  • Operation execution
  • Unsupported answer
  • Not attributable from the response
Stacked bars of error classes at 64K for the IE, MSR and TET memory operations.
(a) by memory operation
Stacked bars of error classes at 64K for the speaker, paralinguistic and environmental evidence types.
(b) by evidence type
05

How VoxMem is built

Questions are planned from an operation, an evidence type and an answer; evidence is realized as spoken dialogue with controlled acoustic cues; histories are built at 8K–64K with the question and evidence fixed; multi-stage quality control validates each step.

The VoxMem construction pipeline: task and evidence design, dialogue and audio realization, variants and context assembly with haystack and filler sessions, and validation into the final benchmark.
Construction pipeline. Histories nest across budgets, so a length comparison holds the item fixed and varies only the surrounding sessions.

Audio is required, by construction

Replacing every user turn with its exact transcript barely changes accuracy on speech semantics, but removes most of the signal for the three audio-native evidence types. Answering them requires the audio itself.

Distractor sessions are labelled per history: haystack sessions share the question's topic or retrieval key and are ruled out only by the selector the question states, and filler sessions pad the history to the target budget.

Accuracy (%) of Gemini-3.7-Flash on the 669 answerable questions at 8K, with full audio versus transcripts only.
Evidence typenFull audioTranscriptΔ
Speaker12669.810.359.5
Paralinguistic22242.83.239.6
Environmental12628.60.827.8
Semantics19575.971.04.9
All66954.923.831.0
06

Get started

The data streams from the Hugging Face Hub; pick a context length and, optionally, a single evidence type. See the README for local models, cluster runs and the output format.

# install
pip install -r requirements.txt

# smoke test: the abstain baseline must score 1.0 on refusal and 0.0 on answerable
python run_voxmembench.py --config 8k_speaker_information \
    --model abstain --allow-abstention --limit 25 --out predictions_smoke.jsonl
python score_voxmembench.py --predictions predictions_smoke.jsonl --judge exact-match

# a real run at 32K, scored by an LLM judge
python run_voxmembench.py --config 32k --model openai:gpt-4o-audio-preview \
    --allow-abstention --out predictions_32k.jsonl
python score_voxmembench.py --predictions predictions_32k.jsonl \
    --judge openai:gpt-4o-mini --out metrics_32k.json
07

BibTeX

If you find VoxMem useful, please cite:

@article{xiao2026voxmem,
  title   = {VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models},
  author  = {Xiao, Yang and Sethu, Vidhyasaharan and Holden, Eun-Jung and Dang, Ting},
  journal = {arXiv preprint arXiv:2609.32607},
  year    = {2026}
}