Current LALMs remain far from reliable spoken conversational memory.
At the 32K budget no evaluated model exceeds 40% overall accuracy; the strongest reaches 38.5%. The five proprietary models average 33.0% and the ten open-weight models 21.9%.
Do audio language models remember not only what was said, but who said it, how it was said, and what was audible?
Every question is posed over a multi-session spoken history. The answer may depend on what a user said, on which speaker said it, on how it was said, or on what could be heard in the background.
Across 15 LALMs, no model exceeds 40% overall accuracy at 32K. Models retain what was said far better than who said it, how, or what was audible, and the gap grows with history length.
Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single continuous recording. Existing benchmarks fall short on all three dimensions: they focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem. We argue that principled memory evaluation requires jointly characterizing the acoustic evidence to be retained and the operations applied to it, and introduce a taxonomy along these two axes. Building on this taxonomy, we present VoxMem: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal tracking, and answer refusal), grounded in multi-session histories and stratified across context budgets from 8K to 64K tokens. Evaluating 15 LALMs, no model exceeds 40% at 32K. Models retain what was said far better than who said it, how, or what was audible, a gap that widens for complex operations, grows with history length, and manifests as qualitatively distinct failure modes across evidence types. VoxMem aims to provide a foundation to measure and drive progress on the full scope of spoken conversational memory.
Acoustic evidence specifies what must be recovered from the history; memory operation specifies how it must be used. Their valid combinations form 15 cells. Select a cell to see details.
Accuracy (%) over all 799 questions at each reference history budget, broken down by answer-critical evidence type. Answer Refusal items are included, with appropriate abstention counted as correct. Click a column header to sort.
| # | Model | Type |
|---|
Each question and its required evidence are held fixed across context budgets, so changes in accuracy across lengths reflect the history, not the question.
At the 32K budget no evaluated model exceeds 40% overall accuracy; the strongest reaches 38.5%. The five proprietary models average 33.0% and the ten open-weight models 21.9%.
At 32K, proprietary models average 55.6% on Speech Semantics, but 32.7% on Speaker Identity, 20.0% on Paralinguistic Cues and 21.9% on Environmental Sound. Open-weight models show the same separation (31.6%, 26.7%, 14.5% and 15.5%).
Multi-session reasoning is relatively robust (41.8% on semantics, 43.1% on speaker identity). Temporal tracking works for semantics (44.5%) but nearly collapses for paralinguistic (3.4%) and environmental (1.2%) evidence: maintaining states carried by vocal delivery or ambient sound is the bottleneck, not temporal reasoning alone. Explore the cells in the matrix above.
Models abstain more successfully on paralinguistic and environmental questions (34.3% and 34.2%) than on semantics and speaker identity (19.9% and 16.2%), the reverse of answerable questions. This suggests that high refusal accuracy on acoustic questions reflects difficulty using acoustic evidence rather than genuine awareness that it is missing.
From 8K to 32K, mean accuracy drops from 40.1% to 33.0% for proprietary models and from 26.6% to 21.9% for open-weight models, and the decline continues to 64K for all four evidence types. At 64K, models retain 70.3% of their 8K accuracy on speech semantics and 69.8% on paralinguistic cues, against 66.5% for speaker identity and 63.2% for environmental sound.
Speaker errors are most often binding failures (48%): the cue is recovered but attached to the wrong person. Paralinguistic errors are dominated by localization failures (63%): the answer-critical vocal cue is often not retrieved at all. Environmental errors split between localization (41%) and unsupported answers (39%).
Questions are planned from an operation, an evidence type and an answer; evidence is realized as spoken dialogue with controlled acoustic cues; histories are built at 8K–64K with the question and evidence fixed; multi-stage quality control validates each step.
Replacing every user turn with its exact transcript barely changes accuracy on speech semantics, but removes most of the signal for the three audio-native evidence types. Answering them requires the audio itself.
Distractor sessions are labelled per history: haystack sessions share the question's topic or retrieval key and are ruled out only by the selector the question states, and filler sessions pad the history to the target budget.
| Evidence type | n | Full audio | Transcript | Δ |
|---|---|---|---|---|
| Speaker | 126 | 69.8 | 10.3 | 59.5 |
| Paralinguistic | 222 | 42.8 | 3.2 | 39.6 |
| Environmental | 126 | 28.6 | 0.8 | 27.8 |
| Semantics | 195 | 75.9 | 71.0 | 4.9 |
| All | 669 | 54.9 | 23.8 | 31.0 |
The data streams from the Hugging Face Hub; pick a context length and, optionally, a single evidence type. See the README for local models, cluster runs and the output format.
# install pip install -r requirements.txt # smoke test: the abstain baseline must score 1.0 on refusal and 0.0 on answerable python run_voxmembench.py --config 8k_speaker_information \ --model abstain --allow-abstention --limit 25 --out predictions_smoke.jsonl python score_voxmembench.py --predictions predictions_smoke.jsonl --judge exact-match # a real run at 32K, scored by an LLM judge python run_voxmembench.py --config 32k --model openai:gpt-4o-audio-preview \ --allow-abstention --out predictions_32k.jsonl python score_voxmembench.py --predictions predictions_32k.jsonl \ --judge openai:gpt-4o-mini --out metrics_32k.json
If you find VoxMem useful, please cite:
@article{xiao2026voxmem,
title = {VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models},
author = {Xiao, Yang and Sethu, Vidhyasaharan and Holden, Eun-Jung and Dang, Ting},
journal = {arXiv preprint arXiv:2609.32607},
year = {2026}
}