Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox
音频大语言模型是听还是读?使用VoxParadox分析和缓解副语言失败
机构 * Institute for Creative Technologies, University of Southern California, Los Angeles, USA(创意技术研究所,南加州大学,洛杉矶,美国)
专题命中 后训练与偏好优化 :LLM(abstract);large language model(abstract);language model(abstract);preference optimization(abstract)
AI总结 针对音频大语言模型在副语言理解上的不足,提出对抗性基准VoxParadox和Prompt-Conditioned Layer Mixer方法,显著提升模型对副语言线索的利用能力。
Comments Accepted as a conference paper at ICML 2026. Project page: https://voxparadox.github.io/