Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt
迈向细粒度时间感知:通过音频侧时间提示进行后训练的大音频-语言模型
机构 * National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China(语音与语言信息处理国家级工程研究中心,中国科学技术大学,合肥,中国) ; ICT Cluster, Singapore Institute of Technology, Singapore(新加坡理工学院ICT集群,新加坡)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
AI总结 本文提出Audio-Side Time Prompt方法,结合强化学习改进大音频-语言模型的时间感知能力,在音频定位、事件检测等任务中取得显著提升。
Comments Submitted to Interspeech 2026