机构
*
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
School of Computing and Information Technology, Great Bay University(大湾区大学计算与信息科技学院)
;
Dongguan Key Laboratory for Intelligence and Information Technology(东莞智能与信息处理重点实验室)
;
Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室)
;
Tencent Youtu Lab(腾讯优图实验室)
;
School of Psychology, Shanghai Jiao Tong University(上海交通大学心理学院)
;
Faculty of Applied Sciences, Macao Polytechnic University(澳门理工学院应用科学学院)
;
The Chinese University of Hong Kong(香港中文大学)
CommentsPreprint. Accepted at NeurIPS 2025 Workshops on SPACE in Vision, Language, and Embodied AI (SpaVLE) as Oral, Embodied World Models for Decision Making (EWM), Aligning Reinforcement Learning Experimentalists and Theorists (ARLET), and Scaling Environments for Agents (SEA)
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
用于以音频为中心任务的音频-语言模型:系统综述
Yi Su, Jisheng Bai, Qisheng Xu, Kele Xu, Yong Dou
机构
*
College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学)
;
School of Communications and Information Engineering, Xi’an University of Posts and Telecommunications(通信与信息工程学院,西安邮电大学)
Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions
视频中矛盾/犹豫识别用于个性化数字健康干预
Manuela González-González, Soufiane Belharbi, Muhammad Osama Zeeshan, Masoumeh Sharafi, Muhammad Haseeb Aslam, Lorenzo Sia, Nicolas Richet, Marco Pedersoli, Alessandro Lameiras Koerich, Simon L Bacon, Eric Granger
机构
*
LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(ETS蒙特利尔大学系统工程系LIVIA实验室)
;
LIVIA, Dept. of Software and IT Engineering, ETS Montreal, Canada(ETS蒙特利尔大学软件与信息工程系LIVIA实验室)
;
Dept. of Health, Kinesiology, & Applied Physiology, Concordia University, Montreal, Canada(康科迪亚大学健康、运动科学与应用生理学系)
;
Montreal Behavioural Medicine Centre, CIUSSS Nord-de-l’Ile-de-Montréal, Canada(蒙特利尔行为医学中心,蒙特利尔北岛卫生与社会服务局)
CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors
CorridorVLA:通过稀疏锚点实现生成动作头的显式空间约束
Dachong Li, ZhuangZhuang Chen, Jin Zhang, Jianqiang Li
机构
*
College of Computer Science and Software Engineering(计算机科学与软件工程学院)
;
National Engineering Laboratory for Big Data System Computing Technology(大数据系统计算技术国家工程实验室)
CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation
CoLA-Flow Policy: 通过连续潜在动作流匹配实现机器人操作的时序一致模仿学习
Wu Songwei, Jiang Zhiduo, Sun Wandong, Xie Guanghu, Zhao Rui, Liu Hong, Liu Yang
机构
*
State Key Laboratory of Robotics and System, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)
;
The University of Sydney(悉尼大学)
;
Honor Device Co., Ltd.(荣耀设备有限公司)
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
检索头能看见图像吗?长上下文视觉语言模型中的多模态检索头
Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao, Yiming Du, Haobo Li, Xiyu Ren, Ginny Wong, Simon See, Lishu Luo, Haodong Duan, Pasquale Minervini, Yangqiu Song
机构
*
HKUST(香港科技大学)
;
University of Edinburgh(爱丁堡大学)
;
CUHK(香港中文大学)
;
NVAITC, NVIDIA, Santa Clara, USA(NVIDIA Santa Clara 分公司)
;
Tsinghua University(清华大学)
Language-guided Medical Image Segmentation with Target-informed Multi-level Contrastive Alignments
基于目标感知多级对比对齐的语言引导医学图像分割
Mingjian Li, Mingyuan Meng, Shuchang Ye, Mingye Zou, Michael Fulham, Lei Bi, Jinman Kim
机构
*
School of Computer Science, The University of Sydney(悉尼大学计算机科学学院)
;
Institute of Translational Medicine, Shanghai Jiao Tong University(上海交通大学转化医学研究院)
;
Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China(北京中关村学院及中关村人工智能研究院)
;
Department of Molecular Imaging, Royal Prince Alfred Hospital(皇家珀斯阿尔弗雷德医院分子成像部)
CommentsAccepted to ICML 2026. This version updates the ICML submission with an optimized model checkpoint. Project page: https://omni-diffusion.github.io
Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation
区域感知多模态大语言模型:基于慢快标记化与伪掩码引导的3D CT报告生成
Sunggu Kyung, Jinyoung Seo, Hyunseok Lim, Dongyeong Kim, Hyungbin Park, Jimin Sung, Jihyun Kim, Wooyoung Jo, Yoojin Nam, Namkug Kim
机构
*
Department of Convergence Medicine, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Republic of Korea(韩国首尔峨山医疗中心蔚山大学医学院融合医学系)
;
University of Ulsan College of Medicine, Seoul, Republic of Korea(韩国首尔蔚山大学医学院)
;
Department of Radiology and Research Institute of Radiology, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Republic of Korea(韩国首尔峨山医疗中心蔚山大学医学院放射科与放射学研究所)
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(信息多媒体国家重点实验室,计算机学院,北京大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Institute for Brain and Intelligence, Fudan University(脑与智能研究院,复旦大学)
;
University of Science and Technology Beijing(北京科技大学)
;
Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
机构
*
Southwestern University of Finance and Economics(西南财经大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Central South University(中南大学)
;
Hithink Research(Hithink研究)
;
Westlake University(西湖大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
University of Manchester(曼彻斯特大学)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
University of Adelaide(阿德莱德大学)
;
Fudan University(复旦大学)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
Chengdu Everimaging Science and Technology Co., Ltd(成都亿联科技有限公司)