Structured Role-Aware Policy Optimization for Multimodal Reasoning
结构化角色感知策略优化用于多模态推理
Bingqing Jiang, Difan Zou
机构
*
School of Computing & Data Science, The University of Hong Kong(计算与数据科学学院,香港大学)
;
School of Computing & Data Science and Institute of Data Science, The University of Hong Kong(计算与数据科学学院和数据科学研究所,香港大学)
机构
*
School of Computer Science and Engineering, Tianjin University of Technology(天津理工大学计算机科学与工程学院)
;
School of Artificial Intelligence, Tianjin University(天津大学人工智能学院)
机构
*
Shanghai University of Engineering Science(上海工程技术大学)
;
Tencent YouTu Lab(腾讯优图实验室)
;
Clinical Research Unit, Zhongshan Hospital of Fudan University(复旦大学中山医院临床研究部)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.CV、cs.AI
DepthPilot: From Controllability to Interpretability in Colonoscopy Video Generation
DepthPilot:从可控性到可解释性在结肠镜视频生成中
Junhu Fu, Ke Chen, Weidong Guo, Shuyu Liang, Jie Xu, Chen Ma, Kehao Wang, Shengli Lin, Zeju Li, Yuanyuan Wang, Yi Guo, Shuo Li
机构
*
College of Biomedical Engineering, Fudan University, Shanghai 200433, China(复旦大学生物医学工程学院)
;
Key Laboratory of Medical Imaging Computing and Computer Assisted Intervention of Shanghai, Shanghai 200032, China(上海医学影像计算与计算机辅助干预重点实验室)
;
Endoscopy Research Institute, Zhongshan Hospital, Fudan University, Shanghai 200032, China(复旦大学中山医院内窥镜研究所)
;
Shanghai Collaborative Innovation Center of Endoscopy, Shanghai 200032, China(上海内窥镜协同创新中心)
;
Department of Biomedical Engineering, Case Western Reserve University, Cleveland, OH 44106, USA(凯斯西储大学生物医学工程系)
;
Department of Computer and Data Science, Case Western Reserve University, Cleveland, OH 44106, USA(凯斯西储大学计算机与数据科学系)
ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers
ViTaPEs: 视觉触觉位置编码用于多模态转换器中的跨模态对齐
Fotios Lygerakis, Ozan Özdenizci, Elmar Rückert
机构
*
Chair of Cyber-Physical Systems(感知系统教授席位)
;
Institute of Machine Learning and Neural Computation(机器学习与神经计算研究所)
;
Technical University of Leoben(莱比锡技术大学)
;
Graz University of Technology(格拉茨技术大学)
Wentao Zhang, Qi Zhang, Mingkun Xu, Mu You, Henghua Shen, Zhongzhi He, Keyan Jin, Derek F. Wong, Tao Fang
机构
*
Business School, Shandong University of Technology(山东理工大学商学院)
;
Faculty of Data Science, City University of Macau(澳门城市大学数据科学学院)
;
Guangdong Institute of Intelligent Science and Technology(广东智能科学与技术研究院)
;
Macau Millennium College(澳门 millennium 学院)
;
Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系)
CommentsThis work is an expanded version of our prior paper published in the IEEE ICASSP 2026 conference arXiv:2512.24947, from 4 to 20+ pages, presenting a well-structured and principled framework, extensive experiments, and deeper insights. Tao Fang is the corresponding author
REVEAL: Multimodal Vision-Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction
REVEAL:多模态视图-语言对齐的视网膜形态学与临床风险预测
Seowung Leem, Lin Gu, Chenyu You, Kuang Gong, Ruogu Fang
机构
*
J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida(佛罗里达大学J. Crayton Pruitt家族生物医学工程系)
;
Research Institute of Electrical Communication, Tohoku University(东北大学电气通信研究所)
;
Department of Applied Mathematics & Statistics, Stony Brook University(石溪大学应用数学与统计学系)
;
Department of Computer Science, Stony Brook University(石溪大学计算机科学系)
CommentsAccepted at the IEEE Engineering in Medicine and Biology Society Annual International Conference (Proceedings of the 48th International Conference), 2026
MAny: Merge Anything for Multimodal Continual Instruction Tuning
MAny:为多模态连续指令微调合并任何内容
Zijian Gao, Wangwang Jia, Xingxing Zhang, Pengfei Qian, Tao Sun, Bo Ding, Yong Dou, Huaimin Wang, Kele Xu
机构
*
College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机科学与技术学院)
;
National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology(国防科技大学并行与分布式计算国家重点实验室)
;
State Key Laboratory of Complex & Critical Software Environment(复杂与关键软件环境国家重点实验室)
;
School of Computer Science, Tsinghua University(清华大学计算机科学与技术系)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.AI、cs.LG
Yueying Li, Fengxiang Wang, Yan Li, Mingshuo Chen, Mengying Zhao, Long Lan
机构
*
College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机科学与技术学院)
;
School of Computer Science, Beijing University of Posts and Telecommunications(北京邮电大学计算机学院)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.CV、cs.AI
Temporal Inversion for Learning Interval Change in Chest X-Rays
时间倒置用于学习胸部X光片的区间变化
Hanbin Ko, Kyungmin Jeon, Doowoong Choi, Chang Min Park
机构
*
Interdisciplinary Program in Bioengineering, Seoul National University Graduate School(首尔大学研究生院生物工程跨学科项目)
;
Integrated Major in Innovative Medical Science, Seoul National University Graduate School(首尔大学研究生院创新医学科学综合专业)
;
Dept. of Radiology, Seoul National University Hospital(首尔大学医院放射科)
;
Seoul National University College of Medicine(首尔大学医学院)
;
Inst. of Medical and Biological Engineering, Seoul National University Medical Research Center(首尔大学医学研究中心医学与生物工程研究所)