arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6878 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6878 篇

2605.23961 2026-05-26 q-bio.BM cs.AI cs.LG 79%

Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation

多模态对齐与偏好优化用于零样本条件RNA生成

Roman Klypa, Alberto Bietti, Sergei Grudinin

机构 * Univ. Grenoble Alpes, CNRS, Grenoble INP, LJK(格勒诺布尔阿尔卑斯大学、法国国家科学研究中心、格勒诺布尔INP、LJK实验室) Center for Computational Mathematics, Flatiron Institute(计算数学中心、Flatiron研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 提出Moirain框架,通过多模态监督微调和直接偏好优化实现条件RNA序列生成,在零样本条件下生成具有高结合亲和力的生物合理RNA序列。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20192 2026-05-25 cs.CL cs.CE cs.CR cs.CY q-fin.CP 79%

Leveraging Large Language Models for Sentiment Analysis: Multi-Modal Analysis of Decentraland's MANA Token

利用大语言模型进行情感分析:Decentraland的MANA代币多模态分析

Xintong Wu, Peiting Tsai, Jing Yuan, Michael Yu, Greg Sun, Luyao Zhang

机构 * University of Pennsylvania(宾夕法尼亚大学) Microsoft(微软) Duke Kunshan University(杜克昆山大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

AI总结 本研究利用基于BERT的大语言模型对Decentraland的Discord社区进行情感分析,并结合多模态金融数据(历史价格、情感分数、交易量和市值)构建LSTM模型,发现多模态模型在MANA代币回报预测中显著优于仅基于价格的基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22090 2026-05-22 cs.AI 79%

A Camera-Cooperative ISAC Framework for Multimodal Non-Cooperative UAVs Sensing

一种用于多模非合作无人机感知的相机协作ISAC框架

Wenfeng Wu, Luping Xiang, Kun Yang

机构 * State Key Laboratory of Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Institute of Intelligent Networks and Communications (NINE)(智能网络与通信研究院) School of Intelligent Software and Engineering, Nanjing University (Suzhou Campus)(南京大学智能软件与工程学院(苏州校区))

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出了一种相机协作ISAC框架,通过多模感知实现高效的无人机波束定向和跟踪,提升了感知精度和资源效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21239 2026-05-21 cs.MM 79%

Multimodal Emotion Recognition with Large Language Models

基于大语言模型的多模态情感识别

Hongrui Zhang, Daiqing Wu, Yangyang Li, Kuien Liu, Yuhui Wang, Yu Zhou, Sicheng Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

AI总结 本文探讨了多模态情感识别中利用大语言模型的范式,总结了现有研究在情感数据增强、多模态情感表示和多模态情感推理三个方向上的进展与挑战,旨在为该领域的发展提供清晰的学术路线图。

Comments Accepted by IJCAI 2026 Survey Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20891 2026-05-21 cs.CV 79%

HDMoE: A Hierarchical Decoupling-Fusion Mixture-of-Experts Framework for Multimodal Cancer Survival Prediction

HDMoE:一种用于多模态癌症生存预测的分层解耦-融合专家混合框架

Huayi Wang, Haochao Ying, Yuyang Xu, Qiyao Zheng, jun wang, Cheng Zhang, Ying Sun, Jian Wu

机构 * Zhejiang University(浙江大学) Xinjiang University(新疆大学) Hangzhou City University(杭州市大学) Sun Yat-sen University Cancer Center(中山大学肿瘤中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出HDMoE框架,通过分层解耦-融合专家混合方法,有效整合多模态医学数据以提高癌症生存预测的准确性,解决了传统方法中特征解耦和融合效果不佳的问题。

Comments 12 pages, HDMoE has been accepted by KDD 2026 AI for Sciences Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20584 2026-05-21 cs.CV 79%

QwenSafe: Multimodal Content Rating Description Identification via Preference-Aligned VLMs

QwenSafe: 通过偏好对齐的视觉语言模型实现多模态内容评级描述识别

Dishanika Denipitiyage, Aruna Seneviratne, Suranga Seneviratne

机构 * University of Sydney(悉尼大学) University of New South Wales(新南威尔士大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出QwenSafe,一种通过联合推理应用元数据和截图来自动识别苹果定义的内容评级描述(CRDs)的视觉语言模型,通过引入metadata2CRD数据构建管道和直接偏好优化(DPO)提升模型预测准确性,实验结果显示QwenSafe在二元CRD分类中显著优于现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13081 2026-05-21 cs.CV 79%

PRA-PoE: Robust Multimodal Alzheimer's Diagnosis with Arbitrary Missing Modalities

PRA-PoE: 基于任意缺失模态的鲁棒多模态阿尔茨海默病诊断

Guangqian Yang, Ye Du, Wenlong Hou, Qian Niu, Shujun Wang

机构 * Department of Biomedical Engineering, The Hong Kong Polytechnic University, Hong Kong SAR, China(生物医学工程系,香港理工大学,香港特别行政区,中国) Department of Technology Management for Innovation, The University of Tokyo, Japan(创新技术管理系,东京大学,日本) Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University, Hong Kong SAR, China(数据科学与人工智能系,香港理工大学,香港特别行政区,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究提出PRA-PoE框架,通过原型锚定表示对齐和不确定性感知专家融合机制,解决多模态学习中模态缺失导致的表示偏移问题,提升了在不同缺失模式下的诊断鲁棒性与准确性。

Comments Early accepted by MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12960 2026-05-21 cs.CL 79%

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging

DiM\textsuperscript{3}: 通过方向和幅度感知融合连接多语言和多模态模型

Zijing Wang, Mingyang Wang, Ercong Nie, Yongkang Liu, Shi Feng, Mengjie Zhao, Daling Wang, Xiaocui Yang, Hinrich Schütze

机构 * Northeastern University, China(东北大学,中国) CIS, LMU Munich, Germany(慕尼黑莱布尼茨大学计算机学院,德国) Munich Center for Machine Learning (MCML), Germany(慕尼黑机器学习中心(MCML),德国) Shanghai Jiao Tong University, China(上海交通大学,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 本研究提出DiM3方法,通过在共享语言模型骨干中融合残差更新,实现多语言和多模态能力的无缝整合,从而提升多语言性能并保持多模态能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20449 2026-05-21 cs.LG cs.AI 79%

LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series

LLM预训练塑造了可泛化的流形:跨模态迁移至时间序列的洞察

Alexis Roger, Prateek Humane, Zhenghan Tai, Gwen Legate, Andrei Mircea, Vasilii Feofanov, Irina Rish

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) University of Toronto(多伦多大学) Concordia University(康科迪亚大学) com(42.com)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

AI总结 研究探讨了语言预训练的Transformer能否成为有效的时序预测器,并揭示了跨模态迁移的机制,指出预训练构建了流形,微调则将数值动态投影到任务相关方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18104 2026-05-19 cs.AI cs.CR 79%

Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction

多模态大语言模型中的安全几何坍缩与自适应漂移修正

Jiahe Guo, Xiangran Guo, Jiaxuan Chen, Weixiang Zhao, Yanyan Zhao, Yutai Hou, Qianchao Wang, Dandan Tu, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Huawei Technologies Co., Ltd(华为技术有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文研究了多模态大语言模型在跨模态安全转移中的不足,提出安全几何坍缩现象,并通过自适应漂移修正方法提升模型安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17949 2026-05-19 cs.CV 79%

SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning

SkyNative: 一种面向遥感视觉证据推理的原生多模态框架

Xiao Yang, Ronghao Fu, Zhiwen Lin, Zhuoran Duan, Jiashun Zhu, Jiasen Hu, Lang Sun, Weipeng Zhang, Jiaqi Liu, Xu Na, Haoran Liu, Weijie Zhang, Bo Yang

机构 * College of Computer Science and Technology, Jilin University, China(吉林大学计算机科学与技术学院) Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education(教育部符号计算与知识工程重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出SkyNative,一种原生多模态框架,通过去除预训练视觉骨干,直接在语言模型token空间中表示图像为原始patch tokens,以提升遥感图像的细粒度空间推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21941 2026-05-19 cs.LG cs.AI 79%

Robust Multimodal Representation Learning in Healthcare

医疗领域鲁棒多模态表征学习

Xiaoguang Zhu, Linxiao Gong, Lianlong Sun, Yang Liu, Haoyu Wang, Jing Liu

机构 * University of California, Davis(加州大学戴维斯分校) HKUST (GZ)(香港科技大学) University of Rochester(罗切斯特大学) Tongji University(同济大学) Georgia Institute of Technology(佐治亚理工学院) Fudan University(复旦大学) The University of British Columbia(不列颠哥伦比亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出双流特征去相关框架,通过结构因果分析处理医疗多模态数据中的系统性偏差,提升模型泛化能力,实验验证在MIMIC-IV、eICU和ADNI数据集上的性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15460 2026-05-18 cs.IR cs.AI 79%

Differentially Private Motif-Preserving Multi-modal Hashing

差分隐私的动机保持多模态哈希

Zehua Cheng, Wei Dai, Jiahao Sun

机构 * Department of Computer Science\ of Oxford Oxford United Kingdom Department of Computer Science\ of Oxford

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);分类 cs.AI

AI总结 本文提出DMP-MH框架,通过去噪后蒸馏方法在保证隐私的前提下保留多模态数据的结构特征,实验表明其在保持隐私的同时提升了检索性能。

Comments 9 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14406 2026-05-15 cs.LG cs.CV 79%

GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation

GeoViSTA:用于多模态环境表示的地理视觉-表格变压器

Yuhao Liu, Sadeer Al-Kindi, Ashok Veeraraghavan, Guha Balakrishnan

机构 * Department of Electrical and Computer Engineering, Rice University(理海大学电气与计算机工程系) Center for Cardiovascular Computational and Precision Health, Department of Cardiology, DeBakey Heart and Vascular Center, Houston Methodist(休斯顿方法主义医疗中心心血管计算与精准健康中心、心内科部门、德贝基心脏和血管中心)

专题命中 多模态训练与对齐 :multimodal(title);cross-modal(abstract);分类 cs.CV

AI总结 GeoViSTA通过融合栅格影像与表格数据,构建统一的地理嵌入表示,提升对环境、社会和健康问题的推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14327 2026-05-15 cs.LG cs.AI 79%

AIM-DDI: A Model-Agnostic Multimodal Integration Module for Drug-Drug Interaction Prediction

AIM-DDI: 一种模型无关的多模态整合模块用于药物-药物相互作用预测

Yerin Park, Sangseon Lee

机构 * Department of Artificial Intelligence, Inha University(人工智能系,Inha大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出AIM-DDI模块,通过共享潜在空间整合多模态信息,提升药物-药物相互作用预测的鲁棒性,尤其在未见过的药物情况下表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22853 2026-05-14 cs.CV 79%

Inference-Time Dynamic Modality Selection for Incomplete Multimodal Classification

推理时动态模态选择用于不完整多模态分类

Siyi Du, Xinzhe Luo, Declan P. O'Regan, Chen Qin

机构 * Department of Electrical and Electronic Engineering & I-X(电气与电子工程系及I-X)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出DyMo框架,通过动态选择并融合可靠恢复的模态,解决不完整多模态学习中的丢弃或填补困境,实验显示其在多种缺失数据场景下优于现有方法。

Comments 27 pages (including appendix), accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12064 2026-05-13 cs.CV 79%

TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

TAR:基于文本语义的跨模态图像配准框架用于光学和SAR图像

Zhuoyu Cai, Dou Quan, Ning Huyan, Pei He, Shuang Wang, Licheng Jiao

机构 * Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education of China, School of Artificial Intelligence, Xidian University(中国教育部智能感知与图像理解重点实验室,西安电子科技大学人工智能学院) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出TAR框架,通过文本语义先验缓解模态差距,提升跨模态特征学习,解决大形变下的光学与SAR图像配准问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11931 2026-05-13 cs.CV 79%

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training

学会思考:通过视觉感知的自我改进训练提升多模态推理

Qihuang Zhong, Liang Ding, Wenjie Xuan, Juhua Liu, Bo Du, Dacheng Tao

机构 * School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence(计算机学院、多媒体软件国家工程研究中心、人工智能研究院) Hubei Key Laboratory of Multimedia(湖北多媒体重点实验室) Network Communication Engineering, Wuhan University, China(网络通信工程、武汉大学,中国) The University of Sydney, Australia(悉尼大学,澳大利亚) Nanyang Technological University, Singapore(南洋理工大学,新加坡)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出VISTA框架,通过视觉感知的自我改进训练提升多模态推理能力,解决数据不平衡和语言先验偏差问题,实验显示在多种训练场景下提升性能。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10500 2026-05-13 cs.CV 79%

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning

视觉增强深度缩放用于多模态潜在推理

Yudong Han, Yong Wang, Zaiquan Yang, Zhen Qu, Liyuan Pan, Xiangxiang Chu

机构 * Beijing Institute of Technology(北京理工大学) AMAP, Alibaba Group(阿里集团AMAP) City University of Hong Kong(香港城市大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Yangtze Delta Region Academy of Beijing Institude of Technology, Jiaxing, China(北京理工大学扬子江地区学院,嘉兴,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出视觉回放模块和路由深度缩放,通过增强视觉感知和细化复杂潜在表示,提升多模态潜在推理的效率与性能。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11716 2026-05-13 cs.AI 79%

SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models

SafeSteer: 多模态大语言模型中的解码级防御机制

Xinyi Zeng, Xue Yang, Jingyuan Zhang, Huanqian Yan, Xiang Chen, Kaiwen Wei, Hankun Kang, Yu Tian

机构 * Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Kuaishou Technology(快手科技) School of Computer Science and Technology, Beihang University(北航计算机科学与技术学院) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Chongqing University(重庆大学) Wuhan University(武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出SafeSteer,通过解码阶段的轻量探针和模态语义对齐向量,提升多模态大语言模型的安全性,实验表明其能提升33.40%的安全性而不需微调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11015 2026-05-13 cs.CR cs.AI 79%

DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization

DCVD:双通道跨模态融合用于联合漏洞检测与定位

Wenxin Tang, Wenbin Li, Junliang Liu, Jingyu Xiao, Xi Xiao, Mingzhe Liu, Jinlong Yang, Xuan Liu, Yuehe Ma, Wang Luo, Qing Li, Lei Wang, Peng Xiangli

机构 * Tsinghua University(清华大学) Hunan University(湖南大学) Dalian Maritime University(大连海事大学) The Chinese University of Hong Kong(香港中文大学) Shenzhen University(深圳大学) Northwestern Polytechnical University(西北工业大学) Shandong University(山东大学) BNU-HKBU United International College(北京师范大学-香港浸会大学联合国际学院) Sun Yat-sen University(中山大学) Peng Cheng Laboratory(鹏城实验室) Guangzhou Intelligence Communications Technology Co., Ltd.(广州智能通信技术有限公司) The Fifth Electronic Research Institute of MIIT(中华人民共和国信息产业部第五电子研究所)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出DCVD框架,通过双通道融合实现功能级检测与语句级定位的联合优化,有效解决单一信息源和缺乏显式监督的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09614 2026-05-12 cs.CV 79%

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

反射锚点用于长链多模态推理中的传播感知视觉保留

Xuan Gong, Hanbo Huang, Hao Zheng, Yiran Zhang, Wenbin Dai, Weishu Zhao, Shiyu Liang

机构 * Shanghai Jiao Tong University(上海交通大学) Lanzhou University(兰州大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出RAPO方法,通过信息论分析优化视觉传播潜力和局部分支空间,提升长链多模态推理的视觉信息保留效果。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09352 2026-05-12 cs.AI 79%

The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?

维特根斯坦表示假说:语言是否是多模态收敛的吸引子?

Zhaoyang Zhang, Run Shao, Dongyue Wu, Jiajie Teng, Chao Tao, Jingdong Chen, Haifeng Li

机构 * Central South University(中南大学) Huazhong University of Science and Technology(华中科技大学) Shanghai Jiao Tong University(上海交通大学) Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 研究探讨了为何不同模态的神经网络收敛于共享表示,发现语言模态对其他模态有显著方向性吸引,提出维特根斯坦表示假说。

Comments 22 pages, 11 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09151 2026-05-12 cs.CV 79%

MultiMedVision: Multi-Modal Medical Vision Framework

多模态医学视觉框架:MultiMedVision

Frank Li, Bardia Khosravi, Mohammadreza Chavoshi, Young Seok Jeon, Theo Dapamede, Hari Trivedi, Janice Newsome, Judy Gichoya

机构 * Emory University(埃默里大学) Yale University(耶鲁大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出MultiMedVision框架,通过稀疏视觉Transformer实现2D/3D医学影像的统一表征学习,无需模态特定适配器,在共享潜在空间中处理混合模态数据,取得2D和3D任务的竞争力表现。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08764 2026-05-12 cs.LG cs.CV eess.IV 79%

Anchoring the Eigengap: Cross-Modal Spectral Stabilization for Sample-Efficient Representation Learning

锚定特征间隙:跨模态谱稳定化以实现样本高效表征学习

Nikhil J. Dhinagar, Vidhi Chhatbar, Chirag Jagad, Pavithra Senthilkumar, Sophia I. Thomopoulos, Mahir H. Khan, Sook-Lei Liew, the ENIGMA-Stroke Recovery Working Group, Paul M. Thompson

机构 * Imaging Genetics Center, Mark & Mary Stevens Neuroimaging & Informatics Institute, Keck School of Medicine, University of Southern California(影像基因中心,马克与玛丽史蒂文斯神经影像与信息学研究所,凯克医学院,南加州大学) Neuroscience Graduate Program, Mark & Mary Stevens Neuroimaging & Informatics Institute, Chan Division of Occupational Science & Occupational Therapy, Biomedical Engineering, University of Southern California(神经科学研究生项目,马克与玛丽史蒂文斯神经影像与信息学研究所,查恩职业科学与职业治疗 division,生物医学工程,南加州大学)

专题命中 多模态训练与对齐 :cross-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出跨模态谱稳定化方法,通过抑制噪声主导方向并保持特征间隙,提升样本效率下的表征学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07274 2026-05-11 cs.AI cs.LG 79%

Structured Role-Aware Policy Optimization for Multimodal Reasoning

结构化角色感知策略优化用于多模态推理

Bingqing Jiang, Difan Zou

机构 * School of Computing & Data Science, The University of Hong Kong(计算与数据科学学院,香港大学) School of Computing & Data Science and Institute of Data Science, The University of Hong Kong(计算与数据科学学院和数据科学研究所,香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出结构化角色感知策略优化(SRPO),通过角色感知的token级信用分配提升多模态推理中的证据基础推理能力,无需外部奖励模型。

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06990 2026-05-11 cs.CV cs.LG 79%

TRAJGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations

TRAJGANR: 通过地理对齐的神经表示进行轨迹导向的城市多模态学习

Maria Despoina Siampou, Gengchen Mai, Ni Lao, Jinmeng Rao, Neha Arora, Cyrus Shahabi, Shushman Choudhury

机构 * Google Research, Mountain View, CA(谷歌研究,山景城,加利福尼亚州) Google LLC, Mountain View, CA(谷歌公司,山景城,加利福尼亚州) Dept. of Computer Science, University of Southern California, Los Angeles, CA(计算机科学系,南加州大学,洛杉矶,加利福尼亚州) SEAI Lab, Dept. of Geography and the Environment, The University of Texas at Austin, Austin, TX(SEAI实验室,地理与环境系,德克萨斯大学奥斯汀分校,奥斯汀,德克萨斯州)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 TRAJGANR提出一种新的多模态自监督学习框架,通过将连续移动模式与静态位置观察对齐,提升城市理解和移动任务的性能,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19316 2026-05-08 cs.CL 79%

KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls

KORE:通过知识导向控制增强大型多模态模型的知识注入

Kailin Jiang, Hongbo Jiang, Ning Jiang, Zhi Gao, Jinhe Bi, Yuchen Ren, Bin Li, Yuntao Du, Lei Liu, Qing Li

机构 * University of Science State Key Laboratory of General Artificial Intelligence, BIGAI Xiamen University Northeast Forestry University Beijing Institute of Technology Ludwig Maximilian University of Munich The University of Sydney C-FAIR\&school of software, Shandong University State Key Lab. for Novel Software Technology, Nanjing University, P.R. China

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 KORE通过知识导向的增强和约束,提升大型多模态模型的知识注入能力,同时保留旧知识。方法利用协方差矩阵和投影初始化,有效减少灾难性遗忘。

Comments ICML 2026, Project Page: https://kore-lmm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05850 2026-05-08 cs.CV 79%

Align3D-AD: Cross-Modal Feature Alignment and Dual-Prompt Learning for Zero-shot 3D Anomaly Detection

Align3D-AD:跨模态特征对齐与双提示学习用于零样本3D异常检测

Letian Bai, Xuanming Cao, Juan Du, Chengyu Tao

机构 * Smart Manufacturing Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)智能制造方向) The Hong Kong University of Science and Technology(香港科技大学) College of Mechanical and Vehicle Engineering, Hunan University(湖南大学机械与车辆工程学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出Align3D-AD框架,通过跨模态特征对齐和双提示学习解决零样本3D异常检测中的领域差距问题,实验表明其在多个数据集上均优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01327 2026-05-08 cs.AI cs.LG 79%

Segment-Aligned Policy Optimization for Multi-Modal Reasoning

基于段落对齐的策略优化用于多模态推理

Lei Gao, Zhuoming Li, Mengxi Jia, Jiakang Yuan, Hongbo Sun, Hao Sun, Xuelong Li

机构 * Fudan University(复旦大学) Southeast University(东南大学) China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd.(中国电信人工智能技术(北京)有限公司) Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出SAPO方法,通过将推理步骤而非token或完整序列作为策略更新的基本单元,提升多模态推理任务的准确性和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏