arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-05-13 至 2026-05-13 共收录 19 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 19 篇

2605.11753 2026-05-13 cs.AI 89%

Towards Visually Grounded Multimodal Summarization via Cross-Modal Transformer and Gated Attention

迈向基于视觉的多模态摘要:通过跨模态Transformer和门控注意力

Abid Ali, Diego Molla-Aliod, Usman Naseem

机构 * School of Computing, Macquarie University(麦考瑞大学计算机学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(title,abstract);分类 cs.AI

AI总结 本文提出SPeCTrA-Sum框架,通过深度视觉处理器和轻量视觉相关性预测器实现文本摘要与图像选择,提升多模态摘要的视觉连贯性与准确性。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11301 2026-05-13 cs.AI cs.CL cs.CV 87%

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?

LatentRouter: 在看到答案之前,我们能否选择合适的多模态模型?

Xueqi Cheng, Yushun Dong

机构 * Department of Computer Science(计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.CL、cs.AI

AI总结 LatentRouter通过多模态效用预测实现多模态大语言模型的路由,通过隐式通信和胶囊修正提升模型选择准确性,在MMR-Bench和VL-RouterBench实验中优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11705 2026-05-13 cs.CV 85%

CAST: Collapse-Aware multi-Scale Topology Fusion for Multimodal Coreset Selection

CAST:面向多模态聚类选择的坍缩感知多尺度拓扑融合

Boran Zhao, Hetian Liu, Zhenxian Hu, Yuqing Yuan, Yu Yan, Pengju Ren

机构 * School of Software Engineering, the National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, National Engineering Research Center for Visual Information and Applications, and Institute of Artificial Intelligence and Robotics(软件工程学院、人机混合增强智能国家重点实验室、视觉信息与应用国家工程研究中心、人工智能与机器人研究院) School of Software Engineering(软件工程学院) XJTU-POLIMI Joint School(西交大-波兰理工联合学院) Faculty of Electronic and Information Engineering(电子与信息工程学院) School of Human Settlements and Civil Engineering(人居与土木工程学院) the National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, National Engineering Research Center for Visual Information and Applications, and Institute of Artificial Intelligence and Robotics(人机混合增强智能国家重点实验室、视觉信息与应用国家工程研究中心、人工智能与机器人研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出CAST框架,通过多尺度拓扑融合解决多模态数据集选择中的跨模态信息失衡和分布不匹配问题,提升跨架构泛化能力和能效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12258 2026-05-13 cs.LG 85%

Instruction Lens Score: Your Instruction Contributes a Powerful Object Hallucination Detector for Multimodal Large Language Models

指令透镜分数:您的指令为多模态大语言模型提供了一个强大的对象幻觉检测器

Runhe Lai, Xinhua Lu, Yanqi Wu, Jinlun Ye, Weijiang Yu, Ruixuan Wang

机构 * School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China(中山大学计算机科学与工程学院,广州,中国) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国) Key Laboratory of Machine Intelligence and Advanced Computing, MOE, Guangzhou, China(机器智能与高级计算关键实验室,教育部,广州,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract,abstract_cn)

AI总结 本文提出InsLen,通过结合校准局部分数和上下文一致性分数,有效检测多模态大语言模型中的对象幻觉,无需额外训练或辅助模型。

Comments Accepted by ICML-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12034 2026-05-13 cs.CV cs.LG cs.MM 84%

Calibrated Multimodal Representation Learning with Missing Modalities

校准的多模态表示学习与缺失模态

Xiaohao Liu, Xiaobo Xia, Jiaheng Wei, Shuo Yang, Xiu Su, See-Kiong Ng, Tat-Seng Chua

机构 * National University of Singapore(国立新加坡大学) University of Science and Technology of China(中国科学技术大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Central South University(中南大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

AI总结 本文提出CalMRL,通过校准缺失模态导致的不完整对齐问题,从锚点偏移角度提供理论见解,结合先验知识和模态间联系,解决优化难题,验证了方法的有效性。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11468 2026-05-13 cs.AI 83%

CAMPA: Efficient and Aligned Multimodal Graph Learning via Decoupled Propagation and Aggregation

CAMPA: 通过解耦传播和聚合实现高效且对齐的多模态图学习

Daohan Su, Hao Liu, Xunkai Li, Yinlin Zhu, Xiong Yongfu, Yi Liu, Hongchao Qin, Rong-Hua Li, Guoren Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 CAMPA通过解耦传播和聚合机制,解决多模态图学习中的模态冲突问题,提升效率和可扩展性,优于现有基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12064 2026-05-13 cs.CV 79%

TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

TAR:基于文本语义的跨模态图像配准框架用于光学和SAR图像

Zhuoyu Cai, Dou Quan, Ning Huyan, Pei He, Shuang Wang, Licheng Jiao

机构 * Key Laboratory of Intelligent Perception and Image Understanding of Ministry of Education of China, School of Artificial Intelligence, Xidian University(中国教育部智能感知与图像理解重点实验室,西安电子科技大学人工智能学院) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出TAR框架,通过文本语义先验缓解模态差距,提升跨模态特征学习,解决大形变下的光学与SAR图像配准问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11931 2026-05-13 cs.CV 79%

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training

学会思考:通过视觉感知的自我改进训练提升多模态推理

Qihuang Zhong, Liang Ding, Wenjie Xuan, Juhua Liu, Bo Du, Dacheng Tao

机构 * School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence(计算机学院、多媒体软件国家工程研究中心、人工智能研究院) Hubei Key Laboratory of Multimedia(湖北多媒体重点实验室) Network Communication Engineering, Wuhan University, China(网络通信工程、武汉大学,中国) The University of Sydney, Australia(悉尼大学,澳大利亚) Nanyang Technological University, Singapore(南洋理工大学,新加坡)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出VISTA框架,通过视觉感知的自我改进训练提升多模态推理能力,解决数据不平衡和语言先验偏差问题,实验显示在多种训练场景下提升性能。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10500 2026-05-13 cs.CV 79%

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning

视觉增强深度缩放用于多模态潜在推理

Yudong Han, Yong Wang, Zaiquan Yang, Zhen Qu, Liyuan Pan, Xiangxiang Chu

机构 * Beijing Institute of Technology(北京理工大学) AMAP, Alibaba Group(阿里集团AMAP) City University of Hong Kong(香港城市大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Yangtze Delta Region Academy of Beijing Institude of Technology, Jiaxing, China(北京理工大学扬子江地区学院,嘉兴,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出视觉回放模块和路由深度缩放,通过增强视觉感知和细化复杂潜在表示,提升多模态潜在推理的效率与性能。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11716 2026-05-13 cs.AI 79%

SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models

SafeSteer: 多模态大语言模型中的解码级防御机制

Xinyi Zeng, Xue Yang, Jingyuan Zhang, Huanqian Yan, Xiang Chen, Kaiwen Wei, Hankun Kang, Yu Tian

机构 * Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) Kuaishou Technology(快手科技) School of Computer Science and Technology, Beihang University(北航计算机科学与技术学院) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Chongqing University(重庆大学) Wuhan University(武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出SafeSteer,通过解码阶段的轻量探针和模态语义对齐向量,提升多模态大语言模型的安全性,实验表明其能提升33.40%的安全性而不需微调。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11015 2026-05-13 cs.CR cs.AI 79%

DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization

DCVD:双通道跨模态融合用于联合漏洞检测与定位

Wenxin Tang, Wenbin Li, Junliang Liu, Jingyu Xiao, Xi Xiao, Mingzhe Liu, Jinlong Yang, Xuan Liu, Yuehe Ma, Wang Luo, Qing Li, Lei Wang, Peng Xiangli

机构 * Tsinghua University(清华大学) Hunan University(湖南大学) Dalian Maritime University(大连海事大学) The Chinese University of Hong Kong(香港中文大学) Shenzhen University(深圳大学) Northwestern Polytechnical University(西北工业大学) Shandong University(山东大学) BNU-HKBU United International College(北京师范大学-香港浸会大学联合国际学院) Sun Yat-sen University(中山大学) Peng Cheng Laboratory(鹏城实验室) Guangzhou Intelligence Communications Technology Co., Ltd.(广州智能通信技术有限公司) The Fifth Electronic Research Institute of MIIT(中华人民共和国信息产业部第五电子研究所)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出DCVD框架,通过双通道融合实现功能级检测与语句级定位的联合优化,有效解决单一信息源和缺乏显式监督的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11477 2026-05-13 cs.CV 77%

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs

LDDR:基于线性DPP的动态分辨率帧采样用于视频MLLMs

Jingfeng Chen, Jiawen Qian, Wendi Deng, Yinuo Guo, Jiaqi Yu, Sicong Leng, Raghuveer Thirukovalluru, Bhuwan Dhingra

机构 * Carnegie Mellon University(卡内基梅隆大学) Individual Researcher(个人研究员) National University Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Duke University(杜克大学)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 本文提出LDDR框架,通过任务条件特征空间中查询感知的DPP帧选择,在有限视觉token预算下提升视频理解,实现3倍速度提升,并通过组DPP重要性度量动态分配分辨率,优于现有基线方法。

Comments 21 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12056 2026-05-13 cs.AI 70%

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models

OmniRefine: 一种面向对齐的协作压缩方法以提高多模态大语言模型的效率

Yuchen Deng, Zidang Cai, Hai-Tao Zheng, Jie Wang, Feidiao Yang, Yuxing Han

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Pengcheng Laboratory(鹏城实验室)

专题命中 多模态训练与对齐 :cross-modal(abstract);audio-visual(abstract);分类 cs.AI

AI总结 本文提出OmniRefine,一种无需训练的两阶段框架,通过跨模态对齐和协作压缩提升多模态大语言模型的推理效率与性能稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11824 2026-05-13 cs.CV cs.AI 62%

REFNet++: Multi-Task Efficient Fusion of Camera and Radar Sensor Data in Bird's-Eye Polar View

REFNet++:多任务高效融合摄像头和雷达传感器数据的鸟瞰极坐标视图

Kavin Chandrasekaran, Sorin Grigorescu, Gijs Dubbelman, Pavol Jancura

机构 * ElektroBit Automotive GmbH Eindhoven University of Technology(埃因霍温理工大学) Transilvania University of Brasov(布拉索夫特拉扬大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出REFNet++,通过统一域对齐摄像头和雷达数据,实现车辆检测和自由空间分割的高效多模态融合,提升环境感知的准确性和效率。

Comments IEEE Intelligent Transportation Systems Conference (ITSC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11559 2026-05-13 cs.CV cs.AI 62%

When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs

当观察不足时:视觉注意结构揭示大语言模型中的幻觉

Fanpu Cao, Xin Zou, Xuming Hu, Hui Xiong

机构 * Thrust of Artificial Intelligence, HKUST (Guangzhou)(人工智能前沿 thrust,香港科技大学(广州)) Department of Computer Science and Engineering, HKUST(计算机科学与工程系,香港科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文通过分析视觉注意的高频结构揭示大语言模型中的幻觉现象,提出LaSCD解码策略,有效减少幻觉并保持模型能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11840 2026-05-13 cs.CV 57%

Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation

选择,而非融合:雷达-相机状态空间模型用于雷达-相机深度估计

Zhangcheng Hou, Tomoaki Ohtsuki

机构 * School of Science and Technology(科学与技术学部)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出雷达-调制选择(RMS)模型,通过在Mamba的选择性扫描中引入雷达信号,提升雷达-相机深度估计的精度与效率,实现在nuScenes数据集上取得最佳性能。

Comments 16 pages, 3 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11799 2026-05-13 cs.CV 57%

SB-BEVFusion: Enhancing the Robustness against Sensor Malfunction and Corruptions

SB-BEVFusion:增强对传感器故障和损坏的鲁棒性

Markus Essl, Marta Moscati, Mubashir Noman, Muhammad Zaigham Zaheer, Usman Naseem, Shah Nawaz, Markus Schedl

机构 * Johannes Kepler University Linz(约翰·凯撒大学林茨分校) MBZUAI(穆罕默德·本·拉希德人工智能研究所) Macquarie University(麦考瑞大学) Linz Institute of Technology(林茨技术学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出SB-BEVFusion框架,通过处理缺失或损坏的相机和激光雷达数据,提升自动驾驶3D目标检测在传感器故障和损坏场景下的鲁棒性。

Comments Accepted at ICIP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11048 2026-05-13 cs.RO cs.AI 57%

ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching

ForceFlow: 通过接触驱动的流匹配学习感知与行动

Shuoheng Zhang, Yifu Yuan, Hongyao Tang, Yan Zheng, Qiaojun Yu, Pengyi Li, Guowei Huang, Helong Huang, Xingyue Quan, Jianye Hao

机构 * Tianjin University(天津大学) Huawei Noah's Ark Lab(华为诺亚实验室) Shanghai AI Lab(上海人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 本文提出ForceFlow框架,通过融合力信号与多模态信息,提升机器人在复杂接触任务中的鲁棒性和泛化能力,实验显示其在六个真实任务中成功率提升37%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11034 2026-05-13 cs.CR cs.AI cs.LG cs.PF 57%

MambaNetBurst: Direct Byte-level Network Traffic Classification without Tokenization or Pretraining

MambaNetBurst:无需分词或预训练的直接字节级网络流量分类

Gayan K. Kulatilleke, Siamak Layeghy, Mahsa Baktashmotlagh, Marius Portmann

机构 * University of Queensland, Brisbane, Australia(昆士兰大学,布里斯班,澳大利亚)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

AI总结 MambaNetBurst基于Mamba-2架构,直接对原始字节进行分类,无需分词或预训练,实现了高效的网络流量分类。

Comments 16 pages, 2 figures. Pareto-optimal frontier. Transformer vs Mamba vs Mamba-2 scaling performance. Code and data available on request

详情

展开后加载摘要…

URL PDF HTML 收藏