arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-05-08 至 2026-05-08 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2602.17419 2026-05-08 cs.CV 89%

EAGLE: Expert-Augmented Attention Guidance for Tuning-Free Industrial Anomaly Detection in Multimodal Large Language Models

EAGLE:专家增强的注意力引导用于多模态大语言模型中的无调优工业异常检测

Xiaomeng Peng, Xilang Huang, Seon Han Choi

机构 * Ewha Womans University(峨山女子大学)

专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract);multimodal(title,abstract);分类 cs.CV

AI总结 EAGLE通过整合专家异常检测器与冻结的MLLM,提出无调优框架,提升多模态大语言模型在工业异常检测中的准确率,且在MVTec-AD和VisA数据集上达到94.4%和88.1%的异常鉴别准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17980 2026-05-08 cs.CV 86%

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding

感知空间:面向高效准确3D场景理解的自我运动感知视频表示

Shuyao Shi, Kang G. Shin

机构 * Department of Computer Science(计算机科学系) University of Michigan(密歇根大学) University of Michigan Ann Arbor(密歇根大学安阿伯分校)

专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出Motion-MLLM框架,结合IMU数据与视觉特征,通过运动-视觉关键帧过滤模块和异构跨模态融合模块,提升3D场景理解与空间推理的效率和准确性。

Comments 22 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06238 2026-05-08 cs.LG cs.AI 83%

Band Together: Untargeted Adversarial Training with Multimodal Coordination against Evasion-based Promotion Attacks

Band Together: 多模态协调的无目标对抗训练对抗基于逃避的推广攻击

Guanmeng Xian, Ning Yang, Philip S. Yu

机构 * Sichuan University(四川大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出UAT-MC方法,通过多模态协调解决逃避攻击中未知目标项的问题,提升系统鲁棒性并保持推荐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05694 2026-05-08 cs.CV 83%

Adaptive Physical-Facial Representation Fusion via Subject-Invariant Cross-Modal Prompt Tuning for Video-Based Emotion Recognition

基于主体不变跨模态提示调优的自适应物理-面部表征融合用于基于视频的情感识别

Xiwen Luo, Jia Li, Rencheng Song, Yu Liu, Juan Cheng

机构 * Department of Biomedical Engineering(生物医学工程系) Anhui Province Key Laboratory of Measuring Theory and Precision Instrument(安徽省测量理论与精密仪器重点实验室) School of Computer Science and Information Engineering(计算机科学与信息工程学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出一种主体不变的跨模态提示调优框架,通过将rPPG波形转换为噪声鲁棒的时间-频率表示,并引入解耦共享-特定适配器以提升跨主体泛化能力,实验证明在MAHNOB-HCI和DEAP基准上优于现有方法。

Comments The source code will be available at https://github.com/MSA-LMC/SCPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06460 2026-05-08 cs.LG 82%

MINER: Mining Multimodal Internal Representation for Efficient Retrieval

MINER:挖掘多模态内部表示以实现高效检索

Weien Li, Rui Song, Zeyu Li, Haochen Liu, Gonghao Zhang, Difan Jiao, Zhenwei Tang, Bowei He, Haolun Wu, Xue Liu, Ye Yuan

机构 * McGill University(麦吉尔大学) MIT - Massachusetts Institute of Technology(麻省理工学院) University of Cambridge(剑桥大学) University of Toronto(多伦多大学) MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(MBZUAI - 摩洛哥 bin Zayed 大学人工智能学院) Mila - Quebec AI Institute(Mila - 加拿大魁北克人工智能研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文提出MINER,通过挖掘Transformer各层内部表示,融合多模态信号生成紧凑嵌入,提升检索性能,优于现有单向量检索器并在部分设置中缩小与晚交互基线的差距。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06073 2026-05-08 cs.LG 82%

PRISM: Iterative Cross-Modal Posterior Refinement for Dynamic Text-Attributed Graphs

PRISM:动态文本属性图的迭代跨模态后验细化

Trimble Chang, Yihang Liu, Mingjing Han, Han Zhang

机构 * College of Artificial Intelligence(人工智能学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract)

AI总结 PRISM通过迭代跨模态后验细化方法,提升动态文本属性图的表示学习能力,有效捕捉节点语义与交互行为的动态依赖关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19316 2026-05-08 cs.CL 79%

KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls

KORE:通过知识导向控制增强大型多模态模型的知识注入

Kailin Jiang, Hongbo Jiang, Ning Jiang, Zhi Gao, Jinhe Bi, Yuchen Ren, Bin Li, Yuntao Du, Lei Liu, Qing Li

机构 * University of Science State Key Laboratory of General Artificial Intelligence, BIGAI Xiamen University Northeast Forestry University Beijing Institute of Technology Ludwig Maximilian University of Munich The University of Sydney C-FAIR\&school of software, Shandong University State Key Lab. for Novel Software Technology, Nanjing University, P.R. China

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

AI总结 KORE通过知识导向的增强和约束,提升大型多模态模型的知识注入能力,同时保留旧知识。方法利用协方差矩阵和投影初始化,有效减少灾难性遗忘。

Comments ICML 2026, Project Page: https://kore-lmm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05850 2026-05-08 cs.CV 79%

Align3D-AD: Cross-Modal Feature Alignment and Dual-Prompt Learning for Zero-shot 3D Anomaly Detection

Align3D-AD:跨模态特征对齐与双提示学习用于零样本3D异常检测

Letian Bai, Xuanming Cao, Juan Du, Chengyu Tao

机构 * Smart Manufacturing Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)智能制造方向) The Hong Kong University of Science and Technology(香港科技大学) College of Mechanical and Vehicle Engineering, Hunan University(湖南大学机械与车辆工程学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出Align3D-AD框架,通过跨模态特征对齐和双提示学习解决零样本3D异常检测中的领域差距问题,实验表明其在多个数据集上均优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01327 2026-05-08 cs.AI cs.LG 79%

Segment-Aligned Policy Optimization for Multi-Modal Reasoning

基于段落对齐的策略优化用于多模态推理

Lei Gao, Zhuoming Li, Mengxi Jia, Jiakang Yuan, Hongbo Sun, Hao Sun, Xuelong Li

机构 * Fudan University(复旦大学) Southeast University(东南大学) China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd.(中国电信人工智能技术(北京)有限公司) Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出SAPO方法,通过将推理步骤而非token或完整序列作为策略更新的基本单元,提升多模态推理任务的准确性和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00699 2026-05-08 cs.CR 78%

STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack

STARE:分步时间对齐与红队引擎用于多模态毒性攻击

Xutao Mao, Liangjie Zhao, Tao Liu, Xiang Zheng, Hongying Zan, Cong Wang

专题命中 多模态训练与对齐 :multi-modal(title);image-text(abstract)

AI总结 STARE通过分步时间对齐和红队引擎,提升多模态毒性攻击的成功率,揭示优化诱导的相位对齐现象,为安全机制提供理论基础。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22991 2026-05-08 cs.LG 78%

Fusion or Confusion? Multimodal Complexity Is Not All You Need

融合还是混淆?多模态复杂性并不都是你需要的

Tillmann Rheude, Roland Eils, Benjamin Wild

机构 * Berlin Institute of Health, Charité - Universitätsmedizin Berlin(柏林健康研究所,柏林查理医院) Intelligent Medicine Institute, Fudan University(复旦大学智能医学研究院) Department of Mathematics and Computer Science, Freie Universität Berlin(柏林自由大学数学与计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

AI总结 本文通过大规模实验挑战多模态学习中复杂架构提升性能的假设,发现增加复杂性常导致混淆而非有效融合,强调需转向方法论严谨性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05709 2026-05-08 cs.AI 70%

Conceal, Reconstruct, Jailbreak: Exploiting the Reconstruction-Concealment Tradeoff in MLLMs

隐藏、重建、突破:在大规模语言模型中利用重建-隐藏权衡

Md Farhamdur Reza, Richeng Jin, Tianfu Wu, Huaiyu Dai

机构 * NC State University(北卡罗来纳州立大学) Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract_cn);分类 cs.AI

AI总结 本文探讨了在多模态大语言模型中利用重建与隐藏的权衡进行意图混淆攻击,提出了一种基于字符移除的变体构造方法,并引入关键词相关的干扰图像以提高攻击效果。

Comments 39 pages, including appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18738 2026-05-08 cs.CL 57%

Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Large Language Models

重新标记,而非替换:扩散大语言模型中的令牌到标记细化

Lin Yao

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) Zhongguancun Academy(中关村学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

AI总结 本文提出Token-to-Mask(T2M)方法,通过重新标记可疑令牌而非覆盖来提升扩散大语言模型的准确性,在AIME 2025和CMATH上分别提升13.33和8.56个百分点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01746 2026-05-08 cs.CV 57%

Point-SRA: Self-Representation Alignment for 3D Representation Learning

点-SRA:用于3D表示学习的自表示对齐

Lintong Wei, Jian Lu, Haozhe Cheng, Jihua Zhu, Kaibing Zhang

机构 * School of Electronics and Information, Xi’an Polytechnic University(西安理工大学电子与信息学院) School of Software, Xi’an Jiaotong University(西安交通大学软件学院) School of Computer Science, Xi’an Polytechnic University(西安理工大学计算机科学学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 Point-SRA通过自蒸馏和概率建模对齐表示,改进3D表示学习,通过不同掩码比例和MeanFlow Transformer实现互补信息提取,优于Point-MAE并在多个任务中取得优异性能。

Comments This is an AAAI 2026 accepted paper titled "Point-SRA: Self-Representation Alignment for 3D Representation Learning", spanning 13 pages in total. The submission includes 7 figures (fig1 to fig7) that visually support the technical analysis

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 2026, Vol. 40, No. 13

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06126 2026-05-08 cs.HC 50%

AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition

AffectGPT-RL:揭示强化学习在开放词汇情绪识别中的作用

Zheng Lian, Fan Zhang, Lan Chen, Yazhou Zhang, Rui Liu, Jinyang Wu, Haoyu Chen, Xiaobai Li, Xiaojiang Peng, Bin He, Jianhua Tao

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 本文提出AffectGPT-RL框架,通过强化学习优化非可微目标,提升开放词汇情绪识别的性能,并在情绪识别和其他任务中取得显著成果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05899 2026-05-08 cs.LG 50%

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading

VisMMOE:利用视觉专家亲和力实现高效的视觉语言MoE卸载

Cheng Xu, Xiaofeng Hou, Jiacheng Liu, Chao Li

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

AI总结 VisMMOE通过剪枝冗余视觉token提升视觉语言MoE模型的卸载效率,通过视觉专家亲和力机制优化专家访问集中度与稳定性,从而提升推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏