arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-04-22 至 2026-04-22 共收录 74 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 7 篇

2604.19148 2026-04-22 cs.RO 50%

Multi-Step Gaussian Process Propagation for Adaptive Path Planning

多步高斯过程传播用于自适应路径规划

Alex Beaudin, Bjørn Andreas Kristiansen, Kristoffer Gryte, Corrado Chiatante, Morten Omholt Alver, Murat Arcak, Tor Arne Johansen

机构 * Department of Electrical Engineering and Computer Sciences, University of California, Berkeley(加州大学伯克利分校电子工程与计算机科学系) Department of Engineering Cybernetics, Norwegian University of Science and Technology (NTNU)(挪威科技与自然大学(NTNU)工程控制系) Department of Electronic Systems, Norwegian University of Science and Technology (NTNU)(挪威科技与自然大学(NTNU)电子系统系)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出基于高斯过程的多步路径规划方法OLAhGP,通过融合多模态环境传感数据和状态输入约束,提升自主水面舰在藻类监测中的路径规划效率与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态训练与对齐 11 篇

2507.09861 2026-04-22 cs.CV cs.AI 91%

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends

基于MLLM的视觉丰富文档理解综述:方法、挑战与新兴趋势

Yihao Ding, Siwen Luo, Yue Dai, Yanbei Jiang, Zechuan Li, Qiang Sun, Geoffrey Martin, Wei Liu, Yifan Peng

机构 * The University of Western Australia(西澳大学) The University of Melbourne(墨尔本大学) Weill Cornell Medicine(韦尔·柯尔医学中心)

专题命中 多模态训练与对齐 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文综述了基于MLLM的视觉丰富文档理解最新进展,探讨了文本、视觉和布局特征的表示与整合技术,以及预训练、指令微调等训练方法,分析了数据稀缺、多页文档处理等挑战及新兴趋势。

Comments Accepted at ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19379 2026-04-22 cs.CV 83%

PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving

PanDA: 无监督领域自适应用于自动驾驶中的多模态3D全景分割

Yining Pan, Shijie Li, Yuchen Wu, Xulei Yang, Na Zhao

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Institute for Infocomm Research (I2R), A*STAR, Singapore(新加坡资讯通信研究院(I2R),A*STAR,新加坡)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出PanDA框架,针对多模态3D全景分割的无监督领域自适应问题,通过不对称多模态增强和双专家伪标签细化模块提升鲁棒性和伪标签完整性,实验表明在多种领域转移场景下超越现有SOTA方法。

Comments Accepted at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19083 2026-04-22 cs.CR cs.AI 83%

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety

ProjLens: 揭示项目器在多模态模型安全中的作用

Kun Wang, Cheng Qian, Miao Yu, Lilan Peng, Liang Lin, Jiaming Zhang, Tianyu Zhang, Yu Cheng, Yang Wang

机构 * University of Science and Technology of China(中国科学技术大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学) Nanyang Technological University(南洋理工大学) Southwest Jiaotong University(西南交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 ProjLens通过分析多模态大语言模型中的后门攻击机制,揭示了项目器在安全漏洞中的关键作用,发现后门注入参数编码于低秩子空间,并通过实验验证了激活机制的差异。

Comments 18 pages ,15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08564 2026-04-22 cs.AI cs.CV cs.LG 81%

How to Teach Large Multimodal Models New Skills

如何向大型多模态模型传授新技能

Zhen Zhu, Yiming Gong, Yao Xiao, Yaoyao Liu, Derek Hoiem

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Google DeepMind(谷歌DeepMind) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文研究了如何在不丧失原有能力的前提下通过序列微调向大型多模态模型传授新技能,提出两种有效的微调方法以平衡学习与遗忘。

Comments In submission. Code is available at https://github.com/jessemelpolio/LMM_CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16161 2026-04-22 cs.CV cs.CL 81%

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

OmniParser V2:结构化思维点用于统一的视觉文本解析及其在多模态大语言模型中的通用性

Wenwen Yu, Zhibo Yang, Jianqiang Wan, Sibo Song, Jun Tang, Wenqing Cheng, Yuliang Liu, Xiang Bai

机构 * School of Information Science and Engineering, East China University of Science and Technology(东华大学信息科学与工程学院) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出OmniParser V2,通过结构化思维点提示方案统一视觉文本解析任务,简化流程并提升性能,在多个数据集上取得最佳结果,并验证其在多模态大语言模型中的通用性。

Comments Accepted by IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19383 2026-04-22 cond-mat.mtrl-sci cs.AI 79%

Multimodal Transformer for Sample-Aware Prediction of Metal-Organic Framework Properties

多模态Transformer用于金属有机框架性质的样本感知预测

Seunghee Han, Jaewoong Lee, Jihan Kim

机构 * Department of Chemical and Biomolecular Engineering, Korea Advanced Institute of Science and Technology(化学与生物分子工程系,韩国科学技术院) Department of Materials Science and Engineering, Korea Advanced Institute of Science and Technology(材料科学与工程系,韩国科学技术院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出EXIT模型,结合MOFid与X射线衍射数据,通过预训练和微调提升对MOF表面面积和孔体积的预测性能,强调实验表征在多孔材料信息学中的价值。

Comments 22 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19264 2026-04-22 cs.CV 79%

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents

DR-MMSearchAgent:多模态搜索代理中的深度推理

Shengqin Wang, Wentao Yan, Huichi Zhou, Yihang Chen, Kun Shao, Zhizhong Zhang, Yuan Xie

机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家) School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家) East China Normal University(东华大学) Shanghai Innovation Institute(上海创新研究院) University College London(伦敦大学学院) Independent Researcher(独立研究者) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出DR-MMSearchAgent框架,通过结构接近性从完整批处理轨迹中提取优势信号,提升不同长度轨迹的生成能力,并采用差异化高斯奖励动态校准交互容忍度,以提高信息可靠性并减少冗余。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18790 2026-04-22 cs.CV 74%

EfficientPENet: Real-Time Depth Completion from Sparse LiDAR via Lightweight Multi-Modal Fusion

EfficientPENet:通过轻量多模态融合实现稀疏LiDAR的实时深度补全

Johny J. Lopez, Md Meftahul Ferdaus, Mahdi Abdelguerfi, Anton Netchaev, Steven Sloan, Ken Pathak, Kendall N. Niles

机构 * Canizaro Livingston Gulf States Center for Environmental Informatics, the University of New Orleans, New Orleans, USA(Canizaro Livingston Gulf States环境信息中心,新奥尔良大学,美国新奥尔良) US Army Corps of Engineers, Engineer Research and Development Center, Vicksburg, Mississippi, USA(美国陆军工程兵团,工程师研究与发展中心,密西西比州维克斯堡,美国)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV

AI总结 本文提出EfficientPENet,通过轻量级多模态融合网络,在稀疏LiDAR和RGB图像上实现实时深度补全,具有更少的参数和更高的速度,同时保持竞争力的精度。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18713 2026-04-22 cs.CV 70%

Align then Refine: Text-Guided 3D Prostate Lesion Segmentation

对齐后再细化:基于文本的3D前列腺病变分割

Cuiling Sun, Linkai Peng, Adam Murphy, Elif Keles, Hiten D. Patel, Ashley Ross, Frank Miller, Baris Turkbey, Andrea Mia Bejar, Halil Ertugrul Aktas, Gorkem Durak, Ulas Bagci

机构 * Department of Radiology, Northwestern University, Chicago, USA(放射科,西北大学,芝加哥,美国) Department of Urology, Northwestern University, Chicago, USA(泌尿科,西北大学,芝加哥,美国) Center for Cancer Research, National Cancer Institute, Bethesda, USA(癌症研究中心,国家癌症研究所,贝塞斯达,美国)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出一种多编码器U-Net架构,通过引入对齐损失、热图损失和置信度门控多头交叉注意力细化模块,提升前列腺病变分割的精度和多模态融合能力。

Comments Accepted to EMBC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19631 2026-04-22 cs.CV 57%

MOSA: Motion-Guided Semantic Alignment for Dynamic Scene Graph Generation

MOSA:基于运动的语义对齐用于动态场景图生成

Xuejiao Wang, Bohao Zhang, Changbo Wang, Gaoqi He

机构 * School of Computer Science and Technology, East China Normal University, Shanghai, China(东华大学计算机科学与技术学院) School of Data Science and Engineering, East China Normal University, Shanghai, China(东华大学数据科学与工程学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出MOSA方法,通过融合运动特征与空间关系,提升动态场景图生成中细粒度关系建模和语义表示利用能力,优化尾部关系学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03337 2026-04-22 cs.CV 57%

Less is More: Token-Efficient Video-QA via Adaptive Frame-Pruning and Semantic Graph Integration

少即是多:通过自适应帧剪枝和语义图集成实现高效的视频问答

Shaoguang Wang, Weiyu Guo, Ziyang Chen, Yijie Xu, Xuming Hu, Hui Xiong

机构 * The Hong Kong University of Science and Technology (Guangzhou), China(香港科技大学(广州)中国) The Hong Kong University of Science and Technology, Hong Kong SAR, China(香港科技大学,香港特别行政区,中国)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

AI总结 本文提出一种结合自适应帧剪枝和轻量级语义图的框架,有效减少视频问答中输入token数量,提升效率并增强基模型性能。

Comments Accepted to CVPR 2026 Findings

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他多模态 2 篇

2604.19489 2026-04-22 cs.CV cs.CY 79%

Seeing Candidates at Scale: Multimodal LLMs for Visual Political Communication on Instagram

大规模识别候选人:用于推特视觉政治沟通的多模态大语言模型

Michael Achmann-Denkler, Mario Haim, Christian Wolff

机构 * Media Informatics Group University of Regensburg(媒体信息学组莱比锡大学) Department of Media and Communication Ludwig-Maximilians-Universität Munich, Germany(媒体与传播系路德维希-马克西米利安大学慕尼黑,德国)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文评估了专用机器学习模型和新兴多模态大语言模型在视觉政治沟通分析中的能力,展示了GPT-4o在识别领先政治人物和计数中的优越性能。

Comments An earlier version was presented at #SMSociety 2024 (London)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09370 2026-04-22 cs.RO 50%

Phase-Aware Policy Learning for Skateboard Riding of Quadruped Robots via Feature-wise Linear Modulation

面向四足机器人滑板骑行的相位感知策略学习

Minsung Yoon, Jeil Jeong, Sung-Eui Yoon

专题命中 其他多模态 :multi-modal(abstract)

AI总结 本文提出相位感知策略学习框架,通过特征级线性调制层提升四足机器人滑板骑行的相位依赖行为捕捉与跨相位知识共享,验证了在仿真中的指令跟踪精度和效率。

Comments ICRA 2026 | Project Page: https://minsungyoon.github.io/projects/papl/ | M. Yoon and J. Jeong contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏