arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45986 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4644 篇

2608.01373 2026-08-04 cs.CR 新提交 78%

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails

《狼来了:多模态安全护栏中安全输入被对抗性误分类为不安全》

Shuo Shi, Rui Yin, Naen Xu, Jiahao Chen, Chunyi Zhou, Tianyu Du, Zhihui Fu, Jun Wang, Zhaoxiang Wang, Shouling Ji

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 该研究针对多模态安全护栏提出不安全诱导攻击,通过不安全语义蒸馏实现84%攻击成功率,揭示了当前多模态安全架构存在安全输入被误判为不安全的可用性漏洞。

Comments Accepted by KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03245 2026-07-07 cs.LG 新提交 78%

PhenoNEST: A Neuro-Symbolic Framework for Ontology-Aware Multimodal Plant Phenotyping and Trait Discovery

PhenoNEST:用于本体感知多模态植物表型分析和性状发现的神经符号框架

Jayant Ghadge, Soumyashree Kar, Surya S. Durbha

机构 * Centre of Studies in Resources Engineering (CSRE)(资源工程研究中心) Indian Institute of Technology Bombay(孟买印度理工学院)

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 提出构建多模态粒状知识图谱框架,以弥合语义鸿沟,通过蒸馏野外记录提取实体和关系,结合模型与本体对齐,经视觉语言模型生成软图,引入节点连接多模态子图,成功映射野外观察,助力小麦育种。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01079 2026-07-02 cs.RO 新提交 78%

Where Am I? Semantic Map Grounding via Vision-Language Models for Multi-Modal Localization

我在哪里?基于视觉-语言模型的语义地图定位用于多模态定位

Suraj Borate, Aarav Shah, Madhu Vadali

机构 * IIT Gandhinagar(印度理工学院甘地讷格尔分校)

专题命中 图文多模态 :multi-modal(title);cross-modal(abstract)

AI总结 将机器人定位重构为语义推理任务,使用微调的Qwen2.5-VL-7B模型融合摄像头、LiDAR和语义地图预测位姿,在室内数据集上达到98.23%位置精度,并展现出跨模态互补性和对新场景的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10742 2026-06-10 cs.CR cs.LG 新提交 78%

MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents

MemVenom:网络代理中多模态记忆的触发式投毒

Yv Zhang, Hao Sun, Hao Fang, Kuofeng Gao, Fan Mo, Bin Chen, Shu-Tao Xia, Yaowei Wang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Peng Cheng Laboratory(鹏城实验室) Tianjin University(天津大学) Shenzhen International Graduate School, Tsinghua University(深圳国际研究生学院,清华大学) Huawei Technologies Ltd.(华为技术有限公司)

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 提出MemVenom框架,针对网络代理的外部记忆系统,通过触发条件检索和攻击诱导,实现黑盒多模态记忆投毒,达到高成功率且不影响良性性能。

Comments Preprint. 27 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09017 2026-04-13 cs.NI 78%

Multimodal Large Language Model Enabled Robust Beamforming for HAP Downlink Communications

多模态大语言模型赋能的HAP下行链路鲁棒波束成形

Xiaoyu Xing, Peng Yang, Guoquan Tao, Dingyi Lu, Zehui Xiong, Xianbin Cao

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 本文提出基于多模态大语言模型的波束成形框架,通过飞行 telemetry 预测HAP姿态并实现鲁棒下行通信,提升用户服务比和总速率22.1%和12.5%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14745 2026-03-17 cs.LG 78%

CAMD: Coverage-Aware Multimodal Decoding for Efficient Reasoning of Multimodal Large Language Models

CAMD:面向多模态大语言模型高效推理的覆盖感知多模态解码

Huijie Guo, Jingyao Wang, Lingyu Si, Jiahuan Zhou, Changwen Zheng, Wenwen Qiang

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 本文提出CAMD,通过动态分配计算资源提升多模态大语言模型的推理效率与可靠性,解决解码方法在易例和难例间的计算不匹配问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06084 2026-03-09 cs.RO 78%

Multimodal Behavior Tree Generation: A Small Vision-Language Model for Robot Task Planning

多模态行为树生成:一种小型视觉-语言模型用于机器人任务规划

Cristiano Battistini, Riccardo Andrea Izzo, Gianluca Bardaro, Matteo Matteucci

机构 * Department of Electronics, Information, and Bioengineering, Politecnico di Milano(电子信息与生物工程系,米兰理工学院)

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 本文提出一种小型视觉-语言模型,通过生成行为树提升机器人任务规划效率,实验证明其在性能和资源消耗上均优于现有闭源模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00854 2026-03-03 cs.LG cs.IR 78%

GeMi: A Graph-based, Multimodal Recommendation System for Narrative Scroll Paintings

GeMi:基于图的多模态推荐系统用于叙事卷轴绘画

Haimonti Dutta, Pruthvi Moluguri, Jin Dai, Saurabh Amarnath Mahindre

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 GeMi是一种基于图的多模态推荐系统,专门用于叙事卷轴绘画,结合多模态内容和用户偏好,为艺术保护和数据存储提供解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15543 2026-02-18 cs.RO 78%

Selective Perception for Robot: Task-Aware Attention in Multimodal VLA

机器人选择性感知:多模态VLA中的任务感知注意力

Young-Chae Son, Jung-Woo Lee, Yoon-Ji Choi, Dae-Kwan Ko, Soo-Chul Lim

机构 * Dongguk University(东国大学)

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 本文提出了一种动态信息融合框架,通过任务感知注意力提升机器人多模态VLA模型的效率与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00030 2026-02-10 cs.LG 78%

RAPTOR-AI for Disaster OODA Loop: Hierarchical Multimodal RAG with Experience-Driven Agentic Decision-Making

RAPTOR-AI用于灾难OODA循环:基于经验驱动的多模态RAG框架

Takato Yasuno

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 RAPTOR-AI通过分层多模态RAG框架,结合经验驱动的代理控制和LoRA适应,提升灾害响应中的检索精度、情境基础性和任务分解准确性。

Comments 8 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08959 2026-01-15 cs.CR cs.SE 78%

Integrating APK Image and Text Data for Enhanced Threat Detection: A Multimodal Deep Learning Approach to Android Malware

整合APK图像和文本数据以增强威胁检测:一种多模态深度学习方法用于Android恶意软件

Md Mashrur Arifin, Maqsudur Rahman, Nasir U. Eisty

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 本文提出一种多模态深度学习方法,通过整合APK图像和文本数据提升Android恶意软件检测效果,系统评估不同图像类型和分辨率,并利用CLIP模型增强分析能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00961 2026-01-14 cs.DL cs.IR 78%

Digital Collections Explorer: An Open-Source, Multimodal Viewer for Searching Digital Collections

数字收藏探索器:一种开源的多模态查看器,用于搜索数字收藏

Ying-Hsiang Huang, Benjamin Charles Germain Lee

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 Digital Collections Explorer是一种开源多模态查看器,利用CLIP技术提升数字收藏的视觉搜索能力,并通过案例研究展示其在文化遗产档案中的应用。

Comments 14 pages, 8 figures, 2 tables

Journal ref Comput. humanit. res. 1 (2025) e14

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15607 2025-12-02 cs.RO 78%

PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation Models

基于偏好强化学习的多模态反馈与轨迹合成:从基础模型出发

Ruiqi Wang, Dezhong Zhao, Ziqin Yuan, Tianyu Shao, Guohua Chen, Dominic Kao, Sungeun Hong, Byung-Cheol Min

机构 * Purdue University(普渡大学) Beijing University of Chemical Technology(北京化工大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Sungkyunkwan University(成均馆大学) Indiana University Bloomington(印第安纳大学布卢明顿分校)

专题命中 图文多模态 :multimodal(title,abstract)

AI总结 PRIMT通过多模态反馈和轨迹合成,利用基础模型解决偏好强化学习中的查询模糊性和信用分配问题,提升机器人复杂行为的学习效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13733 2025-11-26 cs.RO 78%

FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph

基于分层多模态场景图的视觉语言导航:快速与慢速推理

Xiaolin Zhou, Tingyang Xiao, Liu Liu, Yucheng Wang, Maiyue Chen, Xinrui Meng, Xinjie Wang, Wei Feng, Wei Sui, Zhizhong Su

机构 * Horizon Robotics D-Robotics Robotics

专题命中 图文多模态 :multi-modal(title,abstract)

AI总结 FSR-VLN通过分层多模态场景图与快速-慢速推理机制,提升视觉语言导航的效率与准确性,实现低延迟高成功率的导航性能。

Comments Demo video are available at https://horizonrobotics.github.io/robot_lab/fsr-vln/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02113 2025-11-11 cs.IR 78%

Enhancing Multimodal Recommendations with Vision-Language Models and Information-Aware Fusion

Hai-Dang Kieu, Min Xu, Thanh Trung Huynh, Dung D. Le

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11350 2025-11-10 cs.RO 78%

Search-TTA: A Multimodal Test-Time Adaptation Framework for Visual Search in the Wild

Derek Ming Siang Tan, Shailesh, Boyang Liu, Alok Raj, Qi Xuan Ang, Weiheng Dai, Tanishq Duhan, Jimmy Chiun, Yuhong Cao, Florian Shkurti, Guillaume Sartoretti

专题命中 图文多模态 :multimodal(title,abstract)

Comments Accepted for presentation at CORL 2025. Code, models, and data are available at https://search-tta.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19294 2025-10-01 cs.CV cs.AI cs.CL 78%

Object Detection with Multimodal Large Vision-Language Models: An In-depth Review

Ranjan Sapkota, Manoj Karkee

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments First Peer Reviewed Review Paper for Object Detection with Vision-Language Models (VLMs)

Journal ref Information Fusion, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23281 2025-09-30 cs.RO 78%

Preventing Robotic Jailbreaking via Multimodal Domain Adaptation

Francesco Marchiori, Rohan Sinha, Christopher Agia, Alexander Robey, George J. Pappas, Mauro Conti, Marco Pavone

机构 * University of Padova(帕多瓦大学) Stanford University(斯坦福大学) Carnegie Mellon University(卡内基梅隆大学) University of Pennsylvania(宾夕法尼亚大学) Örebro University(奥雷布罗大学) NVIDIA Research(NVIDIA研究)

专题命中 图文多模态 :multimodal(title,abstract)

Comments Project page: https://j-dapt.github.io/. 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19480 2025-09-25 cs.RO cs.LG 78%

OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation

Noriaki Hirose, Catherine Glossop, Dhruv Shah, Sergey Levine

专题命中 图文多模态 :omni-modal(title,abstract)

Comments 9 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18729 2025-09-24 cs.SD 78%

MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning

Haoqin Sun, Chenyang Lyu, Xiangyu Kong, Shiwan Zhao, Jiaming Zhou, Hui Wang, Aobo Kong, Jinghua Zhao, Longyue Wang, Weihua Luo, Kaifu Zhang, Yong Qin

机构 * TMCC, College of Computer Science, Nankai University, Tianjin, China(TMCC,计算机科学学院,南开大学,天津,中国) Alibaba International Digital Commerce(阿里巴巴国际数字商务) University of Exeter(埃克塞特大学)

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00449 2025-09-03 cond-mat.mtrl-sci cond-mat.dis-nn cs.LG 78%

Molecular Identification from AFM images using the IUPAC Nomenclature and Attribute Multimodal Recurrent Neural Networks

Jaime Carracedo-Cosme, Carlos Romero-Muñiz, Pablo Pou, Rubén Pérez

专题命中 图文多模态 :multimodal(title,abstract)

Comments 30 pages, 4 figures, 2 tables, includes supplementary information (with additional 21 pages, 9 figures, 1 table)

Journal ref ACS Appl. Mater. Interfaces 15, 22692-22704 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00302 2025-07-22 cs.LG cond-mat.mtrl-sci 78%

Beyond Atomic Geometry Representations in Materials Science: A Human-in-the-Loop Multimodal Framework

Can Polat, Erchin Serpedin, Mustafa Kurban, Hasan Kurban

机构 * Computer Engineering, Texas A\&M University, College Station, TX 77843, USA College of Science Engineering, Hamad Bin Khalifa University, Doha, Qatar Dept. of Electrical \& Computer Engineering, Texas A\&M University at Qatar, Doha, Qatar Orthotics, Ankara University, Ankara, Turkey

专题命中 图文多模态 :multimodal(title,abstract)

Comments Presented at ICML 2025 Workshop on DataWorld

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13624 2025-04-21 eess.SP 78%

PV-VLM: A Multimodal Vision-Language Approach Incorporating Sky Images for Intra-Hour Photovoltaic Power Forecasting

Huapeng Lin, Miao Yu

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03153 2025-04-07 cs.LG 78%

MORAL: A Multimodal Reinforcement Learning Framework for Decision Making in Autonomous Laboratories

Natalie Tirabassi, Sathish A. P. Kumar, Sumit Jha, Arvind Ramanathan

专题命中 图文多模态 :multimodal(title,abstract)

Comments 9 pages, 14 figures and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13164 2025-04-02 cs.LG 78%

VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning

Yongshuo Zong, Ondrej Bohdal, Timothy Hospedales

专题命中 图文多模态 :multimodal(title,abstract)

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08317 2025-03-24 cs.CL cs.AI cs.CV 78%

Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting

Jiarui Wu, Zhuo Liu, Hangfeng He

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00196 2025-03-05 cs.CV cs.AI cs.CL 78%

DermaSynth: Rich Synthetic Image-Text Pairs Using Open Access Dermatology Datasets

Abdurrahim Yilmaz, Furkan Yuceyalcin, Ece Gokyayla, Donghee Choi, Ozan Erdem, Ali Anil Demircali, Rahmetullah Varol, Ufuk Gorkem Kirabali, Gulsum Gencoglan, Joram M. Posma, Burak Temelkuran

专题命中 图文多模态 :image-text(title);分类 cs.CV、cs.CL、cs.AI

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12662 2025-03-03 cs.CV cs.AI cs.CL 78%

Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models

Shicheng Xu, Liang Pang, Yunchang Zhu, Huawei Shen, Xueqi Cheng

专题命中 图文多模态 :cross-modal(title);分类 cs.CV、cs.CL、cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13904 2025-02-14 cs.LG 78%

Privacy-Preserving Personalized Federated Prompt Learning for Multimodal Large Language Models

Linh Tran, Wei Sun, Stacy Patterson, Ana Milanova

专题命中 图文多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10967 2025-02-13 cs.CV cs.AI cs.CL 78%

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding

Zhanpeng Chen, Mingxiao Li, Ziyang Chen, Nan Du, Xiaolong Li, Yuexian Zou

专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏