arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45986 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4644 篇

2601.08420 2026-01-14 cs.CV 79%

MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP

MMLGNet: 利用CLIP实现遥感数据的跨模态对齐

Aditya Chaudhary, Sneha Barman, Mainak Singha, Ankit Jha, Girish Mishra, Biplab Banerjee

专题命中 图文多模态 :cross-modal(title);multimodal(abstract);分类 cs.CV

AI总结 MMLGNet通过CLIP实现遥感数据的跨模态对齐,利用多模态语言引导网络有效融合光谱、空间和几何信息,提升语义理解能力。

Comments Accepted at InGARSS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06424 2026-01-13 cs.CL 79%

Can a Unimodal Language Agent Provide Preferences to Tune a Multimodal Vision-Language Model?

单模语言代理能否为调整多模态视觉-语言模型提供偏好?

Sazia Tabasum Mim, Jack Morris, Manish Dhakal, Yanming Xiu, Maria Gorlatova, Yi Ding

机构 * Georgia State University(佐治亚州立大学) Duke University(杜克大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出通过单模语言代理提供偏好反馈,提升多模态视觉-语言模型的描述能力,实验显示在准确率上提升13%,且人类偏好匹配率达64.6%。

Comments Accepted to IJCNLP-AACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09818 2026-01-08 cs.CV 79%

ViMoNet: A Multimodal Vision-Language Framework for Human Behavior Understanding from Motion and Video

ViMoNet: 一种多模态视觉-语言框架,用于从运动和视频中理解人类行为

Rajan Das Gupta, Lei Wei, Md Yeasin Rahat, Nafiz Fahad, Abir Ahmed, Liew Tze Hui

机构 * Department of Computer Science, American International University–Bangladesh (AIUB)(美国国际大学-孟加拉国计算机科学系) Faculty of Psychology, Shinawatra University(信武大学心理学系) Faculty of Information Science and Technology, Multimedia University(多媒体大学信息科学与技术系) Department of Information Technology, Washington University of Science & Technology(华盛顿科学与技术大学信息科技系)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 ViMoNet通过整合运动和视频数据,提出多模态视觉-语言框架,有效提升人类行为理解与医疗保健应用潜力。

Comments This is the preprint version of the manuscript. It is currently being prepared for submission to an academic conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02565 2026-01-06 cs.CV 79%

SJTU:Spatial judgments in multimodal models towards unified segmentation through coordinate detection

上海交通大学:通过坐标检测实现多模态模型中的空间判断以实现统一分割

Joongwon Chae, Zhenyu Wang, Peiwu Qin

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 SJTU通过坐标检测实现多模态模型中的空间判断,提升图像分割的准确性和实用性。

Comments A flaw was discovered in the experimental setup. Therefore, we are retracting the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23573 2025-12-30 cs.CV 79%

ProGuard: Towards Proactive Multimodal Safeguard

ProGuard:迈向主动多模态安全防护

Shaohan Yu, Lijun Li, Chenyang Si, Lu Sheng, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) PRLab Nanjing University(南京大学PRLab) Beihang University(北航)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 ProGuard通过强化学习和同义词库奖励机制,实现主动多模态安全防护,显著提升分布外风险检测与描述能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23557 2025-12-30 cs.CR cs.AI 79%

Toward Trustworthy Agentic AI: A Multimodal Framework for Preventing Prompt Injection Attacks

迈向可信的代理AI:一种多模态框架用于防止提示注入攻击

Toqeer Ali Syed, Mishal Ateeq Almutairi, Mahmoud Abdel Moaty

机构 * Faculty of Computer and Information System(计算机与信息系统系) Islamic University of Madinah(麦地那伊斯兰大学) Arab Open University-Bahrain(巴林阿拉伯开放大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种多模态框架,通过溯源感知机制防止代理AI中的提示注入攻击,提升系统安全性和稳定性。

Comments It is accepted in a conference paper, ICCA 2025 in Bahrain on 21 to 23 December

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04939 2025-12-23 cs.CV 79%

Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models

动词幻觉:揭示和评估多模态大语言模型中的动词概念幻觉

Zehao Wang, Xinpeng Liu, Yudonglin Zhang, Xiaoqian Wu, Zhou Fang, Yifan Fang, Junfu Pu, Cewu Lu, Yong-Lu Li

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文首次揭示多模态大语言模型中的动词概念幻觉问题,并提出基于丰富动词知识的调优方法以有效缓解该问题。

Comments Accepted by AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09121 2025-12-19 cs.CV 79%

MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models

MMRel:多模态大语言模型中关系理解的基准测试

Jiahao Nie, Gongjie Zhang, Wenbin An, Yun Xing, Yap-Peng Tan, Alex C. Kot, Shijian Lu

机构 * Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore(南洋理工大学跨学科研究生项目) Alibaba DAMO Academy, Singapore(阿里巴巴达摩院) Xi’an Jiaotong University, China(西安交通大学) Nanyang Technological University, Singapore(南洋理工大学) VinUniversity, Vietnam(文莱大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 MMRel是一个用于评估和提升多模态大语言模型关系理解能力的基准测试,包含大规模高质量的关系数据和对抗性案例,通过实验验证其在提升模型关系理解方面的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15611 2025-12-18 cs.CV 79%

If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions

如果你能描述它,他们就能看到它:从文本描述中跨模态学习视觉概念

Carlo Alberto Barbano, Luca Molinaro, Massimiliano Ciranni, Emanuele Aiello, Vito Paolo Pastore, Marco Grangetto

机构 * University of Turin(都灵大学) University of Genoa(热那亚大学) Politecnico di Torino(都灵理工大学)

专题命中 图文多模态 :cross-modal(title);image-text(abstract);分类 cs.CV

AI总结 本文提出通过文本描述跨模态学习视觉概念的方法,利用知识迁移技术提升视觉-语言模型的零样本性能。

Comments 27 pages. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23661 2025-12-16 cs.CV 79%

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

LLaVA-OneVision-1.5:面向民主化多模态训练的完全开放框架

Xiang An, Yin Xie, Kaicheng Yang, Wenkang Zhang, Xiuwei Zhao, Zheng Cheng, Yirui Wang, Songcen Xu, Changrui Chen, Didi Zhu, Chunsheng Wu, Huajie Tan, Chunyuan Li, Jing Yang, Jie Yu, Xiyao Wang, Bin Qin, Yumeng Wang, Zizhen Yan, Ziyong Feng, Ziwei Liu, Bo Li, Jiankang Deng

机构 * LLaVA-OneVision Community Contributors(LLaVA-OneVision社区贡献者)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 LLaVA-OneVision-1.5通过开放框架和高效训练方法实现了多模态模型的低成本高性能训练。

Comments LLaVA-OneVision-1.5 Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23065 2025-12-16 cs.CL 79%

SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services

SNS-Bench-VL:在社交网络服务中评估多模态大语言模型的基准测试

Hongcheng Guo, Zheyong Xie, Shaosheng Cao, Boyang Wang, Weiting Liu, Anjie Le, Lei Li, Zhoujun Li

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 SNS-Bench-VL是一个用于评估多模态大语言模型在社交网络服务中性能的基准测试,涵盖8个多模态任务,包含4001个问题-答案对,评估25种先进模型,揭示多模态社交理解的挑战。

Comments We found problems in the code while rechecking our implementation. These issues led to noticeable numerical discrepancies, making some of the reported results and conclusions potentially unreliable. Therefore, we request to withdraw this submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11683 2025-12-15 cs.CV 79%

Depth-Copy-Paste: Multimodal and Depth-Aware Compositing for Robust Face Detection

深度复制-粘贴:多模态和深度感知的复合技术用于鲁棒人脸检测

Qiushi Guo

机构 * Coffee AI Lab, Great Wall Motor(GWM), China(咖啡AI实验室,长城汽车(GWM),中国)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 深度复制-粘贴通过多模态和深度感知的增强方法,生成多样且物理一致的训练数据,提升人脸检测的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21451 2025-12-15 cs.CV 79%

MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning

MM-SeR: 多模态自反思用于轻量级图像描述

Junha Song, Yongsik Jo, So Yeon Min, Quanting Xie, Taehwan Kim, Yonatan Bisk, Jaegul Choo

机构 * KAIST(韩国科学技术院) UNIST(全南国立大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 MM-SeR通过轻量级多模态自反思框架,在减少计算成本的同时,实现了与大规模模型相当的图像描述性能,并扩展到长距离视频问答任务。

Comments Project page: https://sites.google.com/view/junha/mm-ser

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00087 2025-12-12 cs.CV 79%

Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data

探索多模态课堂数据中教学活动和话语的自动化识别

Ivo Bueno, Ruikun Hou, Babette Bühler, Tim Fütterer, James Drimalla, Jonathan Kyle Foster, Peter Youngs, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文通过多模态分析方法,实现了课堂活动中教学活动和话语的自动化识别,展示了微调模型在视频和 transcripts 上的高准确率,为可扩展的教师反馈系统提供了基础。

Comments This article has been accepted for publication in the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13886 2025-12-12 cs.CL 79%

Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning

Game-RL: 合成多模态可验证游戏数据以提升VLMs的通用推理能力

Jingqi Tong, Jixin Tang, Hangcheng Li, Yurong Mou, Ming Zhang, Jun Zhao, Yanbo Wen, Fan Song, Jiahao Zhan, Yuyang Lu, Chaoran Tao, Zhiyuan Guo, Jizhou Yu, Tianhao Cheng, Zhiheng Xi, Changhao Jiang, Zhangyue Yin, Yining Zheng, Weifeng Ge, Guanhua Chen, Tao Gui, Xipeng Qiu, Qi Zhang, Xuanjing Huang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 Game-RL通过合成多模态可验证游戏数据提升VLMs的通用推理能力。

Comments 69 pages, 24 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09092 2025-12-11 cs.CV 79%

Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters

解释未见的:多模态视觉-语言推理用于地下矿难中的情境感知

Mizanur Rahman Jewel, Mohamed Elmahallawy, Sanjay Madria, Samuel Frimpong

机构 * Missouri University of Science and Technology(密苏里科学与技术大学) Washington State University(华盛顿州立大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出MDSE框架,通过多模态视觉-语言推理提升地下矿难情境感知能力,实现更准确的描述生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07203 2025-12-09 cs.CV 79%

MMRPT: MultiModal Reinforcement Pre-Training via Masked Vision-Dependent Reasoning

MMRPT:通过遮蔽视觉依赖推理的多模态强化预训练

Xuhui Zheng, Kang An, Ziliang Wang, Yuhang Wang, Faqiang Qian, Yichao Wu

机构 * SenseTime(秒速科技)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 MMRPT通过强化学习提升多模态预训练效果,利用视觉基础推理增强模型对视觉内容的理解能力。

Comments 7 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04895 2025-12-05 cs.AI cs.MA 79%

Chameleon: Adaptive Adversarial Agents for Scaling-Based Visual Prompt Injection in Multimodal AI Systems

Chameleon:基于尺度的视觉提示注入的自适应对抗代理

M Zeeshan, Saud Satti

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 Chameleon是一种自适应对抗框架,通过动态优化图像扰动来暴露和利用多模态AI系统中的缩放漏洞,显著提高攻击成功率并影响决策准确性。

Comments 5 pages, 2 figures, IEEE Transactions on Dependable and Secure Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03657 2025-12-02 cs.CV 79%

Dynamic Multimodal Prototype Learning in Vision-Language Models

视觉-语言模型中的动态多模态原型学习

Xingyu Zhu, Shuo Wang, Beier Zhu, Miaoge Li, Yunfan Li, Junfeng Fang, Zhicai Wang, Dongsheng Wang, Hanwang Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学) The Hong Kong Polytechnic University(香港理工大学) Sichuan University(四川大学) National University of Singapore(新加坡国立大学) Shenzhen University(深圳大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出ProtoMM框架,通过动态更新视觉粒子和多模态原型学习,提升视觉-语言模型在测试时的适应性能。

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22664 2025-12-01 cs.CV 79%

VaMP: Variational Multi-Modal Prompt Learning for Vision-Language Models

VaMP:用于视觉-语言模型的变分多模态提示学习

Silin Cheng, Kai Han

机构 * Visual AI Lab, The University of Hong Kong(香港大学视觉人工智能实验室)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 VaMP提出了一种变分多模态提示学习框架,通过实例条件提示和不确定性建模提升视觉-语言模型在少样本和领域泛化任务中的性能。

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18437 2025-11-25 cs.CV 79%

Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning

基于感知证据的强化学习用于多模态推理

Chi Zhang, Haibo Qiu, Qiming Zhang, Yufei Xu, Zhixiong Zeng, Siqi Yang, Peng Shi, Lin Ma, Jing Zhang

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Meituan Inc(美团公司) The University of Sydney(悉尼大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 PEARL通过将推理锚定在验证的视觉证据上,提升多模态推理能力,有效解决视觉幻觉和奖励黑客问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18314 2025-11-25 cs.LG cs.AI 79%

AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert

AnyExperts: 多模态语言模型中基于需求的专家分配方法

Yuting Gao, Wang Lan, Hengyuan Zhao, Linjiang Huang, Si Liu, Qingpei Guo

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 AnyExperts通过按需分配专家资源,提高多模态MoE模型的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17448 2025-11-24 cs.CV 79%

MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models

MMT-ARD: 多模态多教师对抗蒸馏用于鲁棒视觉-语言模型

Yuqi Li, Junhao Dong, Chuanguang Yang, Shiping Wen, Piotr Koniusz, Tingwen Huang, Yingli Tian, Yew-Soon Ong

机构 * The City University of New York, CUNY(纽约城市大学) Nanyang Technological University(南洋理工大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Technology Sydney(悉尼技术大学) Data61, CSIRO(CSIRO数据61研究所) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 MMT-ARD通过多教师对抗蒸馏提升视觉-语言模型的对抗鲁棒性,实验显示鲁棒精度提升4.32%,训练效率提高2.3倍。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11313 2025-11-24 cs.CV 79%

DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding

DocSLM:一种用于长多模态文档理解的小型视觉-语言模型

Tanveer Hannan, Dimitrios Mallios, Parth Pathak, Faegheh Sardari, Thomas Seidl, Gedas Bertasius, Mohsen Fayyaz, Sunando Sengupta

机构 * Microsoft(微软公司) LMU Munich(慕尼黑大学) MCML FAIR Meta(Meta FAIR) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 DocSLM通过高效的小型视觉-语言模型,在受限内存下实现长多模态文档理解,使用更少的资源达到与先进方法相当的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07864 2025-11-18 cs.CV 79%

Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization

Tiancheng Yang, Lin Zhang, Jiaye Lin, Guimin Hu, Di Wang, Lijie Hu

机构 * MBZUAI Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室) King Abdullah University of Science and Technology(卡迪夫大学科学与技术大学) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院) University of Copenhagen(哥本哈根大学) Tsinghua University(清华大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10974 2025-11-17 cs.CV 79%

Preserving Cross-Modal Consistency for CLIP-based Class-Incremental Learning

Haoran Chen, Houze Xu, Micah Goldblum, Daoguo Dong, Zuxuan Wu

机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究院) Fudan University(复旦大学) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) Columbia University(哥伦比亚大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10074 2025-11-14 cs.CV cs.SY eess.SY 79%

VLF-MSC: Vision-Language Feature-Based Multimodal Semantic Communication System

Gwangyeon Ahn, Jiwan Seo, Joonhyuk Kang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments To appear in the AI4NextG Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19110 2025-11-14 cs.CV 79%

LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models

Zhihui Guo, Xin Man, Hui Xu, Jie Shao, Zhiguo Jiang, Xianchao Zhang, Heng Tao Shen

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07966 2025-11-12 cs.CV 79%

Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection

Shenao Zhao, Pengpeng Liang, Zhoufan Yang

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09997 2025-11-11 cs.CV 79%

Descriptive Image-Text Matching with Graded Contextual Similarity

Jinhyun Jang, Jiyoung Lee, Kwanghoon Sohn

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments This version is incomplete and requires substantial revisions and extensions. We withdraw the paper and plan to submit a thoroughly revised version as a new submission

详情

展开后加载摘要…

URL PDF HTML 收藏