arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-03 至 2026-03-03 共收录 19 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 19 篇

2509.03113 2026-03-03 cs.CV cs.CL 86%

Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection

通过基于梯度的自我反思缓解多模态幻觉

Shan Wang, Maying Shen, Nadine Chang, Chuong Nguyen, Hongdong Li, Jose M. Alvarez

机构 * NVIDIA Australian National University(澳大利亚国立大学) Data61, CSIRO(Data61,CSIRO)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出GACD方法,通过梯度分析缓解多模态模型的幻觉问题,提升输出的视觉基础性。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22171 2026-03-03 cs.HC 82%

A Taxonomy of Human--MLLM Interaction in Early-Stage Sketch-Based Design Ideation

早期阶段基于草图的设计构想中人类与大语言模型交互的分类

Weiyan Shi, Kenny Tsu Wei Choo

专题命中 其他多模态 :MLLM(title,abstract);multimodal(abstract)

AI总结 本文提出了一种分类方法,用于描述人类与大语言模型在早期阶段基于草图的设计构想中的交互模式,揭示了人类与AI角色的动态变化。

Comments Accepted at CHI 2026 Posters

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12880 2026-03-03 cs.AI cs.MM 81%

Has Multimodal Learning Delivered Universal Intelligence in Healthcare? A Comprehensive Survey

多模态学习是否在医疗领域实现了通用智能?一项全面的综述

Qika Lin, Yifan Zhu, Xin Mei, Ling Huang, Jingying Ma, Kai He, Zhen Peng, Erik Cambria, Mengling Feng

机构 * Saw Swee Hock School of Public Health, National University of Singapore(新加坡国立大学公共健康学院) School of Computer Science, Beijing University of Posts and Telecommunications(北京邮电大学计算机学院) School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院) School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

AI总结 本文通过全面调查,指出当前多模态学习在医疗领域尚未实现通用智能,并提出十个潜在研究方向。

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01106 2026-03-03 cs.AI 79%

DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage

DIVA-GRPO:通过难度自适应变体优势增强多模态推理

Haowen Gao, Zhenyu Zhang, Liang Pang, Fangda Guo, Hongjian Dou, Guannan Lv, Shaoguo Liu, Tingting Gao, Huawei Shen, Xueqi Cheng

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, Beijing, China(人工智能安全国家重点实验室,计算技术研究所,中国科学院,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) Kuaishou Technology, Beijing, China(快手科技,北京,中国)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 DIVA-GRPO通过难度自适应变体优势方法提升多模态推理能力,解决GRPO在困难问题上的奖励稀疏性和优势消失问题,提升训练稳定性与推理性能。

Comments Accepted to ICLR 2026. Code and models are available at https://github.com/Siaaaaaa1/DIVA-GRPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00289 2026-03-03 cs.CV 79%

Seeking Necessary and Sufficient Information from Multimodal Medical Data

从多模态医学数据中寻求必要和充分的信息

Boyu Chen, Weiye Bao, Junjie Liu, Michael Shen, Bo Peng, Paul Taylor, Zhu Li, Mengyue Yang

机构 * University College London, London, UK(伦敦大学学院) Imperial College London, London, UK(伦敦帝国学院) Mingdu Tech, China(明都科技) University of Bristol, Bristol, UK(布里斯托大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出通过概率必要性和充分性学习多模态医学数据中的必要和充分特征,以提升模型性能和鲁棒性。

Comments 11 pages, 1 figure. Submitted to MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27492 2026-03-03 cs.CV 79%

ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning

ThinkMorph:多模态交错链式推理中的涌现特性

Jiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li, Michael Qizhe Shieh, Yejin Choi, Ranjay Krishna, Yu Cheng

机构 * National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学) University of Washington(华盛顿大学) Stanford University(斯坦福大学) absolute AI The Chinese University of Hong Kong(香港中文大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 ThinkMorph通过统一模型提升多模态推理性能,展现视觉操控与模式切换等新兴能力。

Comments project page: https://thinkmorph.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02339 2026-03-03 cs.CL 79%

AStar: Boosting Multimodal Reasoning with Automated Structured Thinking

AStar: 通过自动化结构化思维提升多模态推理

Jinyang Wu, Mingkuan Feng, Guocheng Zhai, Shuai Zhang, Zheng Lian, Fangrui Lv, Pengpeng Shao, Ruihan Jin, Zhengqi Wen, Jianhua Tao

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 AStar通过自动化结构化思维提升多模态推理效率,实现更高准确率和更强迁移能力。

Comments Accepted by AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02109 2026-03-03 eess.SP cs.LG 78%

Orchestrating Multimodal DNN Workloads in Wireless Neural Processing

在无线神经处理中协调多模态DNN工作负载

Sai Xu, Kai-Kit Wong, Yanan Du, Hyundong Shin

机构 * Department of Electronic and Electrical Engineering, University College London(电子与电气工程系,伦敦大学学院) Department of Electronic Engineering, Kyung Hee University(电子工程系,庆熙大学) School of Electrical and Electronic Engineering, the University of Sheffield(电气与电子工程学院,谢菲尔德大学) Department of Electronics and Information Convergence Engineering, Kyung Hee University(电子与信息融合工程系,庆熙大学)

专题命中 其他多模态 :multimodal(title,abstract)

AI总结 本文提出O-WiN框架和PACS算法,通过通信-计算流水线优化无线神经处理中多模态DNN工作负载,提升执行效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00565 2026-03-03 cs.CV cs.AI cs.CR 62%

MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs

MIDAS: 多图像分散与语义重建用于对抗多模态大语言模型

Yilian Liu, Xiaojun Jia, Guoshun Nan, Jiuyang Lyu, Zhican Chen, Tao Guan, Shuyuan Luo, Zhongyi Zhai, Yang Liu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学) Guilin University of Electronic Technology(桂林电子科技大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 MIDAS通过多图像分散与语义重建技术,提升对抗多模态大语言模型的劫持性能,达到81.46%的平均攻击成功率。

Journal ref The Fourteenth International Conference on Learning Representations(2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01948 2026-03-03 cs.CV 57%

PreSight: Preoperative Outcome Prediction for Parkinson's Disease via Region-Prior Morphometry and Patient-Specific Weighting

PreSight:通过区域先验形态学和患者特异性加权进行帕金森病术前预后预测

Yand Wang, Chen Zhang, Lanyun Zhu, Yixin Chen, Qunbo Wang, Yutong Bai, Jurgen Germann, Yinghong Wen, Shuai Shao

机构 * Beijing Jiaotong University(北京交通大学) Nanyang Technological University(南洋理工大学) Institute of Medical Technology, Peking University(北京大学医学技术研究院) Beijing Tiantan Hospital, Capital Medical University(北京天坛医院) University Health Network, University of Toronto(多伦多大学健康网络) Suzhou Institute for Advanced Research, University of Science and Technology of China(中国科学技术大学苏州研究院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 PreSight通过结合临床先验与区域自适应形态学,实现帕金森病术前预后预测,提升术后运动改善预测的准确性和临床实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23348 2026-03-03 cs.RO cs.CV 57%

Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts

为拟合物体操纵的物理常识知识而引入分析概念

Jiude Wei, Yuxuan Li, Cewu Lu, Jianhua Sun

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机学院) Shanghai Innovation Institute(上海创新研究院)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出通过引入分析概念,将语义级常识知识接地到物理世界,以提升机器人对关节物体的通用精确操纵能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04490 2026-03-03 cs.CL q-bio.GN 57%

Large Language Models in Bioinformatics: A Survey

大语言模型在生物信息学中的应用:综述

Zhenyu Wang, Zikang Wang, Jiyue Jiang, Pengan Chen, Xiangyu Shi, Yu Li

机构 * The Chinese University of Hong Kong(香港中文大学) Peking University Third Hospital(北京大学第三医院) The Hong Kong Polytechnic University(香港理工大学) The University of Hong Kong(香港大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

AI总结 本文综述了大语言模型在生物信息学中的应用,涵盖基因组序列建模、RNA结构预测等核心方法,并探讨了数据稀缺和跨组学整合等挑战及未来发展方向。

Comments Accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01026 2026-03-03 cs.CV 57%

RaUF: Learning the Spatial Uncertainty Field of Radar

RaUF: 学习雷达的时空不确定性场

Shengpeng Wang, Kuangyu Wang, Wei Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Wuhan University(武汉大学)

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV

AI总结 RaUF通过学习雷达测量的各向异性特性,解决方位模糊和虚假回波问题,提升空间检测的可靠性与不确定性校准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00988 2026-03-03 cs.CV cs.SE 57%

Foundation Models in Remote Sensing: Evolving from Unimodality to Multimodality

遥感中的基础模型:从单模态到多模态的演变

Danfeng Hong, Chenyu Li, Xuyang Li, Gustau Camps-Valls, Jocelyn Chanussot

机构 * School of Automation, Southeast University(自动化学院,东南大学) School of Mathematics, Southeast University(数学学院,东南大学) Aerospace Information Research Institute, Chinese Academy of Sciences(航天信息研究所,中国科学院) Image Processing Laboratory (IPL), Universitat de València(图像处理实验室(IPL),瓦伦西亚大学) Univ. Grenoble Alpes, INRIA, CNRS, Grenoble INP, LJK(格勒诺布尔阿尔卑斯大学,INRIA,CNRS,格勒诺布尔INP,LJK)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文探讨了遥感中基础模型从单模态到多模态的演变,旨在为研究人员提供深入理解与应用指导。

Comments Accepted by IEEE GRSM

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00436 2026-03-03 cs.LG cs.AI 57%

ROKA: Robust Knowledge Unlearning against Adversaries

ROKA: 面对对抗者的鲁棒知识反学习

Jinmyeong Shin, Joshua Tapia, Nicholas Ferreira, Gabriel Diaz, Moayed Daneshyari, Hyeran Jeon

机构 * University of California, Merced(加州大学梅尔塞德斯分校) California State University, East Bay(加州州立大学东湾分校)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 ROKA通过神经愈合机制实现鲁棒的知识反学习,有效对抗间接反学习攻击,同时保护保留数据的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22753 2026-03-03 math.OC math.PR 50%

Enhancing Exploration in Global Optimization by Noise Injection in the Probability Measures Space

通过在概率测度空间中注入噪声增强全局优化的探索

Gaëtan Serré, Pierre Germain, Samuel Gruffaz, Argyris Kalogeratos

专题命中 其他多模态 :multimodal(abstract)

AI总结 本文通过在概率测度空间中注入噪声,提升全局优化中探索和收敛能力,适用于多种动态配置。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18460 2026-03-03 cs.LG 50%

Learning Boltzmann Generators via Constrained Mass Transport

通过约束质量传输学习Boltzmann生成器

Christopher von Klitzing, Denis Blessing, Henrik Schopmans, Pascal Friederich, Gerhard Neumann

机构 * Autonomous Learning Robots, Karlsruhe Institute of Technology(自动化学习机器人,卡尔斯鲁厄大学技术学院) Artificial Intelligence for Materials Sciences, Karlsruhe Institute of Technology(材料科学人工智能,卡尔斯鲁厄大学技术学院)

专题命中 其他多模态 :multimodal(abstract)

AI总结 本研究提出约束质量传输框架,通过约束KL散度和熵衰减来提升Boltzmann生成器的采样效果,有效避免模式崩溃并提高样本效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00313 2026-03-03 nlin.AO physics.comp-ph 50%

Synchronization, Collective Oscillations, and Information Flow in Duplex Networks

同步、集体振荡与双网络中的信息流

Ali Seif, Mina Zarei

专题命中 其他多模态 :multimodal(abstract)

AI总结 研究双网络中部分同步与集体振荡的机制,揭示多模式动态的形成原理。

Comments 26 pages (21 main and 5 Supplementary), 12 figures (8 main and 4 Supplementary)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21029 2026-03-03 cs.LG 50%

FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction

FORCE:通过特征过度依赖校正实现可转移的视觉劫持攻击

Runqi Lin, Alasdair Paren, Suqin Yuan, Muyang Li, Philip Torr, Adel Bibi, Tongliang Liu

机构 * Sydney AI Centre, The University of Sydney(悉尼人工智能中心,悉尼大学) Department of Engineering Science, University of Oxford(工程科学系,牛津大学)

专题命中 其他多模态 :multimodal(abstract)

AI总结 FORCE方法通过校正特征过度依赖,提升视觉劫持攻击的跨模型可转移性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏