arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4868 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4868 篇

2604.26186 2026-04-30 cs.CV cs.HC cs.IR cs.MM 81%

FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing

FASH-iCNN:通过多模态CNN探针使编辑时尚身份可视化

Morayo Danielle Adeyemi, Ryan A. Rossi, Franck Dernoncourt

机构 * Howard University(霍华德大学) Adobe Research(Adobe研究)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 FASH-iCNN通过多模态CNN探针分析87547张Vogue runway图像,揭示时尚品牌、时代和色彩传统,提升时尚AI系统的可解释性。

Comments 5 pages, 4 tables, 1 figure. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10362 2026-04-28 cs.CV cs.AI 81%

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models

视觉漏斗:缓解多模态大语言模型中的上下文盲区

Woojun Jung, Jaehoon Go, Mingyu Jeon, Sunjae Yoon, Junyeong Kim

机构 * Department of AI, Chung-Ang University(Chung-Ang 大学人工智能系)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出视觉漏斗方法,通过上下文锚定和熵缩放投资组合缓解多模态大语言模型中的上下文盲区问题,通过动态调整裁剪尺寸和中心点,提升视觉细节与全局上下文的关联性。

Comments Accepted to CVPR 2026(Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00065 2026-04-21 cs.CL cs.AI 81%

Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models

使用视角词比词汇词对人类更困难,甚至对多模态语言模型更为困难

Dota Tianai Dong, Yifan Luo, Po-Ya Angela Wang, Asli Ozyurek, Paula Rubio-Fernandez

机构 * Max Planck Institute for Psycholinguistics(马克斯·普朗克心理学语言学研究所)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 研究比较了人类和多模态语言模型在使用词汇、所有格和指示词时的认知负荷,发现视角词对两者都更难,且多模态模型在所有格和指示词上表现更差,揭示了其在常识和社交认知能力上的不足。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10973 2026-04-14 cs.AI cs.CL 81%

CFMS: A Coarse-to-Fine Multimodal Synthesis Framework for Enhanced Tabular Reasoning

CFMS:一种用于增强表格推理的粗到细多模态合成框架

Qixian Huang, Hongqiang Lin, Tong Fu, Yingsen Wang, Zhenghui Fu, Qirui Wang, Yiding Sun, Dongxu Zhang

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出CFMS框架,通过粗到细的多模态合成方法提升表格推理能力,实验表明其在处理大表格和小模型时具有鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08476 2026-04-10 cs.CV cs.AI 81%

Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization

可信的GRPO:通过约束策略优化提升多模态语言模型的视觉空间推理

Sai Srinivas Kancheti, Aditya Kanade, Rohit Sinha, Vineeth N Balasubramanian, Tanuja Ganu

机构 * IIT Hyderabad(印度理工学院海得拉巴分校) Microsoft Research(微软研究院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出Faithful GRPO,通过约束策略优化提升多模态模型的空间推理质量,减少推理不一致率并提升视觉 grounding 评分。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25203 2026-03-27 cs.CV cs.CL 81%

Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection

基于概率概念图推理的多模态虚假信息检测

Ruichao Yang, Wei Gao, Xiaobin Zhu, Jing Ma, Hongzhan Lin, Ziyang Luo, Bo-Wen Zhang, Xu-Cheng Yin

机构 * University of Science and Technology Beijing(北京科技大学) Singapore Management University(新加坡管理大学) Hong Kong Baptist University(香港浸会大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出PCGR框架,通过构建概念图进行多模态虚假信息检测,实现可解释且可进化的方法,提升检测准确性和抗新兴操纵能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06663 2026-03-27 cs.CV cs.AI 81%

Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual Prompting

图标记:通过基于图的视觉提示提升多模态语言模型的空间推理能力

Giacomo Frisoni, Lorenzo Molfetta, Mattia Buzzoni, Gianluca Moro

机构 * University of Bologna(博洛尼亚大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出Graph-of-Mark,一种基于图的视觉提示方法,通过在输入图像上叠加场景图来增强多模态语言模型的空间推理能力,实验表明其在视觉问答和定位任务中提升了11个百分点的准确率。

Comments Please cite the definitive, copyrighted, and peer-reviewed version of this article published in AAAI 2026, edited by Sven Koenig et al., AAAI Press, Vol. 40, No. 36, Technical Track, pp. 30726-30734, 2026. DOI: https://doi.org/10.1609/aaai.v40i36.40329

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16786 2026-03-06 cs.LG cs.AI cs.CV 81%

Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach

重新审视多模态KV缓存压缩:一种基于频域的异常KV感知方法

Yaoxin Yang, Peng Ye, Xudong Tan, Chongjun Tu, Maosen Zhao, Jia Hao, Tao Chen

机构 * College of Future Information Technology, Fudan University(未来信息科技学院,复旦大学) Shanghai Innovation Institute(上海创新研究院) The Chinese University of Hong Kong(香港中文大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Zhangjiang Laboratory(张江实验室)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出FlashCache,一种基于频域的异常KV感知KV缓存压缩框架,通过保留关键KV对提升解码效率并降低内存使用。

Comments CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04453 2026-03-06 cs.CL cs.AI cs.LG 81%

Induced Numerical Instability: Hidden Costs in Multimodal Large Language Models

诱导的数值不稳定性:多模态大语言模型中的隐藏成本

Wai Tuck Wong, Jun Sun, Arunesh Sinha

机构 * School of Computing(计算学院) Information Systems, Singapore Management University, Singapore(信息系统,新加坡管理大学,新加坡) Information Systems Department, Rutgers Business School, New Jersey, USA(信息系统系,罗格斯商学院,新泽西,美国)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 研究揭示多模态大语言模型在推理阶段因优化数值不稳定性导致性能退化的新故障模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12880 2026-03-03 cs.AI cs.MM 81%

Has Multimodal Learning Delivered Universal Intelligence in Healthcare? A Comprehensive Survey

多模态学习是否在医疗领域实现了通用智能?一项全面的综述

Qika Lin, Yifan Zhu, Xin Mei, Ling Huang, Jingying Ma, Kai He, Zhen Peng, Erik Cambria, Mengling Feng

机构 * Saw Swee Hock School of Public Health, National University of Singapore(新加坡国立大学公共健康学院) School of Computer Science, Beijing University of Posts and Telecommunications(北京邮电大学计算机学院) School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院) School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

AI总结 本文通过全面调查,指出当前多模态学习在医疗领域尚未实现通用智能,并提出十个潜在研究方向。

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20918 2026-02-25 cs.AI cs.CL 81%

Predicting Sentence Acceptability Judgments in Multimodal Contexts

在多模态上下文中预测句子可接受性判断

Hyewon Jang, Nikolai Ilinykh, Sharid Loáiciga, Jey Han Lau, Shalom Lappin

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了在多模态上下文中,视觉上下文对人类和LLM句子可接受性判断的影响,发现LLM在去除视觉上下文时表现更优,且不同模型的判断分布各异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17124 2026-02-20 cs.CV cs.AI cs.RO 81%

3D Scene Rendering with Multimodal Gaussian Splatting

多模态高斯散射的3D场景渲染

Chi-Shiang Gau, Konstantinos D. Polyzos, Athanasios Bacharis, Saketh Madhuvarasu, Tara Javidi

机构 * UCSD Centers for Machine intelligence, computing, and security (MICS) and Wireless Communications (CWC)(UCSD 机器智能、计算与安全中心(MICS)和无线通信中心) UCSD NVIDIA

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出多模态框架,结合射频传感与高斯散射技术,实现高效且鲁棒的3D场景渲染。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08241 2026-02-10 cs.AI cs.CV 81%

Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs

MLLMs真的能看见吗:在多模态大语言模型中强化视觉注意力

Siqu Ou, Tianrui Wan, Zhiyuan Zhao, Junyu Gao, Xuelong Li

机构 * Shanghai Jiao Tong University(上海交通大学) Northwestern Polytechnical University(西北工业大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 SAYO通过强化学习框架引入区域级视觉注意力奖励,提升多模态大语言模型的视觉聚焦能力,从而在多种推理和感知任务中取得性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04145 2026-02-06 cs.LG cs.CL cs.MM 81%

Training Data Efficiency in Multimodal Process Reward Models

多模态过程奖励模型中的训练数据效率

Jinyuan Li, Chengsong Huang, Langlin Huang, Shaoyang Xu, Haolin Liu, Wenxuan Zhang, Jiaxin Huang

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校) Singapore University of Technology(新加坡科技设计大学) University of Virginia(弗吉尼亚大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM

AI总结 本文提出BIS方法,通过优化标签混合与可靠性,提升多模态过程奖励模型训练的数据效率,仅用10%训练数据即可达到全数据性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07825 2026-02-05 cs.CV cs.AI cs.LG 81%

Deep Multimodal Learning with Missing Modality: A Survey

缺失模态下的深度多模态学习:综述

Renjie Wu, Hu Wang, Hsiang-Ting Chen, Gustavo Carneiro

机构 * The Australian National University(澳大利亚国立大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Adelaide University(阿德莱德大学) The University of Surrey(萨里大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文综述了缺失模态下的多模态学习方法,分析了其动机、技术细节、应用及挑战,为该领域的发展提供了全面的视角。

Comments Accepted by TMLR (Transactions on Machine Learning Research)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11325 2026-02-03 cs.CV cs.AI 81%

Robust MLLM Unlearning via Visual Knowledge Distillation

通过视觉知识蒸馏实现鲁棒的多模态大语言模型去学习

Yuhang Wang, Zhenxing Niu, Haoxuan Ji, Guangyu He, Haichang Gao, Gang Hua

机构 * Xidian University, China(西安电子科技大学) XJTU University, China(西安交通大学) Amazon.com, USA(亚马逊公司)

专题命中 其他多模态 :MLLM(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出通过视觉知识蒸馏实现多模态大语言模型的鲁棒去学习,有效保留文本知识并提升效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14052 2025-12-17 cs.CV cs.CL 81%

HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices

HyperVL: 一种高效的多模态大语言模型用于边缘设备

HyperAI Team, Yuchen Liu, Kaiyang Han, Zhiqiang Xia, Yuhang Dong, Chen Song, Kangyu Tang, Jiaming Xu, Xiushi Feng, WenXuan Yu, Li Peng, Mingyang Wang, Kai Wang, Changpeng Yang, Yang Li, Haoyu Lu, Hao Wang, Bingna Xu, Guangyao Liu, Long Huang, Kaibin Guo, Jinyang Wu, Dan Wu, Hongzhen Wang, Peng Zhou, Shuai Nie, Shande Wang, Runyu Shi, Ying Huang

机构 * HyperAI Team(HyperAI团队) Xiaomi Corporation(小米公司)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 HyperVL是一种为边缘设备优化的高效多模态大语言模型,通过图像分块、视觉分辨率压缩和双一致性学习技术,实现低延迟、低功耗的多模态推理。

Comments Technical report of Xiaomi HyperAI Team

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21574 2025-11-27 cs.CV cs.AI 81%

Multimodal Robust Prompt Distillation for 3D Point Cloud Models

多模态鲁棒提示蒸馏用于3D点云模型

Xiang Gu, Liming Lu, Xu Zheng, Anan Du, Yongbin Zhou, Shuchao Pang

机构 * Xiang Gu 1 , Liming Lu 1 1 1 footnotemark: 1 , Xu Zheng 2,3 , Anan Du 4 , Yongbin Zhou 1 , Shuchao Pang 1 2 2 footnotemark: 2(某大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 多模态鲁棒提示蒸馏通过多模态知识蒸馏提升3D点云模型的鲁棒性,有效防御多种对抗攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20519 2025-11-26 cs.CV cs.AI 81%

Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning

Metis-HOME: 面向多模态推理的混合优化专家混合模型

Xiaohan Lan, Fanfan Liu, Haibo Qiu, Siqi Yang, Delian Ruan, Peng Shi, Lin Ma

机构 * Meituan(美团)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 Metis-HOME通过混合优化专家混合模型,提升多模态推理的复杂推理能力和通用能力,解决推理与泛化之间的矛盾。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19152 2025-11-25 eess.IV cs.AI cs.CV 81%

PSO-UNet: Particle Swarm-Optimized U-Net Framework for Precise Multimodal Brain Tumor Segmentation

PSO-UNet: 基于粒子群优化的U-Net框架用于精确多模态脑肿瘤分割

Shoffan Saifullah, Rafał Dreżewski

机构 * Faculty of Computer Science, AGH University of Krakow(计算机科学系,克拉科夫AGH大学) Department of Informatics, Universitas Pembangunan Nasional Veteran Yogyakarta(信息系,全国 veterans 大学 Yogya 市)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 PSO-UNet通过粒子群优化与U-Net结合,实现多模态脑肿瘤分割的高精度和高效优化。

Comments 9 pages, 6 figures, 4 tables, Gecco 2025 Conference

Journal ref GECCO '25 Companion: Proceedings of the Genetic and Evolutionary Computation Conference Companion, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03206 2025-11-06 cs.CV cs.AI cs.LG 81%

QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models

Kuei-Chun Kao, Hsu Tzu-Yin, Yunqi Hong, Ruochen Wang, Cho-Jui Hsieh

机构 * Department of Computer Science, University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 16 pages

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23118 2025-11-04 cs.CL cs.AI 81%

Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios

Zhongzhen Huang, Linjie Mu, Yakun Zhu, Xiangyu Zhao, Shaoting Zhang, Xiaofan Zhang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07862 2025-10-29 cs.LG cs.AI cs.CV 81%

ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources

Jason Wu, Yuyang Yuan, Kang Yang, Lance Kaplan, Mani Srivastava

机构 * Electrical and Computer Engineering University of California, Los Angeles(电气与计算机工程大学加州大学洛杉矶分校) DEVCOM Army Research Laboratory(国防部陆军研究实验室) University of California, Los Angeles(加州大学洛杉矶分校) Amazon(亚马逊)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to Neurips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11277 2025-10-28 eess.IV cs.AI cs.CV 81%

Macro2Micro: A Rapid and Precise Cross-modal Magnetic Resonance Imaging Synthesis using Multi-scale Structural Brain Similarity

Sooyoung Kim, Joonwoo Kwon, Junbeom Kwon, Jungyoun Janice Min, Sangyoon Bae, Yuewei Lin, Shinjae Yoo, Jiook Cha

机构 * Department of Brain and Cognitive Science, Seoul National University, Seoul, Republic of Korea(脑科学与认知科学系,首尔国立大学) Department of Applied Bioengineering, Seoul National University, Seoul, Republic of Korea(应用生物工程系,首尔国立大学) Department of Psychology, Seoul National University, Seoul, Republic of Korea(心理学系,首尔国立大学) Interdisciplinary Program in Artificial Intelligence, Seoul National University, Seoul, Republic of Korea(人工智能跨学科项目,首尔国立大学) Computational Science Initiative, Brookhaven National Laboratory, Upton, NY, USA(计算科学计划,布鲁赫斯国家实验室) School of Economics, Sogang University(经济学院,成均馆大学) Brookhaven National Laboratory(布鲁赫斯国家实验室)

专题命中 其他多模态 :cross-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments The code will be made available upon acceptance

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15172 2025-10-24 cs.IR cs.CL cs.CV 81%

X-Reflect: Cross-Reflection Prompting for Multimodal Recommendation

Hanjia Lyu, Ryan Rossi, Xiang Chen, Md Mehrab Tanjim, Stefano Petrangeli, Somdeb Sarkhel, Jiebo Luo

机构 * University of Rochester(罗切斯特大学) Adobe Research(Adobe研究)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14340 2025-10-17 eess.IV cs.AI cs.CV cs.LG 81%

A Density-Informed Multimodal Artificial Intelligence Framework for Improving Breast Cancer Detection Across All Breast Densities

Siva Teja Kakileti, Bharath Govindaraju, Sudhakar Sampangi, Geetha Manjunath

专题命中 其他多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01845 2025-10-03 cs.CL cs.CV 81%

Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models

Ece Takmaz, Lisa Bylinina, Jakub Dotlacil

机构 * Utrecht University(乌特勒支大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to the EMNLP 2025 workshop BabyLM: Accelerating language modeling research with cognitively plausible datasets

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19352 2025-09-25 cs.CL cs.AI 81%

TriSPrompt: A Hierarchical Soft Prompt Model for Multimodal Rumor Detection with Incomplete Modalities

Jiajun Chen, Yangyang Wu, Xiaoye Miao, Mengying Zhu, Meng Xi

机构 * Center for Data Science, Zhejiang University(数据科学中心,浙江大学) School of Software Technology, Zhejiang University(软件技术学院,浙江大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14982 2025-09-19 cs.CV cs.CL 81%

Large Multi-modal Models Can Interpret Features in Large Multi-modal Models

Kaichen Zhang, Yifei Shen, Bo Li, Ziwei Liu

机构 * S-Lab, NTU, Singapore(新加坡国立大学S实验室) LMMs-Lab Team(多模态模型实验室团队) Microsoft Research Asia(微软亚洲研究院)

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07742 2025-09-10 cs.HC cs.AI cs.CV 81%

Enhancing Online Learning by Integrating Biosensors and Multimodal Learning Analytics for Detecting and Predicting Student Behavior: A Review

Alvaro Becerra, Ruth Cobos, Charles Lang

机构 * Department of Computer Science Engineering, Universidad Autónoma de Madrid(马德里自治大学计算机科学工程系) Digital Futures Institute, Teachers College Columbia University(哥伦比亚大学教师学院数字未来研究所)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted for publication in Behaviour & Information Technology (Taylor & Francis). Final published version will be available soon at https://www.tandfonline.com/journals/tbit20

详情

展开后加载摘要…

URL PDF HTML 收藏