arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4868 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4868 篇

2502.19674 2025-12-01 cs.CV 79%

Reliable Multimodal Learning Via Multi-Level Adaptive DeConfusion

通过多级自适应去混淆实现可靠的多模态学习

Tong Zhang, Shu Shen, C. L. Philip Chen

机构 * Guangdong Provincial Key Laboratory of Computational AI Models and Cognitive Intelligence(广东省计算人工智能模型与认知智能重点实验室) School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) Pazhou Lab(琶洲实验室) Engineering Research Center of the Ministry of Education on Health Intelligent Perception and Paralleled Digital-Human(教育部健康智能感知与平行数字人工程研究中心)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出多级自适应去混淆方法,通过消除多模态数据中的类间和样本特定混淆,提升多模态模型的分类可靠性。

Comments 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19752 2025-11-26 cs.CV 79%

What You See is (Usually) What You Get: Multimodal Prototype Networks that Abstain from Expensive Modalities

你所见的通常就是你所得到的:一种多模态原型网络,能够避免昂贵的模态

Muchang Bahng, Charlie Berens, Jon Donnelly, Eric Chen, Chaofan Chen, Cynthia Rudin

机构 * Duke University(杜克大学) University of Maine(缅因大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种多模态原型网络,通过避免昂贵的遗传数据来提高物种检测的效率和可解释性。

Comments 19 pages. 16 figures. 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03151 2025-11-26 cs.CL cs.LG 79%

Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)

为何推理重要?多模态推理的进展综述(v1)

Jing Bi, Susan Liang, Xiaofei Zhou, Pinxin Liu, Junjia Guo, Yunlong Tang, Luchuan Song, Chao Huang, Ali Vosoughi, Guangyu Sun, Jinxi He, Jiarui Wu, Shu Yang, Daoan Zhang, Chen Chen, Lianggong Bruce Wen, Zhang Liu, Jiebo Luo, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) University of Central Florida(中央佛罗里达大学) Corning Inc.(康宁公司)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本文综述了多模态推理的进展,探讨了推理在多模态任务中的挑战与优化方法,为未来研究提供指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15478 2025-11-25 cs.CL 79%

Red Teaming Multimodal Language Models: Evaluating Harm Across Prompt Modalities and Models

针对多模态语言模型的红队测试:评估不同提示模态和模型的有害性

Madison Van Doren, Casey Ford

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本研究通过红队测试评估多模态语言模型在不同提示模态下的安全性,发现Pixtral 12B的有害响应率最高,而Claude Sonnet 3.5最安全,凸显了建立多模态安全基准的必要性。

Journal ref AAAI 2026 AIGOV Workshop and EurIPS 2025 Workshop on Unifying Perspectives on Learning Biases

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15600 2025-11-20 cs.CV cs.LG 79%

US-X Complete: A Multi-Modal Approach to Anatomical 3D Shape Recovery

US-X Complete: 一种多模态方法用于解剖三维形状恢复

Miruna-Alexandra Gafencu, Yordanka Velikova, Nassir Navab, Mohammad Farid Azampour

机构 * Computer-Aided Medical Procedures (CAMP), Technical University of Munich, Munich, Germany(计算机辅助医学程序(CAMP),慕尼黑技术大学,慕尼黑,德国) Munich Center for Machine Learning (MCML), Germany(慕尼黑机器学习中心(MCML),德国) Konrad Zuse School of Excellence in Reliable AI (relAI), Germany(康拉德·祖斯卓越可靠人工智能学校(relAI),德国)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 US-X Complete通过结合X光图像与超声数据,提升3D超声中椎体结构的重建精度,实现更完整的椎体可视化。

Comments Accepted at the Workshop on Shape in Medical Imaging at MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11212 2025-11-17 cs.CV 79%

MAFM^3: Modular Adaptation of Foundation Models for Multi-Modal Medical AI

Mohammad Areeb Qazi, Munachiso S Nwadike, Ibrahim Almakky, Mohammad Yaqub, Numan Saeed

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08810 2025-11-13 cs.CV 79%

SIFT-Graph: Benchmarking Multimodal Defense Against Image Adversarial Attacks With Robust Feature Graph

Jingjie He, Weijie Liang, Zihan Shan, Matthew Caesar

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ICCV2025 Workshop, short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08246 2025-11-12 cs.AI 79%

Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning

Ziyu Ma, Chenhui Gou, Yiming Hu, Yong Wang, Xiangxiang Chu, Bohan Zhuang, Jianfei Cai

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06749 2025-11-11 cs.RO cs.CV 79%

Semi-distributed Cross-modal Air-Ground Relative Localization

Weining Lu, Deer Bin, Lian Ma, Ming Ma, Zhihao Ma, Xiangyang Chen, Longfei Wang, Yixiao Feng, Zhouxian Jiang, Yongliang Shi, Bin Liang

机构 * Beijng National Research Center for Information Science and Technology(北京国家信息科学与技术研究中心) Qiyuan Lab(启元实验室) JiangHuai Advanced Technology Center(江淮先进技术中心)

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV

Comments 7 pages, 3 figures. Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11190 2025-11-07 cs.CV 79%

FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models

Shengming Yuan, Xinyu Lyu, Shuailong Wang, Beitao Chen, Jingkuan Song, Lianli Gao

机构 * University of Electronic Science and Technology of China(电子科学与技术大学) Southwestern University of Finance and Economics(西南财经大学) Tongji University(同济大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments 19 pages, 11 figures. Accepted by the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03617 2025-11-06 cs.GR cs.AI 79%

Visualization Biases MLLM's Decision Making in Network Data Tasks

Timo Brand, Henry Förster, Stephen G. Kobourov, Jacob Miller

机构 * Technical University of Munich, Heilbronn, Germany(慕尼黑技术大学)

专题命中 其他多模态 :MLLM(title,abstract);分类 cs.AI

Comments This manuscript was presented at VIS x GenAI, a workshop co-located with IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12926 2025-11-05 eess.SP cs.CV 79%

Task-Oriented Feature Compression for Multimodal Understanding via Device-Edge Co-Inference

Cheng Yuan, Zhening Liu, Jiashu Lv, Jiawei Shao, Yufei Jiang, Jun Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所) School of Electronic and Information Engineering, Harbin Institute of Technology(哈尔滨工业大学电子与信息工程学院) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(香港科技大学电子与计算机工程系) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Mobile Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00997 2025-11-04 cs.CV 79%

MID: A Self-supervised Multimodal Iterative Denoising Framework

Chang Nie, Tianchen Deng, Zhe Liu, Hesheng Wang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(自动化与智能感知学院,上海交通大学) Key Laboratory of System Control and Information Processing, Ministry of Education of China(系统控制与信息处理重点实验室,中华人民共和国教育部)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11807 2025-11-04 cs.CL 79%

Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks?

Simeon Junker, Manar Ali, Larissa Koch, Sina Zarrieß, Hendrik Buschmeier

机构 * Bielefeld University(比勒菲尔德大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

Comments To appear in ACL Findings 2025

Journal ref Findings of the Association for Computational Linguistics: ACL 2025, pp. 24101-24109

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20877 2025-10-27 cs.LG cs.AI 79%

Multimodal Negative Learning

Baoquan Gong, Xiyuan Gao, Pengfei Zhu, Qinghua Hu, Bing Cao

机构 * School of Artificial Intelligence, Tianjin University(人工智能学院,天津大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments Published in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12868 2025-10-21 cs.CV 79%

Computer-Aided Design of Personalized Occlusal Positioning Splints Using Multimodal 3D Data

Agnieszka Anna Tomaka, Leszek Luchowski, Michał Tarnawski, Dariusz Pojda

机构 * Institute of Theoretical and Applied Informatics, Polish Academy of Sciences(理论与应用信息学研究所,波兰科学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14022 2025-10-21 eess.IV cs.CV 79%

I2I-Mamba: Multi-modal medical image synthesis via selective state space modeling

Omer F. Atli, Bilal Kabas, Fuat Arslan, Arda C. Demirtas, Mahmut Yurt, Onat Dalmaz, Tolga Çukur

机构 * Department of Electrical and Electronics Engineering, and National Magnetic Resonance Research Center, Bilkent University(电子工程系和国家磁共振研究中心,比尔肯特大学) Department of Electrical Engineering, Stanford University(电气工程系,斯坦福大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18142 2025-10-07 cs.CV 79%

Autonomous Imagination: Closed-Loop Decomposition of Visual-to-Textual Conversion in Visual Reasoning for Multimodal Large Language Models

Jingming Liu, Yumeng Li, Boyuan Xiao, Yichang Jian, Ziang Qin, Tianjia Shao, Yao-Xiang Ding, Kun Zhou

机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments Published in TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02665 2025-10-06 cs.CL 79%

Self-Improvement in Multimodal Large Language Models: A Survey

Shijian Deng, Kai Wang, Tianyu Yang, Harsh Singh, Yapeng Tian

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Toronto(多伦多大学) University of Notre Dame(诺特丹大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00701 2025-10-02 cs.CV 79%

Graph Integrated Multimodal Concept Bottleneck Model

Jiakai Lin, Jinchang Zhang, Guoyu Lu

机构 * Intelligent Vision and Sensing (IVS) Lab at SUNY Binghamton(智能视觉与感知实验室(IVS)位于纽约州立大学布法罗分校)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12227 2025-09-30 cs.LG cs.AI 79%

Learning to Route: Per-Sample Adaptive Routing for Multimodal Multitask Prediction

Marzieh Ajirak, Oded Bein, Ellen Rose Bowen, Dora Kanellopoulos, Avital Falk, Faith M. Gunning, Nili Solomonov, Logan Grosenick

机构 * Department of Psychiatry, Weill Cornell Medicine(威立·科恩医学部) Feil Family Brain & Mind Research Institute, Weill Cornell Medicine(费尔家族脑与心灵研究研究所)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23475 2025-09-30 cs.CV 79%

Robust Multi-Modal Face Anti-Spoofing with Domain Adaptation: Tackling Missing Modalities, Noisy Pseudo-Labels, and Model Degradation

Ming-Tsung Hsu, Fang-Yu Hsu, Yi-Ting Lin, Kai-Heng Chien, Jun-Ren Chen, Cheng-Hsiang Su, Yi-Chen Ou, Chiou-Ting Hsu, Pei-Kai Huang

机构 * College of Computer and Cyber Security, Fujian Normal University(计算机与网络安全部,福建师范大学) Department of Computer Science, National Tsing Hua University(计算机科学系,国立清华大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21976 2025-09-25 cs.AI 79%

Compression Strategies for Efficient Multimodal LLMs in Medical Contexts

Tanvir A. Khan, Aranya Saha, Ismam N. Swapnil, Mohammad A. Haque

机构 * Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology (BUET)(电子与电气工程系,孟加拉国工程与技术大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07401 2025-09-25 cs.HC cs.AI 79%

Enhancing Higher Education with Generative AI: A Multimodal Approach for Personalised Learning

Johnny Chan, Yuming Li

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments 9 pages, 4 figures, accepted and presented in the 2025 6th International Conference on Advances in Education and Information Technology (AEIT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05644 2025-09-22 cs.CV eess.IV 79%

The Moon's Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction

Tom Sander, Moritz Tenthoff, Kay Wohlfarth, Christian Wöhler

机构 * Image Analysis Group, TU Dortmund University(图象分析组,多特蒙德大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments 48pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00046 2025-09-19 eess.IV cs.CV cs.LG 79%

Mixture of Multicenter Experts in Multimodal AI for Debiased Radiotherapy Target Delineation

Yujin Oh, Sangjoon Park, Xiang Li, Pengfei Jin, Yi Wang, Jonathan Paly, Jason Efstathiou, Annie Chan, Jun Won Kim, Hwa Kyung Byun, Ik Jae Lee, Jaeho Cho, Chan Woo Wee, Peng Shu, Peilong Wang, Nathan Yu, Jason Holmes, Jong Chul Ye, Quanzheng Li, Wei Liu, Woong Sub Koom, Jin Sung Kim, Kyungsang Kim

机构 * Center for Advanced Medical Computing and Analysis (CAMCA), Department of Radiology, Massachusetts General Hospital (MGH) and Harvard Medical School(先进医学计算与分析中心(CAMCA)、放射科、麻省总医院(MGH)和哈佛医学院) Department of Radiation Oncology, Yonsei University College of Medicine(燕京大学医学院放射肿瘤科) Institute for Innovation in Digital Healthcare, Yonsei University(数字医疗创新研究所、燕京大学) Department of Radiation Oncology, Massachusetts General Hospital(麻省总医院放射肿瘤科) Department of Radiation Oncology, Gangnam Severance Hospital(江南松云医院放射肿瘤科) Department of Radiation Oncology, Yongin Severance Hospital(永兴松云医院放射肿瘤科) School of Computing, University of Georgia(佐治亚大学计算机学院) Department of Radiation Oncology, Mayo Clinic(梅奥诊所放射肿瘤科) Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology(金 Jaechul人工智能研究生院、韩国科学技术院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments 12 pages, 5 figures, 4 tables, 1 supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04158 2025-09-12 cs.CV eess.IV 79%

Deep Learning-based Cross-modal Reconstruction of Vehicle Target from Sparse 3D SAR Image

Da Li, Guoqiang Zhao, Chen Yao, Kaiqiang Zhu, Houjun Sun, Jiacheng Bao, Maokun Li

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06617 2025-09-09 eess.IV cs.CV 79%

MM-DINOv2: Adapting Foundation Models for Multi-Modal Medical Image Analysis

Daniel Scholz, Ayhan Can Erdur, Viktoria Ehm, Anke Meyer-Baese, Jan C. Peeken, Daniel Rueckert, Benedikt Wiestler

机构 * Chair for AI for Image-Guided Diagnosis and Therapy, Technical University of Munich (TUM)(人工智能辅助影像诊断与治疗研究所,慕尼黑技术大学) TUM University Hospital, Munich, Germany(慕尼黑技术大学医院) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Chair for AI in Healthcare and Medicine, Technical University of Munich (TUM)(人工智能在医疗与健康领域研究所,慕尼黑技术大学) Department of Radiation Oncology, TUM University Hospital, Munich, Germany(放射肿瘤科,慕尼黑技术大学医院) Chair for Computer Vision and Artificial Intelligence, Technical University of Munich (TUM)(计算机视觉与人工智能研究所,慕尼黑技术大学) Department of Scientific Computing, Florida State University(科学计算系,佛罗里达州立大学) Deutsches Konsortium für Translationale Krebsforschung (DKTK), Partner Site Munich(德国转化癌症研究联盟(DKTK)慕尼黑分部) Institute of Radiation Medicine (IRM), Department of Radiation Sciences (DRS), Helmholtz Center Munich(放射医学研究所(IRM),辐射科学部门(DRS),海德堡中心慕尼黑) Institute for Advanced Study, Technical University of Munich (TUM)(高级研究所,慕尼黑技术大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06333 2025-09-09 cs.CV cs.RO 79%

Multi-Modal Camera-Based Detection of Vulnerable Road Users

Penelope Brown, Julie Stephany Berrio Perez, Mao Shan, Stewart Worrall

机构 * The University of Sydney(悉尼大学) Australian Centre for Robotics(澳大利亚机器人中心)

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06219 2025-09-09 cs.LG cs.MM 79%

MCIGLE: Multimodal Exemplar-Free Class-Incremental Graph Learning

Haochen You, Baojing Liu

机构 * Graduate School of Arts and Sciences(艺术与科学研究生院) Columbia University, New York, USA(哥伦比亚大学) School of Artificial Intelligence(人工智能学院) Hebei Institute of Communications, Shijiazhuang, PR China(河北通信学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.MM

Comments Accepted as a conference paper at KSEM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏