arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4868 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4868 篇

2503.07663 2025-10-23 cs.LG cs.AI 83%

Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMs

Dingkun Zhang, Shuhan Qi, Xinyu Xiao, Kehai Chen, Xuan Wang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Guangdong Provincial Key Laboratory of Novel Security Intelligence Technologies(广东省新型安全智能技术重点实验室)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12783 2025-10-22 cs.CV 83%

Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model

Yiming Shi, Xun Zhu, Kaiwen Wang, Ying Hu, Chenyi Guo, Miao Li, Ji Wu

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) College of AI, Tsinghua University(清华大学人工智能学院) Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16785 2025-10-21 cs.CV 83%

Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs

Jiazhen Liu, Long Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08565 2025-10-10 cs.CV 83%

NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints

Changyao Tian, Hao Li, Gen Luo, Xizhou Zhu, Weijie Su, Hanming Deng, Jinguo Zhu, Jie Shao, Ziran Zhu, Yunpeng Liu, Lewei Lu, Wenhai Wang, Hongsheng Li, Jifeng Dai

机构 * Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Sensetime Research(商汤科技研究院) Nanjing University(南京大学)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025. 22 pages, link: https://github.com/OpenGVLab/NaViL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18682 2025-09-24 cs.MM 83%

Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware Refinement

Beibei Zhang, Yanan Lu, Ruobing Xie, Zongyi Li, Siyuan Xing, Tongwei Ren, Fen Lin

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18198 2025-09-24 cs.AI cs.MA cs.RO 83%

MMCD: Multi-Modal Collaborative Decision-Making for Connected Autonomy with Knowledge Distillation

Rui Liu, Zikang Wang, Peng Gao, Yu Shen, Pratap Tokekar, Ming Lin

机构 * University of Maryland, College Park(马里兰大学 College Park 分校) North Carolina State University(北卡罗来纳州立大学) Adobe Research(Adobe 研究)

专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00717 2025-09-09 cs.NE cs.AI 83%

Advancements in Multimodal Differential Evolution: A Comprehensive Review and Future Perspectives

Dikshit Chauhan, Shivani, Donghwi Jung, Anupam Yadav

机构 * National University of Singapore(新加坡国立大学) Dr. B.R. Ambedkar National Institute of Technology Jalandhar(德拉·B.R. 阿姆贝德卡国家理工学院贾兰德哈尔) Korea University(韩国大学)

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Journal ref Artificial Intelligence Review 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02281 2025-09-08 cs.LG cs.MM 83%

Balanced Multimodal Learning: An Unidirectional Dynamic Interaction Perspective

Shijie Wang, Li Zhang, Xinyan Liang, Yuhua Qian, Shen Hu

机构 * Institute of Big Data Science and Industry(大数据科学与产业研究院) Key Laboratory of Evolutionary Science Intelligence of Shanxi Province(山西省进化智能科学重点实验室) Shanxi University(山西大学)

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12861 2025-08-19 cs.CV 83%

RMMSS: Towards Advanced Robust Multi-Modal Semantic Segmentation with Hybrid Prototype Distillation and Feature Selection

Jiaqi Tan, Xu Zheng, Yang Liu

专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15747 2025-08-19 cs.LG cs.AI 83%

Multi-modal Integration Analysis of Alzheimer's Disease Using Large Language Models and Knowledge Graphs

Kanan Kiguchi, Yunhao Tu, Katsuhiro Ajito, Fady Alnajjar, Kazuyuki Murase

专题命中 其他多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

Comments 38 pages, 8 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01663 2025-08-12 cs.CV 83%

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Xuan Yu, Dayan Guan, Yanfeng Gu

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Code is available at https://github.com/xavier-yu114/Zoom-Refine

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17471 2025-08-06 cs.LG cs.CR cs.CV 83%

Learning New Concepts, Remembering the Old: Continual Learning for Multimodal Concept Bottleneck Models

Songning Lai, Mingqian Liao, Zhangyi Hu, Jiayu Yang, Wenshuo Chen, Hongru Xiao, Jianheng Tang, Haicheng Liao, Yutao Yue

机构 * HKUST(GZ) Deep Interdisciplinary Intelligence Lab(香港科技大学(广州)深度跨学科智能实验室) Wuhan University(武汉大学) Shandong University(山东大学) Tongji University(同济大学) Peking University(北京大学) University of Macau(澳门大学) HKUST(GZ) Institute of Deep Perception Technology, JITRI Deep Interdisciplinary Intelligence Lab(香港科技大学(广州)感知技术研究所,JITRI深度跨学科智能实验室)

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Journal ref ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01594 2025-08-05 cs.CV 83%

CLIMD: A Curriculum Learning Framework for Imbalanced Multimodal Diagnosis

Kai Han, Chongwen Lyu, Lele Ma, Chengxuan Qian, Siqi Ma, Zheng Pang, Jun Chen, Zhe Liu

机构 * School of Computer Science Telecommunication Engineering, \ University, China

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments MICCAI 2025 Early Accept

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13060 2025-06-17 cs.AI cs.LG 83%

Rethinking Explainability in the Era of Multimodal AI

Chirag Agarwal

机构 * University of Virginia(弗吉尼亚大学)

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11550 2025-06-17 cs.LG cs.AI 83%

Improving Multimodal Learning Balance and Sufficiency through Data Remixing

Xiaoyu Ma, Hao Chen, Yongjian Deng

机构 * School of Computer Science and Engineering, Southeast University, Nanjing, China(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及交叉应用重点实验室) College of Computer Science, Beijing University of Technology, Beijing, China(北京理工大学计算机学院)

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments ICML2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00788 2025-06-11 cs.CV 83%

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

Wufei Ma, Luoxin Ye, Celso M de Melo, Jieneng Chen, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments CVPR 2025 highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21302 2025-05-01 cs.CV cs.RO 83%

CMD: Constraining Multimodal Distribution for Domain Adaptation in Stereo Matching

Zhelun Shen, Zhuo Li, Chenming Wu, Zhibo Rao, Lina Liu, Yuchao Dai, Liangjun Zhang

机构 * RAL, Baidu Research(百度研究院) ICT, Chinese Academy of Science(中国科学院信息科技研究所) Nanchang Hangkong University(南昌航空大学) Zhejiang University(浙江大学) Northwestern Polytechnical University(西北工业大学)

专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 13 pages, 5 figures, accepted for publication in Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13971 2025-04-22 cs.CY cs.AI cs.ET cs.NI 83%

The Future of Internet of Things and Multimodal Language Models in 6G Networks: Opportunities and Challenges

Abdelrahman Soliman

机构 * University of Guelph(圭尔夫大学)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16326 2025-03-21 cs.AI 83%

OmniGeo: Towards a Multimodal Large Language Models for Geospatial Artificial Intelligence

Long Yuan, Fengran Mo, Kaiyu Huang, Wenjie Wang, Wangyuxuan Zhai, Xiaoyu Zhu, You Li, Jinan Xu, Jian-Yun Nie

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments 15 pages, Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16384 2025-03-21 cs.CV cs.LG eess.IV 83%

MambaTron: Efficient Cross-Modal Point Cloud Enhancement using Aggregate Selective State Space Modeling

Sai Tarun Inaganti, Gennady Petrenko

专题命中 其他多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted to the Workshop on Image Quality in Computer Vision and Generative AI, WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15045 2025-03-20 cs.CV 83%

DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Wenhui Liao, Jiapeng Wang, Hongliang Li, Chengyu Wang, Jun Huang, Lianwen Jin

专题命中 其他多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05899 2025-03-11 cs.HC cs.AI 83%

Towards Understanding the Use of MLLM-Enabled Applications for Visual Interpretation by Blind and Low Vision People

Ricardo E. Gonzalez Penuela, Ruiying Hu, Sharon Lin, Tanisha Shende, Shiri Azenkot

专题命中 其他多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.AI

Comments 8 pages, 1 figure, 4 tables, to appear at CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11779 2025-02-25 cs.CL cs.AI cs.CV cs.LG cs.MM 83%

MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation

Chenxi Wang, Xiang Chen, Ningyu Zhang, Bozhong Tian, Haoming Xu, Shumin Deng, Huajun Chen

专题命中 其他多模态 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11701 2025-02-24 cs.CL cs.AI cs.CV cs.MM 83%

Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions

Yuhan Fu, Ruobing Xie, Jiazhen Liu, Bangxiang Lan, Xingwu Sun, Zhanhui Kang, Xirong Li

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments The proposed method does not work for up-to-date MLLMs.

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14093 2025-02-13 cs.MM 83%

Routing Experts: Learning to Route Dynamic Experts in Multi-modal Large Language Models

Qiong Wu, Zhaoxi Ke, Yiyi Zhou, Xiaoshuai Sun, Rongrong Ji

专题命中 其他多模态 :multi-modal(title,abstract);MLLM(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04317 2024-12-06 cs.CV 83%

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression

Bo Tong, Bokai Lai, Yiyi Zhou, Gen Luo, Yunhang Shen, Ke Li, Xiaoshuai Sun, Rongrong Ji

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09020 2024-11-08 cs.CL 83%

3M-Health: Multimodal Multi-Teacher Knowledge Distillation for Mental Health Detection

Rina Carines Cabral, Siwen Luo, Josiah Poon, Soyeon Caren Han

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted at CIKM 2024; Code will be made available at https://github.com/adlnlp/3mhealth

Journal ref CIKM '24 (2024) 152-162

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01307 2024-11-05 cs.CL 83%

Can Multimodal Large Language Model Think Analogically?

Diandian Guo, Cong Cao, Fangfang Yuan, Dakui Wang, Wei Ma, Yanbing Liu, Jianhui Fu

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09047 2024-10-14 cs.AI 83%

Modular Multimodal Machine Learning for Extraction of Theorems and Proofs in Long Scientific Documents (Extended Version)

Shrey Mishra, Antoine Gauquier, Pierre Senellart

专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18996 2024-10-01 cs.CL cs.AI cs.CV cs.LG cs.MM 83%

From Linguistic Giants to Sensory Maestros: A Survey on Cross-Modal Reasoning with Large Language Models

Shengsheng Qian, Zuyi Zhou, Dizhan Xue, Bing Wang, Changsheng Xu

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏