arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2505.23830 2025-06-02 cs.CL 79%

EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models

Linglin Jing, Yuting Gao, Zhigang Wang, Wang Lan, Yiwen Tang, Wenhai Wang, Kaipeng Zhang, Qingpei Guo

机构 * Shanghai AI Laboratory(上海人工智能实验室) Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23380 2025-05-30 cs.CV 79%

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Weijia Mao, Zhenheng Yang, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(1 显示实验室,新加坡国立大学) ByteDance(2 字节跳动)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10974 2025-05-30 cs.CV eess.IV 79%

Self-Supervised Enhancement of Forward-Looking Sonar Images: Bridging Cross-Modal Degradation Gaps through Feature Space Transformation and Multi-Frame Fusion

Zhisheng Zhang, Peng Zhang, Fengxiang Wang, Liangli Ma, Fuchun Sun

机构 * College of Electronic Engineering, Naval University of Engineering(电子工程学院,海军工程大学) College of Meteorology and Oceanography, National University of Defense Technology(气象海洋学院,国防科技大学) College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学) Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05214 2025-05-30 eess.IV cs.CV 79%

Gaussian Random Fields as an Abstract Representation of Patient Metadata for Multimodal Medical Image Segmentation

Bill Cassidy, Christian McBride, Connah Kendrick, Neil D. Reeves, Joseph M. Pappachan, Shaghayegh Raad, Moi Hoon Yap

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20904 2025-05-29 cs.CV 79%

HTMNet: A Hybrid Network with Transformer-Mamba Bottleneck Multimodal Fusion for Transparent and Reflective Objects Depth Completion

Guanghu Xie, Yonglong Zhang, Zhiduo Jiang, Yang Liu, Zongwu Xie, Baoshi Cao, Hong Liu

机构 * State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02183 2025-05-29 cs.AI cs.HC cs.LG eess.SP 79%

Multimodal sensor fusion in the latent representation space

Robert J. Piechocki, Xiaoyang Wang, Mohammud J. Bocus

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Under review for Nature Scientific Reports

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19507 2025-05-27 cs.CV cs.LG 79%

Multimodal Machine Translation with Visual Scene Graph Pruning

Chenyu Lu, Shiliang Sun, Jing Zhao, Nan Zhang, Tengfei Song, Hao Yang

机构 * East China Normal University(东华师范大学) Shanghai Jiao Tong University(上海交通大学) Wenzhou University(温州大学) Huawei Technologies Ltd(华为技术有限公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11514 2025-05-27 cs.CL 79%

Investigating Inference-time Scaling for Chain of Multi-modal Thought: A Preliminary Study

Yujie Lin, Ante Wang, Moye Chen, Jingyao Liu, Hao Liu, Jinsong Su, Xinyan Xiao

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) Baidu Inc., Beijing, China(百度公司) Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02810 2025-05-27 cs.LG cs.AI physics.chem-ph q-bio.BM 79%

Mol-LLM: Multimodal Generalist Molecular LLM with Improved Graph Utilization

Chanhui Lee, Hanbum Ko, Yuheon Song, YongJun Jeong, Rodrigo Hormazabal, Sehui Han, Kyunghoon Bae, Sungbin Lim, Sungwoong Kim

机构 * Department of Artificial Intelligence, Korea University(韩国大学人工智能系) Department of Artificial Intelligence, UNIST(UNIST人工智能系) Kim Jaechul Graduate School of AI, KAIST(韩国科学技术院人工智能研究生院) LG AI Research(LG人工智能研究) Department of Statistics, Korea University(韩国大学统计系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18178 2025-05-27 cs.LG cs.AI 79%

Less is More: Multimodal Region Representation via Pairwise Inter-view Learning

Min Namgung, Yijun Lin, JangHyeon Lee, Yao-Yi Chiang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17436 2025-05-26 cs.AI 79%

Scaling Up Biomedical Vision-Language Models: Fine-Tuning, Instruction Tuning, and Multi-Modal Learning

Cheng Peng, Kai Zhang, Mengxian Lyu, Hongfang Liu, Lichao Sun, Yonghui Wu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12453 2025-05-26 cs.MM 79%

Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding

Hanlei Zhang, Qianrui Zhou, Hua Xu, Jianhua Su, Roberto Evans, Kai Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16703 2025-05-23 cs.CL 79%

Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs

Zeping Yu, Sophia Ananiadou

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18047 2025-05-23 cs.CV cs.LG 79%

Progressive Local Alignment for Medical Multimodal Pre-training

Huimin Yan, Xian Yang, Liang Bai, Jiye Liang

专题命中 多模态训练与对齐 :multimodal(title);image-text(abstract);分类 cs.CV

Comments We are currently revising the methodology described in the manuscript to improve its clarity. We have decided to withdraw the current version until a more robust and complete version is ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14715 2025-05-22 eess.IV cs.CV 79%

A Comprehensive Review of Techniques, Algorithms, Advancements, Challenges, and Clinical Applications of Multi-modal Medical Image Fusion for Improved Diagnosis

Muhammad Zubair, Muzammil Hussai, Mousa Ahmad Al-Bashrawi, Malika Bendechache, Muhammad Owais

机构 * Interdisciplinary Research Center for Finance and Digital Economy, King Fahd University of Petroleum and Minerals(金融与数字经济交叉研究中心,国王法赫德石油和矿物大学) Department of Software Engineering, Faculty of Information Technology, Al-Ahliyya Amman University(软件工程系,信息科技学院,阿尔阿赫利亚大学) Department of Information Systems and Operations Management, King Fahd University of Petroleum and Minerals(信息系统与运营管理系,国王法赫德石油和矿物大学) ADAPT Research Centre, School of Computer Science, University of Galway(ADAPT研究中心,计算机科学学院,Galway大学) Department of Mechanical and Nuclear Engineering, Khalifa University(机械与核工程系,哈利法大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments computerized medical imaging and graphics Journal submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11997 2025-05-21 cs.CV 79%

Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality Rebalance

Mingcheng Qu, Guang Yang, Donglin Di, Tonghua Su, Yue Gao, Yang Song, Lei Fan

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) School of Software, Tsinghua University(清华大学软件学院) School of Computer Science and Engineering, UNSW Sydney(新南威尔士大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments accepted by IJCAI2025 Code: https://github.com/MCPathology/MRePath

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09018 2025-05-21 cs.CV cs.LG 79%

Multimodal Fusion of Glucose Monitoring and Food Imagery for Caloric Content Prediction

Adarsh Kumar

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments The manuscript was submitted without proper consideration of institutional policies. Upon review with professor, it was found that the content is subject to licensing restrictions which prohibit public dissemination in its current form. Therefore, I am withdrawing the paper to comply with these requirements

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13175 2025-05-20 cs.AI 79%

Enhancing LLMs for Time Series Forecasting via Structure-Guided Cross-Modal Alignment

Siming Sun, Kai Zhang, Xuejun Jiang, Wenchao Meng, Qinmin Yang

机构 * Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12782 2025-05-20 cs.GR cs.CV cs.IR cs.IT math.IT 79%

AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning

Kai Zhang, Xingyu Chen, Xiaofeng Zhang

机构 * Kai Zhang 1(某机构) Xingyu Chen 3(某机构) Xiaofeng Zhang 2(某机构)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12744 2025-05-20 cs.AI 79%

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation

Weiliang Tang, Dong Jing, Jia-Hui Pan, Zhiwu Lu, Yun-Hui Liu, Li Erran Li, Mingyu Ding, Chi-Wing Fu

机构 * Department of Computer Science(计算机科学系) The Chinese University of Hong Kong(香港中文大学) Gaoling School of Artificial Intelligence(九龙人工智能学院) Renmin University of China(中国人民大学) AWS AI Amazon(AWS人工智能亚马逊) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 17 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09897 2025-05-20 cs.CV 79%

TAMP: Token-Adaptive Layerwise Pruning in Multimodal Large Language Models

Jaewoo Lee, Keyang Xuan, Chanakya Ekbote, Sandeep Polisetty, Yi R. Fung, Paul Pu Liang

机构 * University of North Carolina Chapel Hill(北卡罗来纳大学教堂山分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Massachusetts Institute of Technology(麻省理工学院) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments ACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07372 2025-05-20 eess.IV cs.CV 79%

Multi-modal MRI Translation via Evidential Regression and Distribution Calibration

Jiyao Liu, Shangqi Gao, Yuxin Li, Lihao Liu, Xin Gao, Zhaohu Xing, Junzhi Ning, Yanzhou Su, Xiao-Yong Zhang, Junjun He, Ningsheng Xu, Xiahai Zhuang

机构 * Fudan University(复旦大学) University of Cambridge(剑桥大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Fuzhou University(福州市大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Early accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10862 2025-05-19 cs.CL 79%

Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?

Tairan Fu, Miguel González, Javier Conde, Elena Merino-Gómez, Pedro Reviriego

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments 6 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06592 2025-05-13 cs.CV 79%

Batch Augmentation with Unimodal Fine-tuning for Multimodal Learning

H M Dipu Kabir, Subrota Kumar Mondal, Mohammad Ali Moni

机构 * AI and Cyber Futures Institute, Charles Sturt University, Australia(人工智能与网络未来研究所,查尔斯·斯特劳特大学,澳大利亚) Rural Health Research Institute, Charles Sturt University, Australia(农村健康研究研究所,查尔斯·斯特劳特大学,澳大利亚) School of Computer Science and Engineering, Macau University of Science and Technology, Macao(计算机科学与工程学院,澳门科学技术大学,澳门)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06333 2025-05-13 cs.LG cs.AI 79%

NSF-MAP: Neurosymbolic Multimodal Fusion for Robust and Interpretable Anomaly Prediction in Assembly Pipelines

Chathurangi Shyalika, Renjith Prasad, Fadi El Kalach, Revathy Venkataramanan, Ramtin Zand, Ramy Harik, Amit Sheth

机构 * Artificial Intelligence Institute, University of South Carolina(南卡罗来纳大学人工智能研究所) Clemson Composites Center, Clemson University(克莱姆森大学复合材料中心) Intelligent Circuits, Architectures and Systems Lab, University of South Carolina(南卡罗来纳大学智能电路、架构和系统实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 9 pages, 7 figures, 2 tables, IJCAI 2025 (International Joint Conferences on Artificial Intelligence) Special Track on AI4Tech: AI Enabling Critical Technologies

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05877 2025-05-13 cs.LG cs.AI 79%

Multi-Modal Molecular Representation Learning via Structure Awareness

Rong Yin, Ruyue Liu, Xiaoshuai Hao, Xingrui Zhou, Yong Liu, Can Ma, Weiping Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyberspace Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Xidian University(西安电子科技大学) Renmin University of China(中国人民大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by IEEE Transactions on Image Processing (TIP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20698 2025-05-12 cs.CV cs.IR 79%

MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion

Saron Samuel, Dan DeGenaro, Jimena Guallar-Blasco, Kate Sanders, Oluwaseun Eisape, Tanner Spendlove, Arun Reddy, Alexander Martin, Andrew Yates, Eugene Yang, Cameron Carpenter, David Etter, Efsun Kayi, Matthew Wiesner, Kenton Murray, Reno Kriz

机构 * Stanford University(斯坦福大学) Georgetown University(乔治城大学) Johns Hopkins University(约翰霍普金斯大学) UC Berkeley(伯克利大学) BYU

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07503 2025-05-09 cs.AI cs.LG 79%

Recursive Inference Scaling: A Winning Path to Scalable Inference in Language and Multimodal Systems

Ibrahim Alabdulmohsin, Xiaohua Zhai

机构 * Google Deepmind(谷歌DeepMind)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22367 2025-05-07 q-bio.QM cs.AI cs.LG 79%

MAMMAL -- Molecular Aligned Multi-Modal Architecture and Language

Yoel Shoshan, Moshiko Raboh, Michal Ozery-Flato, Vadim Ratner, Alex Golts, Jeffrey K. Weber, Ella Barkan, Simona Rabinovici-Cohen, Sagi Polaczek, Ido Amos, Ben Shapira, Liam Hazan, Matan Ninio, Sivan Ravid, Michael M. Danziger, Yosi Shamay, Sharon Kurant, Joseph A. Morrone, Parthasarathy Suryanarayanan, Michal Rosen-Zvi, Efrat Hexter

机构 * IBM Research-Israel(IBM研究以色列分公司) IBM TJ Watson Research Center(IBM TJ Watson研究中心) Faculty of Biomedical Engineering(生物医学工程学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02529 2025-05-06 eess.IV cs.CV 79%

RobSurv: Vector Quantization-Based Multi-Modal Learning for Robust Cancer Survival Prediction

Aiman Farooq, Azad Singh, Deepak Mishra, Santanu Chaudhury

机构 * Indian Institute of Technology Jodhpur(印度理工学院朱道尔分校) Indian Institute of Technology Delhi(印度理工学院德里)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏