arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6878 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6878 篇

2409.09135 2025-10-30 cs.AI cs.CL cs.HC cs.LG 81%

Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation

Cheng Charles Ma, Kevin Hyekang Joo, Alexandria K. Vail, Sunreeta Bhattacharya, Álvaro Fernández García, Kailana Baker-Matsuoka, Sheryl Mathew, Lori L. Holt, Fernando De la Torre

机构 * Computer Science Department, Carnegie Mellon University(卡内基梅隆大学计算机科学系) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Neuroscience Institute, Carnegie Mellon University(卡内基梅隆大学神经科学研究所) Department of Psychology, The University of Texas at Austin(德克萨斯大学奥斯汀分校心理学系) Center for Perceptual Systems, The University of Texas at Austin(德克萨斯大学奥斯汀分校感知系统中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 22 pages, first three authors equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24777 2025-10-30 cs.CV cs.AI eess.IV 81%

Cross-Enhanced Multimodal Fusion of Eye-Tracking and Facial Features for Alzheimer's Disease Diagnosis

Yujie Nie, Jianzhang Ni, Yonglong Ye, Yuan-Ting Zhang, Yun Kwok Wing, Xiangqing Xu, Xin Ma, Lizhou Fan

机构 * School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学) Engineering Research Center of Intelligent Unmanned System, Ministry of Education(智能无人机系统工程研究中心,教育部) Department of Psychiatry, The Chinese University of Hong Kong(心理学系,香港中文大学) Department of Electronic Engineering, The Chinese University of Hong Kong(电子工程系,香港中文大学) AICARE Lab, Guangdong Medical University(AICARE实验室,广东医科大学) Department of Neurology, Shandong University of Traditional Chinese Medicine Affiliated Hospital(神经内科,山东中医药大学附属医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 35 pages, 8 figures, and 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22507 2025-10-28 cs.CV cs.AI 81%

GateFuseNet: An Adaptive 3D Multimodal Neuroimaging Fusion Network for Parkinson's Disease Diagnosis

Rui Jin, Chen Chen, Yin Liu, Hongfu Sun, Min Zeng, Min Li, Yang Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments The first two authors contributed equally to this work. Correspondence to: Yang Gao, E-mail: yang.gao@csu.edu.cn

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18279 2025-10-23 cs.CL cs.AI 81%

Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs

Yanhong Li, Zixuan Lan, Jiawei Zhou

机构 * Allen Institute for AI(艾伦人工智能研究所) University of Chicago(芝加哥大学) Stony Brook University(石溪大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings ("Text or Pixels? Evaluating Efficiency and Understanding of LLMs with Visual Text Inputs")

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19257 2025-10-22 cs.CV cs.CL 81%

MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models

Yinan Xia, Yilei Jiang, Yingshui Tan, Xiaoyong Zhu, Xiangyu Yue, Bo Zheng

机构 * Future Lab, Alibaba Group(阿里巴巴集团未来实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00916 2025-10-21 cs.CV cs.AI 81%

Enhancing Osteoporosis Detection: An Explainable Multi-Modal Learning Framework with Feature Fusion and Variable Clustering

Mehdi Hosseini Chagahi, Saeed Mohammadi Dashtaki, Niloufar Delfan, Nadia Mohammadi, Farshid Rostami Pouria, Behzad Moshiri, Md. Jalil Piran, Oliver Faust

机构 * School of Electrical and Computer Engineering, College of Engineering, University of Tehran(塔里班大学电气与计算机工程学院) Department of Epidemiology, Shiraz University of Medical Science(谢尔兹医学科学大学流行病学系) Department of Computer Science and Engineering, Sejong University(世宗大学计算机科学与工程系) School of Computing and Information Science, Anglia Ruskin University(安格利亚 Ruskin 大学计算与信息科学学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13131 2025-10-16 cs.CV cs.MM 81%

OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment

Rongjun Chen, Chengsi Yao, Jinchang Ren, Xianxian Zeng, Peixian Wang, Jun Yuan, Jiawen Li, Huimin Zhao, Xu Lu

机构 * School of Computer Science, Guangdong Polytechnic Normal University(广东 polytechnic 正规大学计算机学院)

专题命中 多模态训练与对齐 :image-text(title);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05839 2025-10-15 cs.MM cs.CV 81%

Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality

Hengyang Zhou, Yiwei Wei, Jian Yang, Zhenyu Zhang

机构 * Nanjing University(南京大学) China University of Petroleum(中国石油大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10406 2025-10-14 cs.CV cs.AI cs.LG 81%

Mesh-Gait: A Unified Framework for Gait Recognition Through Multi-Modal Representation Learning from 2D Silhouettes

Zhao-Yang Wang, Jieneng Chen, Jiang Liu, Yuxiang Guo, Rama Chellappa

机构 * Johns Hopkins University(约翰霍普金斯大学) Advanced Micro Devices, Inc.(先进微器件公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.06145 2025-10-13 cs.CV cs.AI cs.HC cs.LG stat.ML 81%

Improving the Performance of Unimodal Dynamic Hand-Gesture Recognition with Multimodal Training

Mahdi Abavisani, Hamid Reza Vaezi Joze, Vishal M. Patel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 1165-1174

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.06498 2025-10-13 cs.LG cs.AI cs.CV stat.ML 81%

Deep Multimodal Subspace Clustering Networks

Mahdi Abavisani, Vishal M. Patel

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Journal ref IEEE Journal of Selected Topics in Signal Processing, vol. 12, no. 6, pp. 1601-1614, Dec. 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07513 2025-10-10 cs.LG cs.AI cs.CV cs.DB 81%

MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis

Qinghua Liu, Sam Heshmati, Zheda Mai, Zubin Abraham, John Paparrizos, Liu Ren

机构 * The Ohio State University(俄亥俄州立大学) Bosch Research North America(博世北美研究)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24776 2025-09-30 cs.CV cs.AI 81%

VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding

Yizhuo Ding, Mingkang Chen, Zhibang Feng, Tong Xiao, Wanying Qu, Wenqi Shao, Yanwei Fu

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Shenzhen University(深圳大学) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24734 2025-09-30 cs.LG cs.AI cs.CV 81%

A TRIANGLE Enables Multimodal Alignment Beyond Cosine Similarity

Giordano Cicchetti, Eleonora Grassucci, Danilo Comminiello

机构 * Department of Information Engineering, Electronics, and Telecommunications(信息工程、电子与电信系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23109 2025-09-30 cs.AI cs.CV 81%

AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors

Junyang Zhang, Tianyi Zhu, Thierry Tambe

机构 * California Institute of Technology(加州理工学院) Stanford University(斯坦福大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 31 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22729 2025-09-30 cs.CL cs.AI 81%

Multi-Modal Sentiment Analysis with Dynamic Attention Fusion

Sadia Abdulhalim, Muaz Albaghdadi, Moshiur Farazi

机构 * University of Doha for Science and Technology(多哈科学技术大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CL、cs.AI

Comments Paper accepted for presentation at the ACS/IEEE 22nd International Conference on Computer Systems and Applications (AICCSA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22261 2025-09-29 cs.AI cs.CL 81%

InfiMed-Foundation: Pioneering Advanced Multimodal Medical Models with Compute-Efficient Pre-Training and Multi-Stage Fine-Tuning

Guanghao Zhu, Zhitian Hou, Zeyu Liu, Zhijie Sang, Congkai Xie, Hongxia Yang

机构 * The Hong Kong Polytechnic University(香港理工大学) Sun Yat-sen University(中山大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21854 2025-09-29 cs.MM cs.CV 81%

Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization

Songjun Tu, Qichao Zhang, Jingbo Sun, Yuqian Fu, Linjing Li, Xiangyuan Lan, Dongmei Jiang, Yaowei Wang, Dongbin Zhao

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) Pengcheng Laboratory(鹏城实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 12pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17492 2025-09-24 cs.CV cs.AI 81%

Multimodal Medical Image Classification via Synergistic Learning Pre-training

Qinghua Lin, Guang-Hai Liu, Zuoyong Li, Yang Li, Yuting Jiang, Xiang Wu

机构 * College of Biomedical Engineering, Fudan University(复旦大学生物医学工程学院) College of Computer Science and Engineering, Guangxi Normal University(广西师范大学计算机科学与工程学院) Fujian Provincial Key Laboratory of Information Processing and Intelligent Control, School of Computer and Big Data, Minjiang University(福建省信息处理与智能控制重点实验室,闽江学院计算机与大数据学院) Department of Automation Science and Electrical Engineering, Beihang University(北京航空航天大学自动化科学与电气工程学院) Department of Digestive Endoscopy, Fuzhou University Affiliated Provincial Hospital, Provincial Clinical Medical College of Fujian Medical University(福州市大学附属省医院消化内镜科,福建医科大学省临床医学学院) Department of Urology, Fuzhou University Affiliated Provincial Hospital, Provincial Clinical Medical College of Fujian Medical University(福州市大学附属省医院泌尿科,福建医科大学省临床医学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21059 2025-09-23 cs.CV cs.AI cs.CR cs.LG 81%

FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts

Ziyi Zhang, Zhen Sun, Zongmin Zhang, Jihui Guo, Xinlei He

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18042 2025-09-19 cs.CV cs.AI 81%

VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion

Pei Liu, Haipeng Liu, Haichao Liu, Xin Liu, Jinxin Ni, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Li Auto Inc.(Li汽车公司) the School of Aeronautics and Astronautics, Xiamen University(厦门大学航空航天学院) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12234 2025-09-17 cs.LG cs.AI cs.CV eess.IV 81%

Flexible Multimodal Neuroimaging Fusion for Alzheimer's Disease Progression Prediction

Benjamin Burns, Yuan Xue, Douglas W. Scharre, Xia Ning

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Department of Biomedical Informatics(生物医学信息学系) Department of Neurology(神经病学系) Translational Data Analytics Institute(转化数据分析研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at Applications of Medical AI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10408 2025-09-15 cs.CV cs.AI 81%

Multimodal SAM-adapter for Semantic Segmentation

Iacopo Curti, Pierluigi Zama Ramirez, Alioscia Petrelli, Luigi Di Stefano

机构 * University of Bologna(博洛尼亚大学) SINA

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21135 2025-09-15 cs.CV cs.AI 81%

HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection

Harris Song, Tuan-Anh Vu, Sanjith Menon, Sriram Narasimhan, M. Khalid Jawed

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments fix typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07923 2025-09-10 cs.CV cs.AI 81%

Multimodal Contrastive Pretraining of CBCT and IOS for Enhanced Tooth Segmentation

Moo Hyun Son, Juyoung Bae, Zelin Qiu, Jiale Peng, Kai Xin Li, Yifan Lin, Hao Chen

机构 * Department of Computer Science and Engineering(计算机科学与工程系) The Hong Kong University of Science and Technology(香港科学与技术大学) Division of Paediatric Dentistry and Orthodontics, Faculty of Dentistry(牙科学院儿童牙科与正畸科) The University of Hong Kong(香港大学) Delun Dental Hospital(德尔伦牙科医院) Department of Chemical and Biological Engineering(化学与生物工程系) Division of Life Science(生命科学系) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute(香港科技大学深圳-香港协同创新研究院) State Key Laboratory of Nervous System Disorders(神经系统紊乱国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06987 2025-09-10 cs.CV cs.AI 81%

FusWay: Multimodal hybrid fusion approach. Application to Railway Defect Detection

Alexey Zhukov, Jenny Benois-Pineau, Amira Youssef, Akka Zemmari, Mohamed Mosbah, Virginie Taillandier

机构 * University Bordeaux, CNRS, Bordeaux INP, INRIA, LaBRI(波尔多大学、国家科学研究中心、波尔多国立理工学院、INRIA、LaBRI) SNCF, DIR TECHNOLOGIES INNOVATION ET PROJETS GROUPE, IR - DPISF TECH4RAIL - TLI(法国国家铁路公司、技术与项目集团、IR - DPISF TECH4RAIL - TLI)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03961 2025-09-05 cs.CV cs.AI 81%

Multimodal Feature Fusion Network with Text Difference Enhancement for Remote Sensing Change Detection

Yijun Zhou, Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou, Xiaolin Tian, Xudong Jia, Hongsheng Zhang, C. L. Philip Chen

机构 * College of Electronics and Information Engineering, Wuyi University(威怡大学电子与信息工程学院) School of Electronic and Information Engineering and the Key Laboratory of Big Data and Intelligent Robot, Ministry of Education, South China University of Technology(电子与信息工程学院和大数据与智能机器人重点实验室,华南理工大学) State Key Laboratory of Lunar and Planetary Sciences, Macau University of Science and Technology(澳门大学地球和行星科学国家重点实验室) College of Engineering and Computer Science, California State University, Northridge(工程与计算机科学学院,加州大学北岭分校) Department of Geography, The University of Hong Kong(地理系,香港大学) Faculty of Computer Science and Engineering, S(计算机科学与工程学院,S)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00664 2025-09-03 cs.CV cs.AI 81%

Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model

Yifei She, Huangxuan Wu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19574 2025-08-28 cs.CV cs.AI 81%

Multimodal Prototype Alignment for Semi-supervised Pathology Image Segmentation

Mingxi Fu, Fanglei Fu, Xitong Ling, Huaitian Yuan, Tian Guan, Yonghong He, Lianghui Zhu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16744 2025-08-26 cs.LG cs.CL cs.CV 81%

Hyperbolic Multimodal Representation Learning for Biological Taxonomies

ZeMing Gong, Chuanqi Tang, Xiaoliang Huo, Nicholas Pellegrino, Austin T. Wang, Graham W. Taylor, Angel X. Chang, Scott C. Lowe, Joakim Bruslund Haurum

机构 * Simon Fraser University(西蒙弗雷泽大学) University of Waterloo(滑铁卢大学) Vector Institute(向量研究所) University of Guelph(圭尔夫大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所) Aalborg University(奥胡斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏