arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6856 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6856 篇

2506.16895 2025-10-23 cs.CV cs.AI cs.LG 84%

With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You

Fabian Gröger, Shuo Wen, Huyen Le, Maria Brbić

机构 * EPFL(瑞士联邦理工学院) University of Basel(巴塞尔大学) HSLU(苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19136 2025-10-20 cs.CV cs.AI eess.IV 84%

PAD: Phase-Amplitude Decoupling Fusion for Multi-Modal Land Cover Classification

Huiling Zheng, Xian Zhong, Bin Liu, Yi Xiao, Bihan Wen, Xiaofeng Li

机构 * Sanya Science and Education Innovation Park, Wuhan University of Technology, Sanya 572025, China(武汉理工大学三亚科学教育创新园) School of Computer Science and Artificial Intelligence, Wuhan University of Technology, Wuhan 430070, China(武汉理工大学计算机科学与人工智能学院) Key Laboratory of Ocean Circulation and Waves, Institute of Oceanology, Chinese Academy of Sciences, Qingdao 266071, China(中国科学院海洋循环与波浪重点实验室) Hubei Key Laboratory of Transportation Internet of Things, School of Computer Science and Artificial Intelligence, Wuhan University of Technology, Wuhan 430070, China(湖北省交通物联网重点实验室) State Key Laboratory of Maritime Technology and Safety, Wuhan University of Technology, Wuhan 430063, China(武汉理工大学航海技术与安全国家重点实验室) College of Oceanography and Ecological Science, Shanghai Ocean University, Shanghai 201306, China(上海海洋大学海洋科学与生态学院) School of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China(郑州大学计算机与人工智能学院) Rapid-Rich Object Search Lab, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798(南洋理工大学电子与电气工程学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21447 2025-10-13 cs.CV cs.AI 84%

Multimodal Language Models See Better When They Look Shallower

Haoran Chen, Junyan Lin, Xinghao Chen, Yue Fan, Jianfeng Dong, Xin Jin, Hui Su, Jinlan Fu, Xiaoyu Shen

机构 * Zhejiang Gongshang University(浙江工商大学) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室) Institute of Digital Twin, Eastern Institute of Technology, Ningbo(数字孪生研究院,东部技术研究所,宁波) Meituan Inc.(美团公司) National University of Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 6 figures, accepted by EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08589 2025-10-13 cs.CV cs.AI 84%

Beyond CNNs: Efficient Fine-Tuning of Multi-Modal LLMs for Object Detection on Low-Data Regimes

Nirmal Elamon, Rouzbeh Davoudi

机构 * Artificial Creative intelligence (ACI)(人工创意智能(ACI)) Expedia Group(Expedia集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04943 2025-10-01 cs.CV cs.CL 84%

ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding

Jianjiang Yang, Yanshu li, Ziyan Huang

机构 * University of Bristol(布里斯托大学) Brown University(布朗大学) South China University of Technology(华南理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by conference EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22697 2025-09-30 cs.CV cs.AI cs.LG 84%

Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment

Abhiroop Chatterjee, Susmita Ghosh

机构 * Jadavpur University(贾瓦帕尔大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at the IEEE/CVF International Conference on Computer Vision (ICCV 2025), Workshop on Curated Data for Efficient Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19875 2025-09-25 cs.CV cs.AI 84%

Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection

Yunqing Hu, Zheming Yang, Chang Zhao, Wen Ji

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of AI for Industries(工业人工智能研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17136 2025-09-23 cs.CV cs.AI 84%

SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM

Yuhao Tian, Zheming Yang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of AI for Industries(工业人工智能研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16900 2025-09-23 cs.CV cs.AI 84%

ME-Mamba: Multi-Expert Mamba with Efficient Knowledge Capture and Fusion for Multimodal Survival Analysis

Chengsheng Zhang, Linhao Qu, Xiaoyu Liu, Zhijian Song

机构 * Digital Medical Research Center, School of Basic Medical Science, Fudan University, Shanghai 200032, China(复旦大学基础医学学院数字医学研究中心) Shanghai Key Lab of Medical Image Computing(上海医学图像计算重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16618 2025-09-23 cs.CV cs.AI 84%

Surgical-MambaLLM: Mamba2-enhanced Multimodal Large Language Model for VQLA in Robotic Surgery

Pengfei Hao, Hongqiu Wang, Shuaibo Li, Zhaohu Xing, Guang Yang, Kaishun Wu, Lei Zhu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Imperial College London(帝国理工学院伦敦分校) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Early accepted by MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15578 2025-09-22 cs.CV cs.AI 84%

Multimodal Learning for Fake News Detection in Short Videos Using Linguistically Verified Data and Heterogeneous Modality Fusion

Shanghong Li, Chiam Wen Qi Ruth, Hong Xu, Fang Liu

机构 * Singapore University of Social Sciences(新加坡社会科学研究大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13676 2025-09-18 cs.CV cs.AI 84%

Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation

Xiaobo Yang, Xiaojin Gong

机构 * Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19455 2025-09-12 cs.CV cs.AI cs.LG 84%

MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering

Xu Li, Fan Lyu

机构 * Khoury College of Computer Sciences, Northeastern University(东北大学克劳尔计算机科学学院) New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新实验室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06620 2025-09-09 cs.LG cs.AI cs.CL 84%

MedualTime: A Dual-Adapter Language Model for Medical Time Series-Text Multimodal Learning

Jiexia Ye, Weiqi Zhang, Ziyue Li, Jia Li, Meng Zhao, Fugee Tsung

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Technical University of Munich(慕尼黑技术大学) Columbia University(哥伦比亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 6 figure, 3 tables

Journal ref IJCAI 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19475 2025-09-04 cs.CV astro-ph.GA cs.AI cs.LG 84%

GalaxAlign: Mimicking Citizen Scientists' Multimodal Guidance for Galaxy Morphology Analysis

Ruoqi Wang, Haitao Wang, Qiong Luo

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) School of Computer Science and Engineering, Sun Yat-Sen University(中山大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18322 2025-08-27 cs.CV cs.AI 84%

Structures Meet Semantics: Multimodal Fusion via Graph Contrastive Learning

Jiangfeng Sun, Sihao He, Zhonghong Ou, Meina Song

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 9 pages,7 figures,conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04633 2025-08-27 cs.CV cs.AI 84%

M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering

Yanshu Li, Yi Cao, Hongyang He, Qisen Cheng, Xiang Fu, Xi Xiao, Tianyang Wang, Ruixiang Tang

机构 * Brown University(布朗大学) University of Warwick(沃里克大学) Samsung US(三星美国分公司) Boston University(波士顿大学) University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) Rutgers University(罗格斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments COLM 2025, 30 pages, 10 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22483 2025-08-18 cs.LG cs.AI cs.CV 84%

A Closer Look at Multimodal Representation Collapse

Abhra Chaudhuri, Anjan Dutta, Tu Bui, Serban Georgescu

机构 * Fujitsu Research of Europe(富士通欧洲研究)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments International Conference on Machine Learning (ICML) 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16658 2025-08-11 cs.CL cs.AI 84%

Contextual Reinforcement in Multimodal Token Compression for Large Language Models

Naderdel Piero, Zacharias Cromwell, Nathaniel Wainwright, Matthias Nethercott

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10905 2025-08-07 cs.AI cs.CV cs.LG 84%

Learning to Inference Adaptively for Multimodal Large Language Models

Zhuoyan Xu, Khoi Duc Nguyen, Preeti Mukherjee, Saurabh Bagchi, Somali Chaterji, Yingyu Liang, Yin Li

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Purdue University(普渡大学) The University of Hong Kong(香港大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Published at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23402 2025-08-01 cs.CV cs.AI cs.LG 84%

AGA: An adaptive group alignment framework for structured medical cross-modal representation learning

Wei Li, Xun Gong, Jiao Li, Xiaobin Sun

机构 * School of Computing and Artificial Intelligence(计算机与人工智能学院) Southwest Jiaotong University(西南交通大学) Department of Gastroenterology(消化内科部) The Third People’s Hospital of Chengdu(成都第三人民医院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22477 2025-08-01 cs.CV cs.AI 84%

LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks

Hui Liu, Chen Jia, Fan Shi, Xu Cheng, Mengfei Shi, Xia Xie, Shengyong Chen

机构 * Tianjin University of Technology(天津理工大学) Hainan University(海南大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19010 2025-07-31 cs.CV cs.CL 84%

Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection

Md. Mithun Hossain, Md. Shakil Hossain, Sudipto Chaki, M. F. Mridha

机构 * Department of Computer Science and Engineering, Bangladesh University of Business and Technology(计算机科学与工程系,孟加拉国商业技术大学) Department of Computer Science, American International University-Bangladesh(计算机科学系,美国国际大学-孟加拉国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20620 2025-07-29 cs.AI cs.CV 84%

Complementarity-driven Representation Learning for Multi-modal Knowledge Graph Completion

Lijian Li

机构 * Department of Computer and Information Science, University of Macau(计算机与信息科学系,澳门大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03726 2025-07-29 cs.CV cs.CL 84%

Otter: A Multi-Modal Model with In-Context Instruction Tuning

Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Joshua Adrian Cahyono, Jingkang Yang, Ziwei Liu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16854 2025-07-24 cs.CV cs.AI 84%

CLAMP: Contrastive Learning with Adaptive Multi-loss and Progressive Fusion for Multimodal Aspect-Based Sentiment Analysis

Xiaoqiang He

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18877 2025-07-21 cs.AI cs.CV 84%

UniEmoX: Cross-modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion Perception

Chuang Chen, Xiao Sun, Zhi Liu

机构 * School of Artificial Intelligence, Anhui University(安徽大学人工智能学院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥国家综合科学中心人工智能研究所) Anhui Province Key Laboratory of Affective Computing and Advanced Intelligent Machines, School of Computer Science and Information Engineering, Hefei University of Technology(安徽省情感计算与先进智能机器重点实验室,合肥工业大学计算机科学与信息工程学院) Department of Computer and Network Engineering, The University of Electro-Communications(电子通信大学计算机与网络工程系)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13415 2025-07-21 cs.MM cs.AI 84%

SEER: Semantic Enhancement and Emotional Reasoning Network for Multimodal Fake News Detection

Peican Zhu, Yubo Jing, Le Cheng, Bin Chen, Xiaodong Cui, Lianwei Wu, Keke Tang

机构 * School of Artificial Intelligence, Optics, and Electronics (iOPEN), Northwestern Polytechnical University(人工智能、光学与电子学院(iOPEN),西北工业大学) School of Computer Science, Northwestern Polytechnical University(计算机学院,西北工业大学) Unit 93212 of People’s Liberation Army of China(中国人民解放军第九三二一二单位) School of Marine Science and Technology, Northwestern Polytechnical University(海洋科学与技术学院,西北工业大学) Cyberspace Institute of Advanced Technology, Guangzhou University(高级技术网络空间研究院,广州大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

Comments Accepted by SMC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12566 2025-07-18 cs.CV cs.CL 84%

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

Gen Luo, Wenhan Dou, Wenhao Li, Zhaokai Wang, Xue Yang, Changyao Tian, Hao Li, Weiyun Wang, Wenhai Wang, Xizhou Zhu, Yu Qiao, Jifeng Dai

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16803 2025-07-11 cs.CV cs.AI cs.HC cs.LG eess.SP 84%

C3T: Cross-modal Transfer Through Time for Sensor-based Human Activity Recognition

Abhi Kamboj, Anh Duy Nguyen, Minh N. Do

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏