arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6872 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6872 篇

2503.00513 2025-06-17 cs.CV 83%

Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning

Hanxun Yu, Wentong Li, Song Wang, Junbo Chen, Jianke Zhu

机构 * Zhejiang University(浙江大学) Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments CVPR2025, Code Link: https://github.com/hanxunyu/Inst3D-LMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18688 2025-06-17 cs.CR cs.AI cs.LG 83%

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment

Soumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Tianrui Guan, Mengdi Wang, Alvaro Velasquez, Ahmad Beirami, Furong Huang, Dinesh Manocha, Amrit Singh Bedi

机构 * University of Maryland(马里兰大学) Indian Institute of Technology Bombay(印度班加罗尔理工学院) Princeton University(普林斯顿大学) University of Colorado Boulder(科罗拉多大学博尔德分校) Capital One(Capital One公司) University of Central Florida(佛罗里达中央大学)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.AI

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10019 2025-06-17 cs.CV 83%

Dissecting RGB-D Learning for Improved Multi-modal Fusion

Hao Chen, Haoran Zhou, Yunshu Zhang, Zheng Lin, Yongjian Deng

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及跨学科应用国家重点实验室(东南大学)) Tsinghua University(清华大学) College of Computer Science, Beijing University of Technology(北京理工大学计算机学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11691 2025-06-16 cs.CV 83%

DMAF-Net: An Effective Modality Rebalancing Framework for Incomplete Multi-Modal Medical Image Segmentation

Libin Lan, Hongxing Li, Zunhui Xia, Yudong Zhang

机构 * College of Computer Science and Engineering, Chongqing University of Technology(重庆理工大学计算机科学与工程学院) School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 12 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11672 2025-06-16 cs.CV 83%

Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning

Chendi Ge, Xin Wang, Zeyang Zhang, Hong Chen, Jiapei Fan, Longtao Huang, Hui Xue, Wenwu Zhu

机构 * Department of Computer Science and Technology, BNRist, Tsinghua University, Beijing, China(计算机科学与技术系,BNRist,清华大学,北京,中国) Alibaba Group, Hangzhou, China(阿里巴巴集团,杭州,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07138 2025-06-10 cs.CV cs.AI cs.CL cs.MM 83%

Learning Compact Vision Tokens for Efficient Large Multimodal Models

Hao Tang, Chengchao Shen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments The source code and trained weights are available at https://github.com/visresearch/LLaVA-STF

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02308 2025-06-10 cs.LG cs.AI 83%

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping

Xiaojun Shan, Qi Cao, Xing Han, Haofei Yu, Paul Pu Liang

机构 * University of California, San Diego(加州大学圣地亚哥分校) Johns Hopkins University(约翰霍普金斯大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04588 2025-06-05 cs.AI 83%

Zero-shot cross-modal transfer of Reinforcement Learning policies through a Global Workspace

Léopold Maytié, Benjamin Devillers, Alexandre Arnold, Rufin VanRullen

机构 * Airbus AI Research(空客人工智能研究)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

Journal ref Reinforcement Learning Journal, Vol. 3, 2024, pp. 1410-1426

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24176 2025-06-02 cs.MM 83%

ISMAF: Intrinsic-Social Modality Alignment and Fusion for Multimodal Rumor Detection

Zihao Yu, Xiang Li, Jing Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23365 2025-05-30 cs.CV 83%

MCFNet: A Multimodal Collaborative Fusion Network for Fine-Grained Semantic Classification

Yang Qiao, Xiaoyu Zhong, Xiaofeng Gu, Zhiguo Yu

机构 * School of Integrated Circuits, Jiangnan University, Wuxi 214122, China(集成电路学院,江南大学,无锡214122,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12941 2025-05-30 cs.CL cs.LG 83%

HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model

Haiyang Guo, Fanhu Zeng, Ziwei Xiang, Fei Zhu, Da-Han Wang, Xu-Yao Zhang, Cheng-Lin Liu

机构 * School of Advanced Interdisciplinary Sciences, UCAS(UCAS先进交叉科学学院) MAIS, CASIA(CASIA人工智能研究所) School of Artificial Intelligence, UCAS(UCAS人工智能学院) Centre for Artificial Intelligence and Robotics, HKISI-CAS(HKISI-CAS人工智能与机器人中心) FKLPRIU, School of Computer and Information Engineering, Xiamen University of Technology(厦门理工学院计算机与信息工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21079 2025-05-28 cs.CV 83%

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts

Yue Zhang, Yingzhao Jian, Hehe Fan, Yi Yang, Roger Zimmermann

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15576 2025-05-28 cs.RO cs.CV 83%

QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning

Xinyang Tong, Pengxiang Ding, Yiguo Fan, Donglin Wang, Wenjie Zhang, Can Cui, Mingyang Sun, Han Zhao, Hongyin Zhang, Yonghao Dang, Siteng Huang, Shangke Lyu

机构 * MiLAB, Westlake University, Hangzhou, 310030, China(西交利物浦大学微实验室,杭州,310030,中国) Zhejiang University, Hangzhou, 310027, China(浙江大学,杭州,310027,中国) Beijing University of Posts and Telecommunications, Beijing, 100876, China(北京邮电大学,北京,100876,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to ICRA 2025; Github page: https://quart-online.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16916 2025-05-23 cs.CR cs.CV 83%

Backdoor Cleaning without External Guidance in MLLM Fine-tuning

Xuankun Rong, Wenke Huang, Jian Liang, Jinhe Bi, Xun Xiao, Yiming Li, Bo Du, Mang Ye

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13828 2025-05-21 cs.AI 83%

Multimodal RAG-driven Anomaly Detection and Classification in Laser Powder Bed Fusion using Large Language Models

Kiarash Naghavi Khanghah, Zhiling Chen, Lela Romeo, Qian Yang, Rajiv Malhotra, Farhad Imani, Hongyi Xu

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments ASME 2025 International Design Engineering Technical Conferences and Computers and Information in Engineering Conference IDETC/CIE2025, August 17-20, 2025, Anaheim, CA (IDETC2025-168615)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12693 2025-05-20 cs.CV 83%

TACOcc:Target-Adaptive Cross-Modal Fusion with Volume Rendering for 3D Semantic Occupancy

Luyao Lei, Shuo Xu, Yifan Bai, Xing Wei

机构 * School of Software Engineering(软件工程学院) Xi’an Jiaotong University(西安交通大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12251 2025-05-20 cs.CV 83%

SMFusion: Semantic-Preserving Fusion of Multimodal Medical Images for Enhanced Clinical Diagnosis

Haozhe Xiang, Han Zhang, Yu Cheng, Xiongwen Quan, Wanwan Huang

机构 * College of Information and Intelligence, Hunan Agricultural University(湖南农业大学信息与智能学院) College of Artificial Intelligence, Nankai University(南开大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17141 2025-05-19 cs.CV 83%

Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation

Xu Zheng, Haiwei Xue, Jialei Chen, Yibo Yan, Lutao Jiang, Yuanhuiyi Lyu, Kailun Yang, Linfeng Zhang, Xuming Hu

机构 * HKUST(GZ)(香港科技大学(广州)) Tsinghua University(清华大学) Nagoya University(名古屋大学) Hunan University(湖南大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10105 2025-05-16 cs.RO cs.AI 83%

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation

Zibin Dong, Fei Ni, Yifu Yuan, Yinchuan Li, Jianye Hao

机构 * Tianjin University(天津大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03577 2025-05-09 cs.CV 83%

Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models

Xin Zou, Yizhou Wang, Yibo Yan, Yuanhuiyi Lyu, Kening Zheng, Sirui Huang, Junkai Chen, Peijie Jiang, Jia Liu, Chang Tang, Xuming Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19002 2025-04-29 cs.LG cs.CV cs.RO 83%

Deep Learning-Based Multi-Modal Fusion for Robust Robot Perception and Navigation

Delun Lai, Yeyubei Zhang, Yunchong Liu, Chaojie Li, Huadong Mo

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15624 2025-04-29 cs.CV 83%

FaceInsight: A Multimodal Large Language Model for Face Perception

Jingzhi Li, Changjiang Luo, Ruoyu Chen, Hua Zhang, Wenqi Ren, Jianhou Gan, Xiaochun Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区计算机科学与技术学院) Key Laboratory of Education Informatization for Nationalities (Yunnan Normal University), Ministry of Education(民族教育信息化重点实验室(云南师范大学))

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17261 2025-04-25 cs.LG cs.AI 83%

Symbolic Representation for Any-to-Any Generative Tasks

Jiaqi Chen, Xiaoye Zhu, Yue Wang, Tianyang Liu, Xinhui Chen, Ying Chen, Chak Tou Leong, Yifei Ke, Joseph Liu, Yiwen Yuan, Julian McAuley, Li-jia Li

机构 * Stanford University(斯坦福大学) Fellou AI(Fellou人工智能) Fudan University(复旦大学) South China University of Technology(华南理工大学) Cornell University(康奈尔大学) University of California San Diego(加州大学圣地亚哥分校) Wuhan University(武汉大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Hong Kong Polytechnic University(香港理工大学) University of Southern California(南加州大学) Carnegie Mellon University(卡内基梅隆大学) LiveX AI(LiveX人工智能)

专题命中 多模态训练与对齐 :any-to-any(title);multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09448 2025-04-23 cs.CV cs.LG 83%

Bayesian Cross-Modal Alignment Learning for Few-Shot Out-of-Distribution Generalization

Lin Zhu, Xinbing Wang, Chenghu Zhou, Nanyang Ye

专题命中 多模态训练与对齐 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted by AAAI2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13645 2025-04-21 cs.CV cs.LG 83%

Efficient Parameter Adaptation for Multi-Modal Medical Image Segmentation and Prognosis

Numan Saeed, Shahad Hardan, Muhammad Ridzuan, Nada Saadi, Karthik Nandakumar, Mohammad Yaqub

机构 * Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence(计算机视觉系,Mohamed bin Zayed人工智能大学) Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence(机器学习系,Mohamed bin Zayed人工智能大学) Michigan State University(密歇根州立大学) M42 Health(M42健康)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10538 2025-04-16 cs.IR cs.AI 83%

Distilling Transitional Pattern to Large Language Models for Multimodal Session-based Recommendation

Jiajie Su, Qiyong Zhong, Yunshan Ma, Weiming Liu, Chaochao Chen, Xiaolin Zheng, Jianwei Yin, Tat-Seng Chua

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02430 2025-04-11 cs.CV 83%

FOLDER: Accelerating Multi-modal Large Language Models with Enhanced Performance

Haicheng Wang, Zhemeng Yu, Gabriele Spadaro, Chen Ju, Victor Quétu, Shuai Xiao, Enzo Tartaglione

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15011 2025-04-08 cs.CV 83%

CrossOver: 3D Scene Cross-Modal Alignment

Sayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath, Iro Armeni

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Project Page: https://sayands.github.io/crossover/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23721 2025-04-01 cs.LG cs.AI 83%

Unimodal-driven Distillation in Multimodal Emotion Recognition with Dynamic Fusion

Jiagen Li, Rui Yu, Huihao Huang, Huaicheng Yan

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23456 2025-04-01 cs.CV 83%

CADFormer: Fine-Grained Cross-modal Alignment and Decoding Transformer for Referring Remote Sensing Image Segmentation

Maofu Liu, Xin Jiang, Xiaokang Zhang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏