arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2408.05914 2025-10-22 cs.CV 79%

Learning Collaborative Knowledge with Multimodal Representation for Polyp Re-Identification

Suncheng Xiang, Jiale Guan, Shilun Cai, Jiacheng Ruan, Dahong Qian

机构 * Shanghai Jiao Tong University(上海交通大学) Zhongshan Hospital of Fudan University(复旦大学中山医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17394 2025-10-21 cs.LG cs.CV 79%

MILES: Modality-Informed Learning Rate Scheduler for Balancing Multimodal Learning

Alejandro Guerra-Manzanares, Farah E. Shamout

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted and presented at the 2025 International Joint Conference on Neural Networks (IJCNN'25). The paper was awarded an honorable mention (best 4 papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17289 2025-10-21 cs.CL 79%

Addressing Antisocial Behavior in Multi-Party Dialogs Through Multimodal Representation Learning

Hajar Bakarou, Mohamed Sinane El Messoussi, Anaïs Ollagnier

机构 * Universit\'e C \ te d'Azur, CNRS, Inria, I3S Sophia Antipolis France Universit\'e C \ te d'Azur, CNRS, Inria, I3S

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17078 2025-10-21 cs.CV 79%

Towards a Generalizable Fusion Architecture for Multimodal Object Detection

Jad Berjawi, Yoann Dupas, Christophe C'erin

机构 * Université Grenoble Alpes(格勒诺布尔大学) Université Sorbonne Paris Nord(巴黎-萨克勒大学) INRIA(法国国家信息与自动化研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 8 figures, accepted at ICCV 2025 MIRA Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08217 2025-10-21 cs.AI 79%

Quantum Federated Learning for Multimodal Data: A Modality-Agnostic Approach

Atit Pokharel, Ratun Rahman, Thomas Morris, Dinh C. Nguyen

机构 * Department of Electrical and Computer Engineering, The University of Alabama in Huntsville(电气与计算机工程系,阿拉巴马大学亨茨维尔分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments This paper was presented at BEAM with CVPR 2025

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 545-554. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15944 2025-10-21 cs.LG cs.AI 79%

Lyapunov-Stable Adaptive Control for Multimodal Concept Drift

Tianyu Bell Pan, Mengdi Zhu, Alexa Jordyn Cole, Ronald Wilson, Damon L. Woodard

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Florida Institute of National Security(佛罗里达国家安全研究所) Applied Artificial Intelligence Group(应用人工智能组) University of Florida(佛罗里达大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21514 2025-10-21 cs.CV 79%

G$^{2}$D: Boosting Multimodal Learning with Gradient-Guided Distillation

Mohammed Rakib, Arunkumar Bagavathi

机构 * Oklahoma State University(俄克拉荷马州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15317 2025-10-20 cs.AI 79%

VERITAS: Leveraging Vision Priors and Expert Fusion to Improve Multimodal Data

Tingqiao Xu, Ziru Zeng, Jiayu Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15026 2025-10-20 cs.CV 79%

MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning

Mattia Segu, Marta Tintore Gazulla, Yongqin Xian, Luc Van Gool, Federico Tombari

机构 * Google(谷歌) ETH Zurich(苏黎世联邦理工学院) INSAIT, Sofia University, St. Kliment Ohridski(INSAIT,索菲亚大学,圣克莱孟·奥赫里茨基)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14344 2025-10-17 cs.CR cs.AI 79%

BinCtx: Multi-Modal Representation Learning for Robust Android App Behavior Detection

Zichen Liu, Shao Yang, Xusheng Xiao

机构 * Arizona State University(亚利桑那州立大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20401 2025-10-17 cs.CV cs.RO 79%

SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment

Binod Singh, Sayan Deb Sarkar, Iro Armeni

机构 * Technical University of Munich(慕尼黑技术大学) Stanford University(斯坦福大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Project Page: https://singhbino3d.github.io/sgpp/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12245 2025-10-15 cs.LG cs.AI 79%

MoRA: On-the-fly Molecule-aware Low-Rank Adaptation Framework for LLM-based Multi-Modal Molecular Assistant

Tao Yin, Xiaohong Zhang, Jiacheng Zhang, Li Huang, Zhibin Zhang, Yuansong Zeng, Jin Xie, Meng Yan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02575 2025-10-15 cs.CL cs.CR cs.LG 79%

Cross-Modal Safety Alignment: Is textual unlearning all you need?

Trishna Chakraborty, Erfan Shayegani, Zikui Cai, Nael Abu-Ghazaleh, M. Salman Asif, Yue Dong, Amit K. Roy-Chowdhury, Chengyu Song

机构 * University of California, Riverside(加州大学河滨分校)

专题命中 多模态训练与对齐 :cross-modal(title);multi-modal(abstract);分类 cs.CL

Comments Accepted by EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11112 2025-10-14 cs.CV 79%

Multimodal Disease Progression Modeling via Spatiotemporal Disentanglement and Multiscale Alignment

Chen Liu, Wenfang Yao, Kejing Yin, William K. Cheung, Jing Qin

机构 * School of Nursing, The Hong Kong Polytechnic University(香港理工大学护理学院) Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10524 2025-10-14 cs.CV 79%

Unified Open-World Segmentation with Multi-Modal Prompts

Yang Liu, Yufei Yin, Chenchen Jing, Muzhi Zhu, Hao Chen, Yuling Xi, Bo Feng, Hao Wang, Shiyu Li, Chunhua Shen

机构 * Zhejiang University(浙江大学) Hangzhou Dianzi University(杭州电子科技大学) Zhejiang University of Technology(浙江工业大学) Apple(苹果公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06538 2025-10-14 cs.CL 79%

Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model

Xinyue Lou, You Li, Jinan Xu, Xiangyu Shi, Chi Chen, Kaiyu Huang

机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education(大数据与人工智能交通联合实验室(北京交通大学)) School of Computer Science and Technology, Beijing Jiaotong University(计算机科学与技术学院,北京交通大学) Tsinghua University(清华大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07979 2025-10-13 cs.CV 79%

Visual Representation Alignment for Multimodal Large Language Models

Heeji Yoon, Jaewoo Jung, Junwan Kim, Hyungyu Choi, Heeseong Shin, Sangbeom Lim, Honggyu An, Chaehyun Kim, Jisang Han, Donghyun Kim, Chanho Eom, Sunghwan Hong, Seungryong Kim

机构 * KAIST AI(韩国科学技术院人工智能研究所) New York University(纽约大学) Chung-Ang University(Chung-Ang 大学) Korea University(韩国大学) ETH Zürich(苏黎世联邦理工学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Project Page: https://cvlab-kaist.github.io/VIRAL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08457 2025-10-10 cs.CL 79%

ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping

Shuang Chen, Yue Guo, Yimeng Ye, Shijue Huang, Wenbo Hu, Haoxi Li, Manyuan Zhang, Jiayu Chen, Song Guo, Nanyun Peng

机构 * University of California, Los Angeles(加州大学洛杉矶分校) The Hong Kong University of Science and Technology(香港科技大学) Columbia University(哥伦比亚大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06938 2025-10-09 quant-ph cs.AI 79%

Expressive and Scalable Quantum Fusion for Multimodal Learning

Tuyen Nguyen, Trong Nghia Hoang, Phi Le Nguyen, Hai L. Vu, Truong Cong Thang

机构 * University of Technology Sydney(悉尼技术大学) The University of Aizu(御所大学) Washington State University(华盛顿州立大学) Hanoi University of Science and Technology(河内科学技术大学) Monash University(莫纳什大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 22 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05184 2025-10-08 cs.AI 79%

Representation Potentials of Foundation Models for Multimodal Alignment: A Survey

Jianglin Lu, Hailing Wang, Yi Xu, Yizhou Wang, Kuo Yang, Yun Fu

机构 * Department of Electrical and Computer Engineering, Northeastern University(东北大学电气与计算机工程系) Khoury College of Computer Science, Northeastern University(东北大学科赫里计算机科学学院)

专题命中 多模态训练与对齐 :multimodal(title);cross-modal(abstract);分类 cs.AI

Journal ref The 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03232 2025-10-06 cs.CV 79%

LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models

Ci-Siang Lin, Min-Hung Chen, Yu-Yang Sheng, Yu-Chiang Frank Wang

机构 * Graduate Institute of Communication Engineering, National Taiwan University, Taiwan(台湾国立台湾大学通信工程研究所) NVIDIA

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01780 2025-10-03 cs.CR cs.AI cs.CY cs.LG 79%

Secure Multi-Modal Data Fusion in Federated Digital Health Systems via MCP

Aueaphum Aueawatthanaphisut

机构 * School of Information, Computer, and Communication Technology(信息、计算机与通信技术学院) Sirindhorn International Institute of Technology, Thammasat University(泰国朱拉安吞国际技术学院,泰国 Thammasat 大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments 6 pages, 8 figures, 7 equations, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00894 2025-10-02 cs.AI 79%

FusionAdapter for Few-Shot Relation Learning in Multimodal Knowledge Graphs

Ran Liu, Yuan Fang, Xiaoli Li

机构 * Singapore Management University(新加坡管理大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Archived paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00652 2025-10-02 cs.CV 79%

OTTER: Open-Tagging via Text-Image Representation for Multi-modal Understanding

Jieer Ouyang, Xiaoneng Xiang, Zheng Wang, Yangkai Ding

机构 * Huawei Singapore Research Center(华为新加坡研究中心)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at ICDM 2025 BigIS Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00515 2025-10-02 cs.CV 79%

Efficient Multi-modal Large Language Models via Progressive Consistency Distillation

Zichen Wen, Shaobo Wang, Yufa Zhou, Junyuan Zhang, Qintong Zhang, Yifeng Gao, Zhaorun Chen, Bin Wang, Weijia Li, Conghui He, Linfeng Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12299 2025-10-01 cs.CR cs.AI 79%

QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety

Taegyeong Lee, Jeonghwa Yoo, Hyoungseo Cho, Soo Yong Kim, Yunho Maeng

机构 * FnGuide Inc.(FnGuide公司) Safe Generative AI Lab, MODULABS(MODULABS安全生成AI实验室) A.I.MATICS Inc.(A.I.MATICS公司) Ewha Womans University(成均馆大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

Comments Accept to ACLW 2025 (WOAH); fix typo

Journal ref ACL Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25037 2025-09-30 cs.CL 79%

GateMABSA: Aspect-Image Gated Fusion for Multimodal Aspect-based Sentiment Analysis

Adamu Lawan, Haruna Yunusa

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments 6 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23677 2025-09-30 cs.CV 79%

MSD-KMamba: Bidirectional Spatial-Aware Multi-Modal 3D Brain Segmentation via Multi-scale Self-Distilled Fusion Strategy

Dayu Tan, Ziwei Zhang, Yansan Su, Xin Peng, Yike Dai, Chunhou Zheng, Weimin Zhong

机构 * Key Laboratory of Intelligent Computing and Signal Processing, Ministry of Education, Anhui University(智能计算与信号处理重点实验室,教育部,安徽大学) Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education, East China University of Science and Technology(能源化工过程智能制造重点实验室,教育部,东华大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19904 2025-09-30 cs.RO cs.MM eess.SP 79%

WildFusion: Multimodal Implicit 3D Reconstructions in the Wild

Yanbaihui Liu, Boyuan Chen

机构 * Duke University(杜克大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.MM

Comments Our project website is at: http://generalroboticslab.com/WildFusion

Journal ref 2025 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21268 2025-09-26 cs.CV 79%

MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources

Sicong Leng, Jing Wang, Jiaxi Li, Hao Zhang, Zhiqiang Hu, Boqiang Zhang, Yuming Jiang, Hang Zhang, Xin Li, Lidong Bing, Deli Zhao, Wei Lu, Yu Rong, Aixin Sun, Shijian Lu

机构 * Nanyang Technological University(南洋理工大学) DAMO Academy, Alibaba Group(阿里巴巴集团达摩院) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏