arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6887 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6887 篇

2503.16069 2025-06-30 cs.CV 79%

Disentangled and Interpretable Multimodal Attention Fusion for Cancer Survival Prediction

Aniek Eijpe, Soufyan Lakbir, Melis Erdal Cesur, Sara P. Oliveira, Sanne Abeln, Wilson Silva

机构 * AI Technology for Life, Department of Information and Computing Sciences, Department of Biology, Utrecht University, Utrecht, The Netherlands(AI技术与生命、信息与计算科学系、生物学系、乌得勒支大学、乌得勒支、荷兰) Computational Pathology group, Department of Pathology, The Netherlands Cancer Institute, Amsterdam, The Netherlands(计算病理组、病理学系、荷兰癌症研究所、阿姆斯特丹、荷兰)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 1 figure, 3 tables. Preprint submitted and accepted to MICCAI 2025. This preprint has not undergone peer review or any post-submission improvements or corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21393 2025-06-27 cs.AI 79%

TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding

Junwen Zhang, Pu Chen, Yin Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 43 pages and 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21237 2025-06-27 cs.CV 79%

DiMPLe -- Disentangled Multi-Modal Prompt Learning: Enhancing Out-Of-Distribution Alignment with Invariant and Spurious Feature Separation

Umaima Rahman, Mohammad Yaqub, Dwarikanath Mahapatra

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学) Khalifa University(卡比拉大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21018 2025-06-27 cs.CV 79%

LASFNet: A Lightweight Attention-Guided Self-Modulation Feature Fusion Network for Multimodal Object Detection

Lei Hao, Lina Xu, Chang Liu, Yanni Dong

机构 * School of Geophysics and Geomatics, China University of Geosciences, Wuhan(地质物理与地质信息学院,中国地质大学(武汉)) School of Resource and Environmental Sciences, Wuhan University(资源与环境科学学院,武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19324 2025-06-25 cs.CV 79%

Memory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning

Mingcheng Qu, Guang Yang, Donglin Di, Yue Gao, Tonghua Su, Yang Song, Lei Fan

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) School of Software, Tsinghua University(清华大学软件学院) School of Computer Science and Engineering, UNSW Sydney(新南威尔士大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments accepted by MICCAI2025 code: https://github.com/MCPathology/M2Surv

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03507 2025-06-24 cs.CV 79%

Mineral segmentation using electron microscope images and spectral sampling through multimodal graph neural networks

Samuel Repka, Bořek Reich, Fedor Zolotarev, Tuomas Eerola, Pavel Zemčík

机构 * Lappeenranta-Lahti University of Technology(拉佩宁塔-拉赫蒂技术大学) Brno University of Technology - Faculty of Information Technology(布拉格技术大学-信息科技学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16879 2025-06-24 eess.IV cs.CV 79%

Ultra-high resolution multimodal MRI densely labelled holistic structural brain atlas

José V. Manjón, Sergio Morell-Ortega, Marina Ruiz-Perez, Boris Mansencal, Edern Le Bot, Marien Gadea, Enrique Lanuza, Gwenaelle Catheline, Thomas Tourdias, Vincent Planche, Rémi Giraud, Denis Rivière, Jean-François Mangin, Nicole Labra-Avila, Roberto Vivo-Hernando, Gregorio Rubio, Fernando Aparici, Maria de la Iglesia-Vaya, Pierrick Coupé

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17202 2025-06-23 cs.CV 79%

UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation

Teng Li, Quanfeng Lu, Lirui Zhao, Hao Li, Xizhou Zhu, Yu Qiao, Jun Zhang, Wenqi Shao

机构 * HKUST(香港科技大学) Shanghai AI Laboratory(上海人工智能实验室) SJTU(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Code: https://github.com/tliby/UniFork

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17136 2025-06-23 cs.CV 79%

Semi-Supervised Multi-Modal Medical Image Segmentation for Complex Situations

Dongdong Meng, Sheng Li, Hao Wu, Guoping Wang, Xueqing Yan

机构 * School of Physics, Peking University, Beijing, China(北京大学物理系) School of Computer Science, Peking University, Beijing, China(北京大学计算机系) Department of Radiotherapy, Peking University Cancer Hospital, Beijing, China(北京大学肿瘤医院放疗科)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 2 figures, accepted at MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04443 2025-06-23 cs.AI 79%

POV Learning: Individual Alignment of Multimodal Models using Human Perception

Simon Werner, Katharina Christ, Laura Bernardy, Marion G. Müller, Achim Rettinger

机构 * Trier University(特里尔大学) University of Innsbruck(因斯布鲁克大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01124 2025-06-23 eess.IV cs.CV 79%

Cross-modality Attention Adapter: A Glioma Segmentation Fine-tuning Method for SAM Using Multimodal Brain MR Images

Xiaoyu Shi, Shurong Chai, Yinhao Li, Jingliang Cheng, Jie Bai, Guohua Zhao, Yen-Wei Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18014 2025-06-19 cs.AI 79%

Coupled Mamba: Enhanced Multi-modal Fusion with Coupled State Space Model

Wenbing Li, Hang Zhou, Junqing Yu, Zikai Song, Wei Yang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10794 2025-06-18 cs.CV 79%

Distraction is All You Need for Multimodal Large Language Model Jailbreaking

Zuopeng Yang, Jiluan Fan, Anli Yan, Erdun Gao, Xin Lin, Tao Li, Kanghua Mo, Changyu Dong

机构 * Guangzhou University(广州大学) Shanghai Jiao Tong University(上海交通大学) Australian Institute for Machine Learning, The University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2025 highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13705 2025-06-17 cs.LG cs.AI 79%

TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning

Junru Zhang, Lang Feng, Xu Guo, Yuhan Wu, Yabo Dong, Duanqing Xu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00942 2025-06-17 cs.CL 79%

Experiential Semantic Information and Brain Alignment: Are Multimodal Models Better than Language Models?

Anna Bavaresco, Raquel Fernández

机构 * Institute for Logic, Language and Computation University of Amsterdam(逻辑、语言与计算研究所 阿姆斯特丹大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to CoNLL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16282 2025-06-16 cs.LG cs.AI 79%

Understanding the Emergence of Multimodal Representation Alignment

Megan Tjandrasuwita, Chanakya Ekbote, Liu Ziyin, Paul Pu Liang

机构 * Massachusetts Institute of Technology, USA(麻省理工学院) NTT Research, USA(NTT研究)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments To appear as a poster in ICML 2025. 21 pages, 22 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11454 2025-06-13 cs.LG cs.AI q-bio.QM 79%

Elucidating the Design Space of Multimodal Protein Language Models

Cheng-Yen Hsieh, Xinyou Wang, Daiheng Zhang, Dongyu Xue, Fei Ye, Shujian Huang, Zaixiang Zheng, Quanquan Gu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments ICML 2025 Spotlight; Project Page: https://bytedance.github.io/dplm/dplm-2.1/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08716 2025-06-11 eess.IV cs.CV 79%

Enhancing Synthetic CT from CBCT via Multimodal Fusion: A Study on the Impact of CBCT Quality and Alignment

Maximilian Tschuchnig, Lukas Lamminger, Philipp Steininger, Michael Gadermayr

机构 * Salzburg University of Applied Sciences(萨尔茨堡应用科学大学) MedPhoton GmbH(MedPhoton公司) University of Salzburg(萨尔茨堡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Data is open source. Code will be provided on acceptance. Paper currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07077 2025-06-10 cs.CR cs.AI 79%

Dual-Priv Pruning : Efficient Differential Private Fine-Tuning in Multimodal Large Language Models

Qianshan Wei, Jiaqi Li, Zihan You, Yi Zhan, Kecen Li, Jialin Wu, Xinfeng Li Hengjun Liu, Yi Yu, Bin Cao, Yiwen Xu, Yang Liu, Guilin Qi

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06730 2025-06-10 cs.CR cs.AI 79%

Fuse and Federate: Enhancing EV Charging Station Security with Multimodal Fusion and Federated Learning

Rabah Rahal, Abdelaziz Amara Korba, Yacine Ghamri-Doudane

机构 * Rabah Rahal Networks and Systems Laboratory (LRS), Badji Mokhtar Annaba University, Algeria(拉比·拉尔实验室,巴吉·莫克塔尔安纳巴大学,阿尔及利亚) Abdelaziz Amara korba Networks and Systems Laboratory (LRS), Badji Mokhtar Annaba University, Algeria(阿卜杜勒阿齐兹·阿马拉·科尔巴实验室,巴吉·莫克塔尔安纳巴大学,阿尔及利亚) L3I, University of La Rochelle, France(L3I,拉罗歇尔大学,法国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05890 2025-06-09 cs.CV 79%

Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation

Yiheng Li, Yang Yang, Zichang Tan, Huan Liu, Weihua Chen, Xu Zhou, Zhen Lei

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) CAIR, HKISI, Chinese Academy of Sciences(中国科学院计算机辅助研究部) School of Computer Science and Engineering, the Faculty of Innovation Engineering, M.U.S.T(慕斯科技大学计算机科学与工程学院) Sangfor Technologies Inc.(Sangfor技术有限公司) Beijing Jiaotong University(北京交通大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13088 2025-06-09 cs.CV cs.LG 79%

Cross-modal feature fusion for robust point cloud registration with ambiguous geometry

Zhaoyi Wang, Shengyu Huang, Jemil Avers Butt, Yuanzhou Cai, Matej Varga, Andreas Wieser

机构 * ETH Zürich, Institute of Geodesy and Photogrammetry(苏黎世联邦理工学院测绘与摄影测量研究所) Atlas optimization GmbH(Atlas优化公司) University of Zürich(苏黎世大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments To appear in the ISPRS Journal of Photogrammetry and Remote Sensing. 19 pages, 14 figures

Journal ref ISPRS J. Photogramm. Remote Sens. 227 (2025) 31-47

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15334 2025-06-09 cs.CV 79%

Modality-Fair Preference Optimization for Trustworthy MLLM Alignment

Songtao Jiang, Yan Zhang, Ruizhe Chen, Tianxiang Hu, Yeying Jin, Qinglin He, Yang Feng, Jian Wu, Zuozhu Liu

机构 * Zhejiang University(浙江大学) ByteDance(字节跳动) National University of Singapore(新加坡国立大学) Angelalign Inc., China(中国Angelalign公司)

专题命中 多模态训练与对齐 :MLLM(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07987 2025-06-06 cs.AI 79%

Universal Adversarial Attack on Aligned Multimodal LLMs

Temurbek Rahmatullaev, Polina Druzhinina, Nikita Kurdiukov, Matvey Mikhalchuk, Andrey Kuznetsov, Anton Razzhigaev

机构 * AIRI MSU(莫斯科国立大学) HSE University(俄罗斯高等经济大学) Skoltech(斯克里普钦科技大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Added benchmarks, baselines, author, appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03675 2025-06-05 cs.CV 79%

BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation

Jialei Chen, Xu Zheng, Danda Pani Paudel, Luc Van Gool, Hiroshi Murase, Daisuke Deguchi

机构 * Graduate School of Informatics, Nagoya University(名古屋大学信息学研究科) AI Thrust, The Hong Kong University of Science and Technology(香港科学与技术大学人工智能 thrust) INSAIT, Sofia University(索菲亚大学INSAIT)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01456 2025-06-03 q-bio.GN cs.AI cs.LG q-bio.NC 79%

GenDMR: A dynamic multimodal role-swapping network for identifying risk gene phenotypes

Lina Qin, Cheng Zhu, Chuqi Zhou, Yukun Huang, Jiayi Zhu, Ping Liang, Jinju Wang, Yixing Huang, Cheng Luo, Dezhong Yao, Ying Tan

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 31 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01097 2025-06-03 cs.CV 79%

Generic Token Compression in Multimodal Large Language Models from an Explainability Perspective

Lei Lei, Jie Gu, Xiaokang Ma, Chu Tang, Jingmin Chen, Tong Xu

机构 * University of Science and Technology of China(中国科学技术大学) Rightly Robotics(Rightly机器人公司) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00813 2025-06-03 cs.CV cs.LG 79%

TIME: TabPFN-Integrated Multimodal Engine for Robust Tabular-Image Learning

Jiaqi Luo, Yuan Yuan, Shixin Xu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00466 2025-06-03 eess.AS cs.SD 79%

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

Cunhang Fan, Ying Chen, Jian Zhou, Zexu Pan, Jingjing Zhang, Youdian Gao, Xiaoke Yang, Zhengqi Wen, Zhao Lv

机构 * School of Computer Science and Technology, Anhui University(安徽大学计算机科学与技术学院) Alibaba group(阿里巴巴集团) Department of Automation, Tsinghua University(清华大学自动化系) Beijing National Research Center for lnformation Science and Technology, Tsinghua University(清华大学信息科学与技术国家研究中心)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 eess.AS

Comments Accepted to IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00365 2025-06-03 cs.CV eess.SP 79%

Feature Fusion and Knowledge-Distilled Multi-Modal Multi-Target Detection

Ngoc Tuyen Do, Tri Nhu Do

机构 * School of Information and Communications, Hanoi University of Science and Technology(信息与通信学院,河内科学技术大学) Telecom Neural Detection Lab, Polytechnique Montréal(电信神经检测实验室,蒙特利尔理工学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏