arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2606.31394 2026-07-03 cs.LG cs.AI cs.CV q-bio.QM 新提交 76%

Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images

解决AI中的叠加问题以实现可解释性与患者-神经元图像的跨模态对齐

Jisung Park, Seohyeon Kang, Daeun Yoo, Eunsu Lee, Seoin Cho, Wooyeop Choi, Ian Choi, James R. Evan, Daesoo Kim, Sonia Gandhi, Minee L. Choi

机构 * KAIST(韩国科学技术院) Konyang University(建阳大学) Chang Gung University(长庚大学) UCL Queen Square Institute of Neurology & The Francis Crick Institute(伦敦大学学院皇后广场神经病学研究所与弗朗西斯·克里克研究所)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV、cs.AI

AI总结 利用稀疏自编码器解决高维生物数据中神经网络表示空间的叠加问题,恢复几何保真度,并通过Gromov-Wasserstein最优传输实现图像与单细胞RNA测序数据的跨模态对齐。

Comments 10 pages, 7 figures (plus 14 in appendix), 1 table, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03626 2026-06-03 cs.CV cs.AI cs.CY 76%

TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics

TurtleAI:海龟图形学中视觉编程的多模态模型基准测试

Chao Wen, Jacqueline Staub, Adish Singla

机构 * MPI-SWS(马克斯·普朗克研究所-斯图加特)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

AI总结 提出TurtleAI基准,包含823个基于海龟图形学真实任务的视觉编程任务,评估20多个多模态模型发现成功率低于30%,并通过少量种子样本生成合成数据微调Qwen2-VL-72B提升约20%性能。

Comments ACL Findings 2026 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07250 2026-05-11 cs.CV cs.AI 76%

Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment

难以阅读,却容易被破解:视觉退化如何绕过大语言模型的安全对齐

Zhixue Song, Boyan Han, Yiwei Wang, Chi Zhang

机构 * AGI Lab, Westlake University, China(西lake大学AGI实验室,中国) University of California, Merced, USA(加州大学默塞德分校,美国)

专题命中 多模态训练与对齐 :MLLM(title);分类 cs.CV、cs.AI

AI总结 研究发现视觉退化会削弱大语言模型的安全防御,提出结构化认知卸载策略以缓解风险。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06912 2026-04-09 cs.CV cs.AI 76%

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models

Q-Zoom:基于查询的自适应感知以实现高效多模态大语言模型

Yuheng Shi, Xiaohuan Pei, Linfeng Wen, Minjing Dong, Chang Xu

机构 * University of Sydney(悉尼大学) Sun Yat-sen University(中山大学) City University of Hong Kong(香港城市大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

AI总结 Q-Zoom通过高效粗到细的框架优化多模态大语言模型的高分辨率感知,提升推理速度并保持精度,实验表明其在文档识别和高分辨率场景中表现优异。

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05623 2026-03-09 cs.CV cs.AI 76%

Post Fusion Bird's Eye View Feature Stabilization for Robust Multimodal 3D Detection

融合后鸟瞰图特征稳定化用于鲁棒多模态3D检测

Trung Tien Dong, Dev Thakkar, Arman Sargolzaei, Xiaomin Lin

机构 * Department of Electrical Engineering, University of South Florida(佛罗里达州立大学电气工程系) Department of Mechanical Engineering, University of South Florida(佛罗里达州立大学机械工程系)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

AI总结 本文提出了一种轻量级模块PFS,通过稳定特征统计、抑制传感器退化区域和残差校正,提升多模态3D检测在域转移和传感器故障下的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21506 2025-12-29 cs.LG cs.AI cs.CL cs.HC 76%

MotionTeller: Multi-modal Integration of Wearable Time-Series with LLMs for Health and Behavioral Understanding

MotionTeller: 多模态整合可穿戴时间序列与大语言模型用于健康和行为理解

Aiwei Zhang, Arvind Pillai, Andrew Campbell, Nicholas C. Jacobson

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CL、cs.AI

AI总结 MotionTeller通过整合可穿戴时间序列与大语言模型,实现高精度的自然语言行为摘要生成,提升健康和行为理解的效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18245 2025-12-23 cs.CV cs.AI 76%

Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image

光谱偏差与跨模态语义一致性学习用于超光谱图像的目标检测

Xiao He, Chang Tang, Xinwang Liu, Wei Zhang, Zhimin Gao, Chuankun Li, Shaohua Qiu, Jiangfeng Xu

机构 * School of Computer, Wuhan University(武汉大学计算机学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) school of computer, National University of Defense Technology(国防科技大学计算机学院) Shandong Provincial Key Laboratory of Computer Networks, Shandong Computer Science Center (National Supercomputing Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(山东省计算机网络重点实验室、山东省计算机科学中心(国家超级计算中心济南中心)、齐鲁工业大学(山东省科学院)) School of Computer and Artificial Intelligence, Zhengzhou University(郑州大学计算机与人工智能学院) School of Information and Communication Engineering, North University of China(北方大学信息与通信工程学院) National Key Laboratory of Electromagnetic Energy, Naval University of Engineering(电磁能国家重点实验室、海军工程大学) Hexagon AB

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV、cs.AI

AI总结 本文提出SDCM网络,通过光谱偏差与跨模态语义一致性学习,提升超光谱图像目标检测的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01606 2025-10-03 cs.IR cs.AI cs.CL 76%

Bridging Collaborative Filtering and Large Language Models with Dynamic Alignment, Multimodal Fusion and Evidence-grounded Explanations

Bo Ma, LuYao Liu, Simon Lau, Chandler Yuan, and XueY Cui, Rosie Zhang

机构 * Department of Software \& Microelectronics, Peking University, Beijing, China Economic Law School, China University of Political Science Financial Media, Peking University, ChangSha, China

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24165 2025-09-30 cs.CV cs.AI 76%

LatXGen: Towards Radiation-Free and Accurate Quantitative Analysis of Sagittal Spinal Alignment Via Cross-Modal Radiographic View Synthesis

Moxin Zhao, Nan Meng, Jason Pui Yin Cheung, Chris Yuk Kwan Tang, Chenxi Yu, Wenting Zhong, Pengyu Lu, Chang Shi, Yipeng Zhuang, Teng Zhang

机构 * Department of Orthopaedics and Traumatology, The University of Hong Kong(香港大学骨科与创伤学系) Department of Joint Surgery, Shandong Provincial Hospital Affiliated to Shandong First Medical University(山东省第一医科大学附属山东省人民医院骨科)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV、cs.AI

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23442 2025-09-30 eess.IV cs.AI cs.CV cs.LG eess.SP 76%

S$^3$F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network

Md. Saiful Bari Siddiqui, Mohammed Imamul Hassan Bhuiyan

机构 * Department of Computer Science and Engineering, BRAC University(计算机科学与工程系,布拉克大学) Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology(电气与电子工程系,孟加拉国工程与技术大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments Submitted to IEEE Journal of Biomedical and Health Informatics (JBHI). This preprint includes few additional details not present in the journal submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15250 2025-09-23 cs.CV cs.AI 76%

Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruning

Wenda Qin, Andrea Burns, Bryan A. Plummer, Margrit Betke

机构 * Boston University(波士顿大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

Comments Accepted to EMNLP 2025. Data and code to be released at https://github.com/wdqin/VLN-NAP

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19668 2025-09-22 eess.SP cs.AI cs.CL cs.LG 76%

SuPreME: A Supervised Pre-training Framework for Multimodal ECG Representation Learning

Mingsheng Cai, Jiuming Jiang, Wenhao Huang, Che Liu, Rossella Arcucci

机构 * The University of Edinburgh(爱丁堡大学) Imperial College London(帝国理工学院) Shenzhen Yinwang Intelligent Technology Co., Ltd(深圳英伟达智能技术有限公司)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CL、cs.AI

Comments Findings of The 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09018 2025-08-13 cs.CV cs.CL cs.LG 76%

Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions

Moran Yanuka, Assaf Ben Kish, Yonatan Bitton, Idan Szpektor, Raja Giryes

机构 * Tel Aviv University(特拉维夫大学) Google Research(谷歌研究)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.CL

Comments Accepted to NAACL 2025

Journal ref Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics, Human Language Technologies, Long Papers, pp. 10497-10518

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02409 2025-08-05 cs.CV cs.AI 76%

Hydra: Accurate Multi-Modal Leaf Wetness Sensing with mm-Wave and Camera Fusion

Yimeng Liu, Maolin Gan, Huaili Zeng, Li Liu, Younsuk Dong, Zhichao Cao

机构 * Michigan State University(密歇根州立大学) Tsinghua University(清华大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments In Proceedings of ACM MobiCom (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19679 2025-07-29 cs.CV cs.AI 76%

Efficient Learning for Product Attributes with Compact Multimodal Models

Mandar Kulkarni

机构 * Flipkart Data Science(Flipkart数据科学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05020 2025-07-11 cs.CV cs.AI 76%

Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision

Soham Walimbe, Britty Baby, Vinkle Srivastav, Nicolas Padoy

机构 * University of Strasbourg, CNRS, INSERM, ICube, UMR7357, Strasbourg, France(斯特拉斯堡大学,法国国家科学研究中心(CNRS),法国国家卫生研究院(INSERM),ICube,UMR7357,斯特拉斯堡) Institute of Image-Guided Surgery, IHU Strasbourg, Strasbourg, France(影像引导手术研究所,斯特拉斯堡IHU,斯特拉斯堡)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09952 2025-06-12 cs.CV cs.AI 76%

UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting

Ziyi Wang, Yanran Zhang, Jie Zhou, Jiwen Lu

机构 * Department of Automation, Tsinghua University(自动化系,清华大学)

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV、cs.AI

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12425 2025-04-30 eess.IV cs.AI cs.CV q-bio.QM 76%

Multi-stage intermediate fusion for multimodal learning to classify non-small cell lung cancer subtypes from CT and PET

Fatih Aksu, Fabrizia Gelardi, Arturo Chiti, Paolo Soda

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

Journal ref Pattern Recognition Letters 193 (2025) 86-93

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16516 2025-04-28 cs.CV cs.AI 76%

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation

Junrong Yue, Yifan Zhang, Chuan Qin, Bo Li, Xiaomin Lie, Xinlei Yu, Wenxin Zhang, Zhendong Zhao

机构 * City University of Hong Kong, Dongguan Campus(香港城市大学东莞校区) The University of Melbourne(墨尔本大学) Tsinghua University(清华大学) Baidu Inc.(百度公司) University of Chinese Academy of Science(中国科学院大学) National University of Singapore(新加坡国立大学)

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments 11 pages, 4 figures, Submitted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18278 2025-04-01 cs.CV cs.AI 76%

TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model

Cheng Yang, Yang Sui, Jinqi Xiao, Lingyi Huang, Yu Gong, Chendi Li, Jinghua Yan, Yu Bai, Ponnuswamy Sadayappan, Xia Hu, Bo Yuan

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09838 2025-01-20 cs.CV cs.AI eess.IV 76%

CrossModalityDiffusion: Multi-Modal Novel View Synthesis with Unified Intermediate Representation

Alex Berian, Daniel Brignac, JhihYang Wu, Natnael Daba, Abhijit Mahalanobis

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments Accepted in the 2025 WACV workshop GeoCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02649 2025-01-07 cs.CV cs.AI 76%

Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features

Haixu Liu, Penghao Jiang, Zerui Tao, Muyan Wan, Qiuzhuang Sun

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments CVPR GeolifeCLEF

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00346 2025-01-03 cs.CV cs.AI cs.LG 76%

CNC: Cross-modal Normality Constraint for Unsupervised Multi-class Anomaly Detection

Xiaolei Wang, Xiaoyang Wang, Huihui Bai, Eng Gee Lim, Jimin Xiao

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV、cs.AI

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14424 2024-12-20 cs.CV cs.AI cs.LG 76%

FedPIA -- Permuting and Integrating Adapters leveraging Wasserstein Barycenters for Finetuning Foundation Models in Multi-Modal Federated Learning

Pramit Saha, Divyanshu Mishra, Felix Wagner, Konstantinos Kamnitsas, J. Alison Noble

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments Accepted for publication in AAAI 2025 (Main Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10452 2024-12-17 eess.IV cs.AI cs.CV 76%

Structurally Consistent MRI Colorization using Cross-modal Fusion Learning

Mayuri Mathur, Anav Chaudhary, Saurabh Kumar Gupta, Ojaswa Sharma

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV、cs.AI

Comments 9 pages, 6 figures, 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.04215 2024-10-28 cs.CV cs.MM 76%

Multimodal Engagement Analysis from Facial Videos in the Classroom

Ömer Sümer, Patricia Goldberg, Sidney D'Mello, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.MM

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.04449 2024-10-01 cs.CV cs.AI cs.LG 76%

Multi-modal Masked Siamese Network Improves Chest X-Ray Representation Learning

Saeed Shurrab, Alejandro Guerra-Manzanares, Farah E. Shamout

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

Comments Under review

Journal ref Scientific Reports 14 (2024) 22516

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09465 2024-08-20 cs.CV cs.AI 76%

MedMAP: Promoting Incomplete Multi-modal Brain Tumor Segmentation with Alignment

Tianyi Liu, Zhaorui Tan, Muyin Chen, Xi Yang, Haochuan Jiang, Kaizhu Huang

专题命中 多模态训练与对齐 :multi-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17620 2024-07-26 cs.CV cs.AI 76%

CoMoTo: Unpaired Cross-Modal Lesion Distillation Improves Breast Lesion Detection in Tomosynthesis

Muhammad Alberb, Marawan Elbatel, Aya Elgebaly, Ricardo Montoya-del-Angel, Xiaomeng Li, Robert Martí

专题命中 多模态训练与对齐 :cross-modal(title);分类 cs.CV、cs.AI

Comments ADSMI @ MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11996 2024-06-12 cs.LG cond-mat.mes-hall cond-mat.mtrl-sci cond-mat.soft cs.AI cs.CL 76%

Accelerating Scientific Discovery with Generative Knowledge Extraction, Graph-Based Representation, and Multimodal Intelligent Graph Reasoning

Markus J. Buehler

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏