arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6856 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6856 篇

2508.07871 2025-12-10 cs.CV 85%

CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning

CATP:面向高效增强多模态上下文学习的上下文自适应令牌剪枝

Yanshu Li, Jianjiang Yang, Zhennan Shen, Ligong Han, Haoyan Xu, Ruixiang Tang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 CATP通过上下文自适应令牌剪枝方法,提升多模态上下文学习的效率和性能,减少冗余令牌带来的影响。

Comments 14 pages, 12 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14997 2025-12-09 cs.CV cs.LG 85%

Language Integration in Fine-Tuning Multimodal Large Language Models for Image-Based Regression

基于图像回归的多模态大语言模型微调中的语言整合

Roy H. Jennings, Genady Paikin, Roy Shaul, Evgeny Soloveichik

机构 * Samsung Israel R&D Center(三星以色列研发中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出RvTC方法,通过灵活的区间法替代传统词汇受限分类,提升多模态大语言模型在图像回归任务中的性能。

Comments WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18541 2025-12-05 cs.AI 85%

Align$^2$LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation

Align$^2$LLaVA: 多模态指令编纂的级联人类与大语言模型偏好对齐

Hongzhe Huang, Jiang Liu, Zhewen Yu, Li Cai, Dian Jiao, Wenqiao Zhang, Siliang Tang, Juncheng Li, Hao Jiang, Haoyuan Li, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Alibaba(阿里巴巴)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);MLLM(abstract);分类 cs.AI

AI总结 Align$^2$LLaVA通过级联人类与LLM偏好对齐方法,有效压缩多模态指令数据,提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09060 2025-10-29 cs.LG cs.AI q-bio.GN 85%

Multimodal 3D Genome Pre-training

Minghao Yang, Pengteng Li, Yan Liang, Qianyi Cai, Zhihang Zheng, Shichen Zhang, Pengfei Zhang, Zhi-An Huang, Hui Xiong

机构 * Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou), China(人工智能研究所,香港科技大学(广州)) School of Artificial Intelligence, South China Normal University, China(人工智能学院,华南师范大学) Thrust of Bioscience and Biomedical Engineering, The Hong Kong University of Science and Technology (Guangzhou), China(生物科学与生物医学工程研究所,香港科技大学(广州)) Department of Computer Science, City University of Hong Kong (Dongguan), China(计算机科学系,香港城市大学(东莞)) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR, China(计算机科学与工程系,香港科技大学香港特别行政区)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15948 2025-10-21 cs.AI cs.CR 85%

VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search

MingSheng Li, Guangze Zhao, Sichen Liu

机构 * Independent Researcher(独立研究者) Harbin Institute of Technology(哈尔滨工业大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15653 2025-08-25 cs.CV 85%

MapKD: Unlocking Prior Knowledge with Cross-Modal Distillation for Efficient Online HD Map Construction

Ziyang Yan, Ruikai Li, Zhiyong Cui, Bohan Li, Han Jiang, Yilong Ren, Aoyong Li, Zhenning Li, Sijia Wen, Haiyang Yu

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13387 2025-08-20 cs.AI 85%

SPANER: Shared Prompt Aligner for Multimodal Semantic Representation

Thye Shan Ng, Caren Soyeon Han, Eun-Jung Holden

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02479 2025-08-05 cs.CV 85%

Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding

Xinquan Yu, Wei Lu, Xiangyang Luo

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12945 2025-07-18 cs.CV 85%

Analysis of Image-and-Text Uncertainty Propagation in Multimodal Large Language Models with Cardiac MR-Based Applications

Yucheng Tang, Yunguan Fu, Weixi Yi, Yipei Wang, Daniel C. Alexander, Rhodri Davies, Yipeng Hu

机构 * Department of Medical Physics and Biomedical Engineering, University College London, UK(医学物理与生物医学工程系,伦敦大学学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

Comments It is accepted by 28th International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07424 2025-07-11 cs.CV 85%

Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning

Jingjing Jiang, Chao Ma, Xurui Song, Hanwang Zhang, Jun Luo

机构 * Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23224 2025-06-30 cs.CL 85%

MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence Calibration

Zhitao He, Sandeep Polisetty, Zhiyuan Fan, Yuchen Huang, Shujin Wu, Yi R. Fung

机构 * Hong Kong University of Science and Technology(香港科技大学) UMass Amherst(马萨诸塞大学阿姆赫斯特分校) University of Southern California(南加州大学)

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CL

Comments 18 pages, ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11878 2025-06-10 cs.CL cs.CE q-fin.CP 85%

Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications

Jimin Huang, Mengxi Xiao, Dong Li, Zihao Jiang, Yuzhe Yang, Yifei Zhang, Lingfei Qian, Yan Wang, Xueqing Peng, Yang Ren, Ruoyu Xiang, Zhengyu Chen, Xiao Zhang, Yueru He, Weiguang Han, Shunian Chen, Lihang Shen, Daniel Kim, Yangyang Yu, Yupeng Cao, Zhiyang Deng, Haohang Li, Duanyu Feng, Yongfu Dai, VijayaSai Somasundaram, Peng Lu, Guojun Xiong, Zhiwei Liu, Zheheng Luo, Zhiyuan Yao, Ruey-Ling Weng, Meikang Qiu, Kaleb E Smith, Honghai Yu, Yanzhao Lai, Min Peng, Jian-Yun Nie, Jordan W. Suchow, Xiao-Yang Liu, Benyou Wang, Alejandro Lopez-Lira, Qianqian Xie, Sophia Ananiadou, Junichi Tsujii

机构 * The Fin AI(Fin AI) Wuhan University(武汉大学) Columbia University(哥伦比亚大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Nanjing University(南京大学) Rensselaer Polytechnic Institute(罗格斯理工学院) The University of Manchester(曼彻斯特大学) Stevens Institute of Technology(史泰斯理工学院) National University of Singapore(新加坡国立大学) University of Florida(佛罗里达大学) University of Montreal(蒙特利尔大学) Yale University(耶鲁大学) New York University(纽约大学) Harvard University(哈佛大学) NVIDIA(英伟达) Artificial Intelligence Research Centre(人工智能研究中心) Archimedes/Athena Research Centre(Archimedes/Athena研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments 33 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10917 2025-05-20 cs.CV 85%

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Mingxiao Li, Na Su, Fang Qu, Zhizhou Zhong, Ziyang Chen, Yuan Li, Zhaopeng Tu, Xiaolong Li

机构 * Tencent(腾讯)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02441 2025-05-06 cs.AI 85%

MSFNet-CPD: Multi-Scale Cross-Modal Fusion Network for Crop Pest Detection

Jiaqi Zhang, Zhuodong Liu, Kejian Yu

机构 * College of Computer and Information Science, Southwest University(西南大学计算机与信息科学学院) School of Economics and Management, Beijing Jiaotong University(北京交通大学经济管理学院) School of Computer Science and Technology, Donghua University(东华大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);image-text(abstract);分类 cs.AI

Comments Accepted to IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10479 2025-04-22 cs.CV 85%

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Xuehui Wang, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, Hao Li, Jiahao Wang, Nianchen Deng, Songze Li, Yinan He, Tan Jiang, Jiapeng Luo, Yi Wang, Conghui He, Botian Shi, Xingcheng Zhang, Wenqi Shao, Junjun He, Yingtong Xiong, Wenwen Qu, Peng Sun, Penglong Jiao, Han Lv, Lijun Wu, Kaipeng Zhang, Huipeng Deng, Jiaye Ge, Kai Chen, Limin Wang, Min Dou, Lewei Lu, Xizhou Zhu, Tong Lu, Dahua Lin, Yu Qiao, Jifeng Dai, Wenhai Wang

机构 * Shanghai AI Laboratory(上海人工智能实验室) SenseTime Research(商汤科技研究院) Tsinghua University(清华大学) Nanjing University(南京大学) Fudan University(复旦大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19546 2025-04-17 cs.CV 85%

MMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training

Biao Wu, Yutong Xie, Zeyu Zhang, Minh Hieu Phan, Qi Chen, Ling Chen, Qi Wu

专题命中 多模态训练与对齐 :cross-modal(title);multimodal(abstract);multi-modal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10351 2025-04-15 cs.CV 85%

Multimodal Representation Learning Techniques for Comprehensive Facial State Analysis

Kaiwen Zheng, Xuri Ge, Junchen Fu, Jun Peng, Joemon M. Jose

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV

Comments Accepted by ICME2025

Journal ref ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15166 2025-04-15 cs.CV cs.AI cs.CL cs.LG cs.MM 85%

Machine Unlearning in Hyperbolic vs. Euclidean Multimodal Contrastive Learning: Adapting Alignment Calibration to MERU

Àlex Pujol Vidal, Sergio Escalera, Kamal Nasrollahi, Thomas B. Moeslund

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19425 2025-03-25 cs.CV 85%

Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment

Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Sanath Narayan, Ankit Singh, Noel E. O'Connor

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted CVPR 2025; First two authors contributed equally;

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09498 2025-03-13 cs.LG cs.CV 85%

Towards Robust Multimodal Representation: A Unified Approach with Adaptive Experts and Alignment

Nazanin Moradinasab, Saurav Sengupta, Jiebei Liu, Sana Syed, Donald E. Brown

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09698 2025-01-14 cs.IR cs.AI 85%

Harnessing Multimodal Large Language Models for Multimodal Sequential Recommendation

Yuyang Ye, Zhi Zheng, Yishan Shen, Tianshu Wang, Hengruo Zhang, Peijun Zhu, Runlong Yu, Kai Zhang, Hui Xiong

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14520 2025-01-09 cs.CV 85%

Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

Han Zhao, Min Zhang, Wei Zhao, Pengxiang Ding, Siteng Huang, Donglin Wang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted to the Thirty-Ninth AAAI Conference on Artificial Intelligence (AAAI-25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20455 2024-12-31 cs.CV 85%

Cross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection

Ayush Ghadiya, Purbayan Kar, Vishal Chudasama, Pankaj Wasnik

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Accepted to CVPR'24 MULA Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.10994 2024-12-19 cs.CL cs.AI cs.CV cs.MM 85%

Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

Dingjie Song, Wenjun Wang, Shunian Chen, Xidong Wang, Michael Guan, Benyou Wang

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to COLING 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04521 2024-10-08 cs.CV 85%

MC-CoT: A Modular Collaborative CoT Framework for Zero-shot Medical-VQA with LLM and MLLM Integration

Lai Wei, Wenkai Wang, Xiaoyu Shen, Yu Xie, Zhihao Fan, Xiaojin Zhang, Zhongyu Wei, Wei Chen

专题命中 多模态训练与对齐 :MLLM(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 21 pages, 14 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19130 2024-10-01 eess.IV cs.AI cs.LG 85%

Multi-modal Cross-domain Self-supervised Pre-training for fMRI and EEG Fusion

Xinxu Wei, Kanhao Zhao, Yong Jiao, Nancy B. Carlisle, Hua Xie, Gregory A. Fonzo, Yu Zhang

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13979 2024-06-21 eess.IV cs.CV cs.LG 85%

Knowledge-driven Subspace Fusion and Gradient Coordination for Multi-modal Learning

Yupei Zhang, Xiaofei Wang, Fangliangzi Meng, Jin Tang, Chao Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.14065 2024-04-18 cs.CV 85%

AsymFormer: Asymmetrical Cross-Modal Representation Learning for Mobile Platform Real-Time RGB-D Semantic Segmentation

Siqi Du, Weixi Wang, Renzhong Guo, Ruisheng Wang, Yibin Tian, Shengjun Tang

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01735 2024-04-03 cs.IR cs.MM 85%

CIRP: Cross-Item Relational Pre-training for Multimodal Product Bundling

Yunshan Ma, Yingzhi He, Wenjun Zhong, Xiang Wang, Roger Zimmermann, Tat-Seng Chua

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.MM

Comments arXiv preprint, 10 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10331 2023-11-20 eess.IV cs.CV 85%

Leveraging Multimodal Fusion for Enhanced Diagnosis of Multiple Retinal Diseases in Ultra-wide OCTA

Hao Wei, Peilun Shi, Guitao Bai, Minqing Zhang, Shuangle Li, Wu Yuan

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏