arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2512.07687 2025-12-09 cs.CL cs.CV 73%

HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs

HalluShift++: 通过内部表示转移弥合语言与视觉,解决多模态大语言模型中的层级幻觉

Sujoy Nath, Arkaprabha Basu, Sharanya Dasgupta, Swagatam Das

机构 * Netaji Subhash Engineering College (NSEC)(奈尔贾伊·萨布哈工程学院) TCG Crest Electronics and Communication Sciences Unit (ECSU)(电子与通信科学单位) Indian Statistical Institute(印度统计研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

AI总结 HalluShift++通过分析MLLM内部表示转移,解决多模态大语言模型中的层级幻觉问题,提升幻觉检测的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07170 2025-12-09 cs.CV cs.AI 73%

Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach

迈向统一的语义和可控图像融合:一种扩散变换器方法

Jiayang Li, Chengjie Jiang, Junjun Jiang, Pengwei Liang, Jiayi Ma, Liqiang Nie

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Electronic Information School, Wuhan University(武汉大学电子信息学院)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 DiTFuse通过融合图像与自然语言指令,实现端到端、语义感知的图像融合,统一了多种融合任务并在多个基准测试中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01558 2025-11-25 cs.CV cs.AI 73%

VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval

VideoLights: 用于联合视频亮点检测和时刻检索的特征细化与跨任务对齐变换器

Dhiman Paul, Md Rizwan Parvez, Nabeel Mohammed, Shafin Rahman

机构 * North South University(北南大学) Qatar Computing Research Institute (QCRI)(卡塔尔计算研究 institute)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 VideoLights通过引入特征细化、跨模态融合和联合任务反馈机制,提升视频亮点检测与时刻检索的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11126 2025-11-17 cs.CL cs.CV 73%

Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion

Yi Shi, Wenlong Meng, Zhenyuan Guo, Chengkun Wei, Wenzhi Chen

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05553 2025-11-11 cs.CV cs.AI 73%

EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning

Xinyan Cai, Shiguang Wu, Dafeng Chi, Yuzheng Zhuang, Xingyue Quan, Jianye Hao, Qiang Guan

机构 * Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00801 2025-11-11 cs.CV cs.AI cs.RO 73%

Environment-Driven Online LiDAR-Camera Extrinsic Calibration

Zhiwei Huang, Jiaqi Li, Hongbo Zhao, Xiao Ma, Ping Zhong, Xiaohu Zhou, Wei Ye, Rui Fan

机构 * Department of Control Science & Engineering, the College of Electronic & Information Engineering, Tongji University(控制科学与工程系,电子与信息工程学院,同济大学) School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学) Beijing Institute of Aerospace Control Devices(北京航天控制器件研究所) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04601 2025-11-07 cs.CV cs.MM 73%

PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning

Yicheng Xiao, Yu Chen, Haoxuan Ma, Jiale Hong, Caorui Li, Lingxiang Wu, Haiyun Guo, Jinqiao Wang

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02228 2025-11-05 cs.CV cs.AI 73%

Collaborative Attention and Consistent-Guided Fusion of MRI and PET for Alzheimer's Disease Diagnosis

Delin Ma, Menghui Zhou, Jun Qi, Yun Yang, Po Yang

机构 * School of Software Yunnan University(软件学院 云南大学) Department of Computer Science The University of Sheffield(计算机科学系 剑桥大学) Department of Computing Xian JiaoTong-Liverpool University(计算系 西交利物浦大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01463 2025-11-04 cs.CV cs.AI cs.GR 73%

HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA

Lei Hu, Yongjing Ye, Shihong Xia

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 5figures. The Thirty-Ninth Annual Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01284 2025-11-04 cs.CV cs.AI 73%

Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions

Karma Phuntsho, Abdullah, Kyungmi Lee, Ickjai Lee, Euijoon Ahn

机构 * James Cook University(詹姆斯库克大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01082 2025-11-04 cs.CV cs.AI cs.LG 73%

GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction

Narges Ghasemi, Amir Ziashahabi, Salman Avestimehr, Cyrus Shahabi

机构 * of Computer Science, University of Southern California, Los Angeles, CA, USA Computer Engineering, University of Southern California, Los Angeles, CA, USA

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted to IEEE International Conference on Data Mining (ICDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21501 2025-10-27 cs.CV cs.AI 73%

GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs

Guanghao Zheng, Bowen Shi, Mingxing Xu, Ruoyu Sun, Peisen Zhao, Zhibo Zhang, Wenrui Dai, Junni Zou, Hongkai Xiong, Xiaopeng Zhang, Qi Tian

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Inc.(华为公司)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14497 2025-10-03 cs.CV cs.CL 73%

Efficient Whole Slide Pathology VQA via Token Compression

Weimin Lyu, Qingqiao Hu, Kehan Qi, Zhan Shi, Wentao Huang, Saumya Gupta, Chao Chen

机构 * Stony Brook University(石溪大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00032 2025-10-02 eess.SP cs.AI cs.CL cs.LG q-bio.NC 73%

WaveMind: Towards a Conversational EEG Foundation Model Aligned to Textual and Visual Modalities

Ziyi Zeng, Zhenyang Cai, Yixi Cai, Xidong Wang, Junying Chen, Rongsheng Wang, Yipeng Liu, Siqi Cai, Benyou Wang, Zhiguo Zhang, Haizhou Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22853 2025-09-30 q-bio.QM cs.AI cs.CL cs.LG 73%

Patient-specific Biomolecular Instruction Tuning

Irsyad Adam, Zekai Chen, David Laub, Shaun Porwal, Arda Pekis, Kevin Brown

机构 * Standard Model Biomedicine(标准模型生物医学) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17747 2025-09-23 cs.CV cs.AI 73%

Dual-View Alignment Learning with Hierarchical-Prompt for Class-Imbalance Multi-Label Classification

Sheng Huang, Jiexuan Yan, Beiyan Liu, Bo Liu, Richang Hong

机构 * Ministry of Education Key Laboratory of Dependable Service Computing in Cyber Physical Society(教育部可信服务计算网络社会重点实验室) School of Big Data and Software Engineering(大数据与软件工程学院) School of Computer Science and Information Engineering(计算机科学与信息工程学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments accepted by IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16892 2025-09-23 cs.CV cs.AI 73%

Learning from Gene Names, Expression Values and Images: Contrastive Masked Text-Image Pretraining for Spatial Transcriptomics Representation Learning

Jiahe Qian, Yaoyu Fang, Ziqiao Weng, Xinkun Wang, Lee A. Cooper, Bo Zhou

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15667 2025-09-22 cs.CL cs.SD eess.AS 73%

VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion

Dimitrios Damianos, Leon Voukoutis, Georgios Paraskevopoulos, Vassilis Katsouros

机构 * Institute for Speech and Language Processing, Athena Research Center, Greece(语音与语言处理研究所,亚特兰蒂斯研究中心,希腊)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14067 2025-09-22 cs.CV cs.AI 73%

VLA-Mark: A cross modal watermark for large vision-language alignment model

Shuliang Liu, Qi Zheng, Jesse Jiaxi Xu, Yibo Yan, Junyan Zhang, He Geng, Aiwei Liu, Peijie Jiang, Jia Liu, Yik-Cheung Tam, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) University of Toronto(多伦多大学) Ant Group, Alibaba(蚂蚁集团,阿里巴巴) New York University Shanghai(纽约大学上海分校)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by the main conference, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04139 2025-09-09 cs.CV cs.AI cs.ET cs.LG cs.RO 73%

Driver-Net: Multi-Camera Fusion for Assessing Driver Take-Over Readiness in Automated Vehicles

Mahdi Rezaei, Mohsen Azarmi

机构 * Institute for Transport Studies, Computer Vision and Machine Learning Group, University of Leeds(交通研究 institute,计算机视觉和机器学习组,莱斯特大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Journal ref 2025 IEEE Intelligent Vehicles Symposium (IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17524 2025-08-26 cs.CV cs.AI 73%

OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation

Xingxin He, Aurora Rofena, Ruimin Feng, Haozhe Liao, Zhaoye Zhou, Albert Jang, Fang Liu

机构 * Athinoula A. Martinos Center for Biomedical Imaging(阿提诺拉A.马丁诺斯生物医学成像中心) Harvard Medical School(哈佛医学院) Massachusetts General Hospital(麻省总医院) University Campus Bio-Medico of Rome(罗马生物医学大学校园)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04107 2025-08-20 cs.CV cs.AI 73%

Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder

Jingchao Wang, Zhijian Wu, Dingjiang Huang, Yefeng Zheng, Hong Wang

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06566 2025-08-12 cs.CV cs.AI 73%

Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features

Manish Kansana, Elias Hossain, Shahram Rahimi, Noorbakhsh Amiri Golilarz

机构 * Department of Computer Science and Engineering, Mississippi State University(计算机科学与工程系,密苏里州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04449 2025-08-07 cs.CV cs.CL 73%

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay

Jun Zhang, Desen Meng, Zhengming Zhang, Zhenpeng Huang, Tao Wu, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) China Mobile Research Institute(中国移动研究院) Shanghai AI Lab(上海AI实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICCV 2025; Code released at https://github.com/MCG-NJU/p-MoD

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10887 2025-07-29 cs.CV cs.AI 73%

Point Cloud Self-supervised Learning via 3D to Multi-view Masked Learner

Zhimin Chen, Xuewei Chen, Xiao Guo, Yingwei Li, Longlong Jing, Liang Yang, Bing Li

机构 * Clemson University(克莱姆森大学) Michigan State University(密歇根州立大学) Johns Hopkins University(约翰霍普金斯大学) The City University of New York(纽约城市大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11129 2025-07-18 cs.CV cs.AI cs.LG 73%

MMOne: Representing Multiple Modalities in One Scene

Zhifeng Gu, Bing Wang

机构 * Spatial Intelligence Group, The Hong Kong Polytechnic University(香港理工大学空间智能组)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11661 2025-07-17 cs.CL cs.AI 73%

Partitioner Guided Modal Learning Framework

Guimin Hu, Yi Xin, Lijie Hu, Zhihong Zhu, Hasti Seifi

机构 * Guangdong University of Technology(广东工业大学) University of Copenhagen(哥本哈根大学) Nanjing University(南京大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Tencent(腾讯) Arizona State University(亚利桑那州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments acm multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04291 2025-06-23 cs.CL cs.CV 73%

Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models

Saketh Bachu, Erfan Shayegani, Rohit Lal, Trishna Chakraborty, Arindam Dutta, Chengyu Song, Yue Dong, Nael Abu-Ghazaleh, Amit K. Roy-Chowdhury

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICML 2025 as a spotlight poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16933 2025-06-05 cs.LG cs.CL cs.CV 73%

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Zebin You, Shen Nie, Xiaolu Zhang, Jun Hu, Jun Zhou, Zhiwu Lu, Ji-Rong Wen, Chongxuan Li

机构 * Gaoling School of AI, Renmin University of China(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理研究重点实验室) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心,教育部) Ant Group(蚂蚁集团)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments Project page and codes: \url{https://ml-gsai.github.io/LLaDA-V-demo/}

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14562 2025-05-21 cs.SD cs.MM eess.AS 73%

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

Parthasaarathy Sudarsanam, Irene Martín-Morató, Tuomas Virtanen

机构 * Audio Research Group, Tampere University(塔尔皮莱大学音频研究组)

专题命中 多模态训练与对齐 :multimodal(abstract);audio-visual(abstract);分类 cs.MM、eess.AS

Comments Accepted to European Signal Processing Conference (EUSIPCO 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏