arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2508.06146 2025-08-11 cs.CV 70%

Text-guided Visual Prompt DINO for Generic Segmentation

Yuchen Guan, Chong Sun, Canmiao Fu, Zhipeng Huang, Chun Yuan, Chen Li

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) WeChat AI, Tencent Inc.(微信AI,腾讯公司)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05213 2025-08-08 cs.CV 70%

Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation

Jianming Liu, Wenlong Qiu, Haitao Wei

机构 * School of Artificial Intelligence(人工智能学院) School of Digital Industry(数字产业学院) Jiangxi Normal University(江西师范大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 10 pages,Accepted at ACMMM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01150 2025-08-05 cs.CV 70%

OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding

Dianyi Yang, Xihan Wang, Yu Gao, Shiyang Liu, Bohan Ren, Yufeng Yue, Yi Yang

机构 * School of Automation, Beijing Institute of Technology, Beijing, China(自动化学院,北京理工大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments IROS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02331 2025-08-05 cs.CV cs.SD 70%

VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection

Hao Cheng, Zhiwei Zhao, Yichao He, Zhenzhen Hu, Jia Li, Meng Wang, Richang Hong

机构 * Hefei University of Technology(合肥工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Source code and pre-trained models will be available at https://github.com/MSA-LMC/VAEmo

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19903 2025-08-05 cs.CV 70%

Scaling Vision Pre-Training to 4K Resolution

Baifeng Shi, Boyi Li, Han Cai, Yao Lu, Sifei Liu, Marco Pavone, Jan Kautz, Song Han, Trevor Darrell, Pavlo Molchanov, Hongxu Yin

机构 * UC Berkeley(加州大学伯克利分校) NVIDIA(英伟达)

专题命中 多模态训练与对齐 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments CVPR 2025. Project Page: https://nvlabs.github.io/PS3

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19442 2025-07-31 cs.AI cs.DC 70%

A Survey on Large Language Model Acceleration based on KV Cache Management

Haoyang Li, Yiming Li, Anxin Tian, Tianhao Tang, Zhanchao Xu, Xuejia Chen, Nicole Hu, Wei Dong, Qing Li, Lei Chen

机构 * Department of Computing, The Hong Kong Polytechnic University(计算系,香港理工大学) Department of Computer Science and Engineering(计算机科学与工程系) Department of Computer Science and Technology(计算机科学与技术系) Department of Computing and Data Science(计算与数据科学系)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

Comments Accepted to TMLR 2025. The revised version incorporates more papers and has been further polished

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13803 2025-07-21 cs.CV 70%

GRAM-MAMBA: Holistic Feature Alignment for Wireless Perception with Adaptive Low-Rank Compensation

Weiqi Yang, Xu Zhou, Jingfu Guan, Hao Du, Tianyu Bai

机构 * Hefei University of Technology(合肥工业大学) Sun Yat-sen University(中山大学) Jilin University(吉林大学) Durham University(杜伦大学)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09081 2025-07-15 cs.CV 70%

From Physics to Foundation Models: A Review of AI-Driven Quantitative Remote Sensing Inversion

Zhenyu Yu, Mohd Yamani Idna Idris, Hua Wang, Pei Wang, Junyi Chen, Kun Wang

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05934 2025-07-09 cs.AI 70%

BlueLM-2.5-3B Technical Report

Baojiao Xiong, Boheng Chen, Chengzhi Wang, Daxiong Luo, Dongsheng Xu, Dongyang Liu, Fan Yang, Fangyuan Li, Fei Teng, Feng Wang, Fukang Qin, Fuquan Peng, Guanxin Tan, Guozhi Wang, Haibo Yu, Haohao Gao, Heng Liu, Hongbo Yang, Hongjian Zou, Houzheng Shen, Hu Meng, Huan Li, Hui Tan, Jiali Chen, Jianzhao Chen, Jinliang Zhu, Kai Wang, Lei Wu, Liangbing Liu, Liuyang Bian, Liyan He, Long Liu, Peiwen Li, Penggang Shi, Qi Ding, Rui Hu, Shuai Cao, Shuai Ren, Shuang Peng, Teng Xie, Weiji Chen, Weilin Xiang, Weixin Wu, Xi Yin, Xiaoxin Chen, Xu Chen, Yafei Wen, Yan Hu, Yanzhou Yang, Yina Xie, Yinghao Chen, Yixuan Liao, Yu Geng, Yuanjiang Ouyang, Yuanzhuo Yang, Yuehua He, Yushuai Peng, Zhaoxiong Wang, Zheng Wang, Zhibo Zhou, Ziyang Wu

机构 * vivo AI Lab(vivo人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04999 2025-07-08 cs.CV 70%

Robust Incomplete-Modality Alignment for Ophthalmic Disease Grading and Diagnosis via Labeled Optimal Transport

Qinkai Yu, Jianyang Xie, Yitian Zhao, Cheng Chen, Lijun Zhang, Liming Chen, Jun Cheng, Lu Liu, Yalin Zheng, Yanda Meng

机构 * University of Exeter(埃克塞特大学) University of Liverpool(利物浦大学) Chinese Academy of Sciences(中国科学院) The University of Hong Kong(香港大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) Dalian University of Technology(大连理工大学) A*STAR

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06002 2025-07-04 cs.CV 70%

Task-Adapter++: Task-specific Adaptation with Order-aware Alignment for Few-shot Action Recognition

Congqi Cao, Peiheng Han, Yueran zhang, Yating Yu, Qinyi Lv, Lingtong Min, Yanning zhang

机构 * School of Computer Science, Northwestern Polytechnical University(计算机科学学院,西北工业大学) School of Electronics and Information, Northwestern Polytechnical University(电子与信息学院,西北工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments extended work of Task-Adapter

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13337 2025-07-03 cs.CV 70%

Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens

Qihang Fan, Huaibo Huang, Mingrui Chen, Ran He

机构 * MAIS & NLPR, Institute of Automation, Chinese Academy of Sciences, Beijing, China(自动化研究所,中国科学院,北京) School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学人工智能学院,北京)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.24124 2025-07-02 cs.LG cs.CV 70%

Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives

Sixun Dong, Wei Fan, Teresa Wu, Yanjie Fu

机构 * Arizona State University(亚利桑那州立大学) University of Oxford(牛津大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Code: https://github.com/Ironieser/TimesCLIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21803 2025-06-30 eess.SP cs.AI cs.LG 70%

From Token to Rhythm: A Multi-Scale Approach for ECG-Language Pretraining

Fuying Wang, Jiacheng Xu, Lequan Yu

机构 * School of Computing and Data Science, The University of Hong Kong, Hong Kong SAR, China(计算与数据科学学院,香港大学,香港特别行政区,中国)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18943 2025-06-25 cs.CV 70%

From Pixels and Words to Waves: A Unified Framework for Spectral Dictionary vLLMs

Andrew Kiruluta, Priscilla Burity

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14696 2025-06-19 cs.CV 70%

YOLOv11-RGBT: Towards a Comprehensive Single-Stage Multispectral Object Detection Framework

Dahang Wan, Rongsheng Lu, Yang Fang, Xianli Lang, Shuangbao Shu, Jingjing Chen, Siyuan Shen, Ting Xu, Zecong Ye

机构 * School of Instrument Science and Opto-electronics Engineering, Anhui Province Key Laboratory of Measuring Theory and Precision Instrument(仪器科学与光电工程学院、安徽省测量理论与精密仪器重点实验室) Hefei University of Technology(合肥工业大学) School of Information Engineering(信息工程学院) Engineering University of PAP(工程大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 29 pages, 8 figures . The errors in the first version have been corrected, and no new version will be submitted in the near future. The next version will include more experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11412 2025-06-16 cs.CY cs.AI 70%

The Strategic Imperative for Healthcare Organizations to Build Proprietary Foundation Models

Naresh Tiwari

专题命中 多模态训练与对齐 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03601 2025-06-13 cs.CV 70%

GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Yu Liu, Kai Chen, Ping Luo

机构 * The University of Hong Kong(香港大学) Shanghai AI Laboratory(上海人工智能实验室) Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments ECCV2024-Workshop, Camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07214 2025-06-10 cs.CV cs.CR 70%

Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation

Zhiyuan Zhong, Zhen Sun, Yepang Liu, Xinlei He, Guanhong Tao

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07086 2025-06-10 cs.CL 70%

Representation Decomposition for Learning Similarity and Contrastness Across Modalities for Affective Computing

Yuanhe Tian, Pengsen Cheng, Guoqing Jin, Lei Zhang, Yan Song

机构 * University of Washington(华盛顿大学) Sichuan University(四川大学) People’s Daily Online(人民日报网络) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03555 2025-06-05 cs.CV 70%

WIFE-Fusion:Wavelet-aware Intra-inter Frequency Enhancement for Multi-model Image Fusion

Tianpei Zhang, Jufeng Zhao, Yiming Zhu, Guangmang Cui

机构 * National Natural Science Foundation of China(中华人民共和国国家自然科学基金委员会)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19036 2025-06-02 cs.CV 70%

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs

Hongliang Li, Jiaxin Zhang, Wenhui Liao, Dezhi Peng, Kai Ding, Lianwen Jin

机构 * South China University of Technology(华南理工大学) Intsig Information Co., Ltd.(Intsig信息有限公司) INTSIG-SCUT Joint Lab on Document Analysis and Recognition(INTSIG-SCUT文档分析与识别联合实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04405 2025-05-27 cs.IR cs.AI 70%

Universal Item Tokenization for Transferable Generative Recommendation

Bowen Zheng, Hongyu Lu, Yu Chen, Wayne Xin Zhao, Ji-Rong Wen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14471 2025-05-20 cs.CV 70%

Integrating Extra Modality Helps Segmentor Find Camouflaged Objects Well

Chengyu Fang, Chunming He, Longxiang Tang, Yuelin Zhang, Chenyang Zhu, Yuqi Shen, Chubin Chen, Guoxia Xu, Xiu Li

机构 * Tsinghua University(清华大学) Duke University(杜克大学) The Chinese University of Hong Kong(香港中文大学) Nanjing University of Posts and Telecommunications(南京邮电大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 18 pages, 8 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10088 2025-05-16 cs.CV 70%

MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models

Yuncheng Guo, Xiaodong Gu

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Due to the limitation "The abstract field cannot be longer than 1,920 characters", the abstract appearing here is slightly shorter than that in the PDF file

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06920 2025-05-13 cs.CV 70%

Bi-directional Self-Registration for Misaligned Infrared-Visible Image Fusion

Timing Li, Bing Cao, Pengfei Zhu, Bin Xiao, Qinghua Hu

机构 * College of Intelligence and Computing(智能与计算学院) Tianjin University(天津大学) School of Computer Science and Technology(计算机科学与技术学校) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18856 2025-04-29 cs.CV 70%

Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation

Shahad Albastaki, Anabia Sohail, Iyyakutti Iyappan Ganapathi, Basit Alawode, Asim Khan, Sajid Javed, Naoufel Werghi, Mohammed Bennamoun, Arif Mahmood

机构 * Department of Computer Science(计算机科学系) ARIC Khalifa University of Science and Technology(科技大学) Information Technology University of the Punjab(旁遮普信息科技大学) University of the Western Australia(西澳大学)

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04616 2025-04-22 cs.CV 70%

Assessing and Learning Alignment of Unimodal Vision and Language Models

Le Zhang, Qian Yang, Aishwarya Agrawal

机构 * Mila - Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) Canada CIFAR AI Chair(加拿大CIFAR人工智能 chair)

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments CVPR 2025 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00654 2025-04-02 cs.CV 70%

QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA

Shuai Li, Jian Xu, Xiao-Hui Li, Chao Deng, Lin-Lin Huang

专题命中 多模态训练与对齐 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16149 2025-03-21 eess.IV cs.CV 70%

Selective Complementary Feature Fusion and Modal Feature Compression Interaction for Brain Tumor Segmentation

Dong Chen, Boyue Zhao, Yi Zhang, Meng Zhao

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏