arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4651 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2506.05439 2025-09-22 cs.CV cs.AI cs.CL 67%

LLMs Can Compensate for Deficiencies in Visual Representations

Sho Takishita, Jay Gala, Abdelrahman Mohamed, Kentaro Inui, Yova Kementchedjhieva

机构 * Fujitsu Limited(富士通有限公司) MBZUAI Tohoku University(东北大学) RIKEN(日本研究机构)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23759 2025-09-18 cs.CL cs.AI cs.CV cs.LG 67%

Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint

Heekyung Lee, Jiaxin Ge, Tsung-Han Wu, Minwoo Kang, Trevor Darrell, David M. Chan

机构 * POSTECH University of California, Berkeley(加州大学伯克利分校)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16146 2025-09-16 cs.CV cs.AI cs.CL cs.LG 67%

Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation

Zhenglin Hua, Jinghan He, Zijun Yao, Tianxu Han, Haiyun Guo, Yuheng Jia, Junfeng Fang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University)(东南大学新一代人工智能技术及其交叉应用关键实验室) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Wuhan University of Technology(武汉理工大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17243 2025-09-03 cs.CV cs.AI cs.CL 67%

CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models

Zicong Tang, Ziyang Ma, Suqing Wang, Zuchao Li, Lefei Zhang, Hai Zhao, Yun Li, Qianren Wang

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) School of Computer Science, Wuhan University(武汉大学计算机学院) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机学院) Cognitive AI Lab, Shanghai Huawei Technologies, China(上海华为技术有限公司认知人工智能实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02865 2025-08-13 eess.IV cs.AI cs.CL cs.CV 67%

VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge

Zihan Li, Diping Song, Zefeng Yang, Deming Wang, Fei Li, Xiulan Zhang, Paul E. Kinahan, Yu Qiao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Washington(华盛顿大学) Shenzhen Institutes of Advanced Technology(深圳先进技术研究所) Chinese Academy of Sciences(中国科学院) State Key Laboratory of Ophthalmology(眼科学国家重点实验室) Zhongshan Ophthalmic Center(中山眼科中心) Sun Yat-sen University(中山大学) Guangdong Provincial Key Laboratory of Ophthalmology and Visual Science(广东省眼科学与视觉科学重点实验室) Guangdong Provincial Clinical Research Center for Ocular Diseases(广东省眼科临床研究中心) Department of Bioengineering(生物工程系) Department of Radiology(放射科)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by IEEE TPAMI, 14 pages, 15 tables, 4 figures with Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19795 2025-07-25 cs.CL cs.AI cs.CV 67%

VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks

Juhwan Choi, Junehyoung Kwon, JungMin Yun, Seunguk Yu, YoungBin Kim

机构 * AITRICS Seoul(AITRICS首尔) Chung-Ang University(Chung-ang 大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICCV 2025 Workshop on Curated Data for Efficient Learning (CDEL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01790 2025-07-03 cs.CL cs.AI cs.CV cs.LG 67%

How Do Vision-Language Models Process Conflicting Information Across Modalities?

Tianze Hua, Tian Yun, Ellie Pavlick

机构 * Brown University(布朗大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments All code and resources are available at: https://github.com/ethahtz/vlm_conflicting_info_processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19303 2025-06-25 cs.RO 67%

Robotic Perception with a Large Tactile-Vision-Language Model for Physical Property Inference

Zexiang Guo, Hengxiang Chen, Xinheng Mai, Qiusang Qiu, Gan Ma, Zhanat Kappassov, Qiang Li, Nutan Chen

机构 * College of Big Data and Internet, Shenzhen Technology University, China(大数据与互联网学院,深圳科技大学,中国) Sino-German College of Intelligent Manufacturing, Shenzhen Technology University, China(中德智能制造学院,深圳科技大学,中国) Robotics Department, Institute of Smart Systems and Artificial Intelligence (ISSAI), Nazarbayev University, Kazakhstan(机器人系,智能系统与人工智能研究所(ISSAI),纳扎尔巴耶夫大学,哈萨克斯坦) Foundation Robotics Labs, Germany(基础机器人实验室,德国)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract)

Comments This paper has been accepted by the 2025 International Conference on Climbing and Walking Robots (CLAWAR). These authors contributed equally to this work: Zexiang Guo, Hengxiang Chen, Xinheng Mai

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15734 2025-06-23 cs.AI cs.CL cs.CR cs.CV cs.LG 67%

The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models

Peiyuan Tang, Haojie Xin, Xiaodong Zhang, Jun Sun, Qin Xia, Zijiang Yang

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院) School of Computing and Information Systems, Singapore Management University(新加坡管理学院计算与信息系统学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 23 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14100 2025-06-18 cs.RO cs.SY eess.SY 67%

A Hierarchical Test Platform for Vision Language Model (VLM)-Integrated Real-World Autonomous Driving

Yupeng Zhou, Can Cui, Juntong Peng, Zichong Yang, Juanwu Lu, Jitesh H Panchal, Bin Yao, Ziran Wang

机构 * Purdue University(普渡大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05429 2025-06-09 cs.CV cs.AI cs.CL cs.LG 67%

Coordinated Robustness Evaluation Framework for Vision-Language Models

Ashwin Ramesh Babu, Sajad Mousavi, Vineet Gundecha, Sahand Ghorbanpour, Avisek Naug, Antonio Guillen, Ricardo Luna Gutierrez, Soumyendu Sarkar

机构 * Hewlett Packard Enterprise (Hewlett Packard Labs)(惠普企业公司(惠普实验室))

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted: IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05399 2025-06-09 cs.CV cs.AI cs.CL 67%

Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation

Israa A. Albadarneh, Bassam H. Hammo, Omar S. Al-Kadi

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 31 pages, 15 figures, 6 tables

Journal ref Computer Science Review, Vol. 58, pp. 100766, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15037 2025-06-04 cs.LG 67%

Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

Huanyu Zhang, Chengzu Li, Wenshan Wu, Shaoguang Mao, Yifan Zhang, Haochen Tian, Ivan Vulić, Zhang Zhang, Liang Wang, Tieniu Tan, Furu Wei

机构 * Microsoft Research(微软研究院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Language Technology Lab, University of Cambridge(剑桥大学语言技术实验室) Nanjing University(南京大学)

专题命中 图文多模态 :multimodal(abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13928 2025-06-04 cs.CV cs.AI cs.CL cs.LG 67%

Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images

Shengguang Wu, Fan-Yun Sun, Kaiyue Wen, Nick Haber

机构 * Stanford University(斯坦福大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ACL 2025 Main. Project Website: https://s-vco.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22045 2025-05-29 cs.MM cs.CV cs.SD eess.AS 67%

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Le Xu, Chenxing Li, Yong Ren, Yujie Chen, Yu Gu, Ruibo Fu, Shan Yang, Dong Yu

机构 * Tencent AI Lab(腾讯AI实验室) Institute of Automation, Chinese Academy of Sciences(自动化研究所)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted by INTERSPEECH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19091 2025-05-27 cs.CL cs.AI cs.CV cs.LG 67%

ReadBench: Measuring the Dense Text Visual Reading Ability of Vision-Language Models

Benjamin Clavié, Florian Brand

机构 * Artificial Intelligence and Intelligent Information Systems, University of Trier(人工智能与智能信息系统,特里尔大学) German Research Center for Artificial Intelligence (DFKI), Trier(德国人工智能研究中心(DFKI),特里尔)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13487 2025-05-23 cs.CL cs.AI cs.CV cs.LG 67%

Transferring Textual Preferences to Vision-Language Understanding through Model Merging

Chen-An Li, Tzu-Han Lin, Yun-Nung Chen, Hung-yi Lee

机构 * National Taiwan University(台湾大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04568 2025-05-20 cs.CV cs.AI cs.CL cs.LG 67%

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision

Giorgio Giannone, Ruoteng Li, Qianli Feng, Evgeny Perevodchikov, Rui Chen, Aleix Martinez

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11221 2025-05-19 cs.LG 67%

Sample Efficient Reinforcement Learning via Large Vision Language Model Distillation

Donghoon Lee, Tung M. Luu, Younghwan Lee, Chang D. Yoo

机构 * Robotics Program KAIST(韩国釜山科学技术院机器人计划) Electrical Engineering KAIST(韩国釜山科学技术院电子工程)

专题命中 图文多模态 :multimodal(abstract);multimodal foundation model(abstract)

Comments 5 pages, ICASSP 2025. The first two authors are equally contributed

Journal ref ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09656 2025-05-16 q-bio.QM 67%

VIGIL: Vision-Language Guided Multiple Instance Learning Framework for Ulcerative Colitis Histological Healing Prediction

Zhengxuan Qiu, Bo Peng, Xiaoying Tang, Jiankun Wang, Qin Guo

专题命中 图文多模态 :multi-modal(abstract);image-text(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06869 2025-05-06 cs.LG cs.AI cs.CL cs.CV 67%

Impact of Noisy Supervision in Foundation Model Learning

Hao Chen, Zihan Wang, Ran Tao, Hongxin Wei, Xing Xie, Masashi Sugiyama, Bhiksha Raj, Jindong Wang

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 18 pages, 10 figures, 6 tables, preprint. arXiv admin note: substantial text overlap with arXiv:2309.17002

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01790 2025-05-06 cs.CV cs.CL cs.MM 67%

Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos

Markos Stamatakis, Joshua Berger, Christian Wartena, Ralph Ewerth, Anett Hoppe

机构 * TIB – Leibniz Information Centre for Science and Technology(蒂宾根-莱比锡信息科学与技术研究中心) Hochschule Hannover – Data(汉诺威高等学院-数据) H Institute for Applied Data Science(汉诺威应用数据科学研究所) L3S Research Center – Leibniz University Hannover(L3S研究中心-汉诺威莱比锡大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments 12 pages (excluding references), 8 tables, 1 equation

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09738 2025-04-15 cs.CV cs.AI cs.LG cs.MM 67%

Automatic Detection of Intro and Credits in Video using CLIP and Multihead Attention

Vasilii Korolkov, Andrey Yanchenko

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 22 pages, 11 figures, submitted as a preprint. ArXiv preprint only, not submitted to a journal yet

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15485 2025-04-09 cs.CV cs.AI cs.CL cs.LG 67%

TULIP: Towards Unified Language-Image Pretraining

Zineng Tang, Long Lian, Seun Eisape, XuDong Wang, Roei Herzig, Adam Yala, Alane Suhr, Trevor Darrell, David M. Chan

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments (v2) Clarified fine-tuning process, updated appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01916 2025-04-03 cs.CV cs.AI cs.CL 67%

FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs

Mothilal Asokan, Kebin Wu, Fatima Albreiki

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01901 2025-04-03 cs.CV cs.AI cs.CL cs.RO 67%

Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

Haochen Wang, Yucheng Zhao, Tiancai Wang, Haoqiang Fan, Xiangyu Zhang, Zhaoxiang Zhang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01324 2025-04-03 cs.CV cs.AI cs.CL 67%

On Data Synthesis and Post-training for Visual Abstract Reasoning

Ke Zhu, Yu Wang, Jiangjiang Liu, Qunyi Xie, Shanshan Liu, Gang Zhang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23388 2025-04-01 cs.CV cs.AI cs.LG cs.MM 67%

COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptation

Fanding Huang, Jingyan Jiang, Qinting Jiang, Hebei Li, Faisal Nadeem Khan, Zhi Wang

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02946 2025-03-31 cs.CV cs.AI cs.LG cs.MM 67%

Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis

Po-Hsuan Huang, Jeng-Lin Li, Chin-Po Chen, Ming-Ching Chang, Wei-Chao Chen

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by WACV2025

Journal ref https://openaccess.thecvf.com/content/WACV2025/papers/Huang_Who_Brings_the_Frisbee_Probing_Hidden_Hallucination_Factors_in_Large_WACV_2025_paper.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14559 2025-03-20 cs.LG cs.AI cs.CL cs.CV 67%

Squeeze Out Tokens from Sample for Finer-Grained Data Governance

Weixiong Lin, Chen Ju, Haicheng Wang, Shengchao Hu, Shuai Xiao, Mengting Chen, Yuheng Jiao, Mingshuai Yao, Jinsong Lan, Qingwen Liu, Ying Chen

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏