arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46073 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2505.05229 2025-10-29 cs.CV cs.MM 62%

Does CLIP perceive art the same way we do?

Andrea Asperti, Leonardo Dessì, Maria Chiara Tonetti, Nico Wu

机构 * Dept. of Informatics (DISI) University of Bologna(信息学院(DISI)博洛尼亚大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Journal ref Proceedings of IEEE International Conference on Content-Based Multimedia Indexing (IEEE CBMI 2025), Dublin, Ireland, 22-24 October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22370 2025-10-28 cs.RO cs.AI cs.CV cs.LG cs.SE 62%

BLIP-FusePPO: A Vision-Language Deep Reinforcement Learning Framework for Lane Keeping in Autonomous Vehicles

Seyed Ahmad Hosseini Miangoleh, Amin Jalal Aghdasian, Farzaneh Abdollahi

机构 * Department of Electrical Engineering, Amirkabir University of Technology (Tehran Polytechnic)(电气工程系,阿米尔卡比尔技术大学(德黑兰理工大学))

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments https://github.com/Amin-A96/BLIP-FusePPO-A-Vision-Language-Deep-Reinforcement-Learning-Framework-for-Lane-Keeping-in-Autonomous.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22045 2025-10-28 cs.CV cs.AI 62%

VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT

Hyeonsu Kang, Emily Bao, Anjan Goswami

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle - Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21806 2025-10-28 cs.CV cs.AI 62%

Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval

Jiaao Yu, Mingjie Han, Tao Gong, Jian Zhang, Man Lan

机构 * School of Computer Science and Technology, East China Normal University, China(上海师范大学计算机科学与技术学院) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15235 2025-10-24 cs.CV cs.CL 62%

ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding

Jialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai, Xinghao Chen

机构 * Peking University(北京大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22651 2025-10-24 cs.CV cs.CL cs.LG 62%

Sherlock: Self-Correcting Reasoning in Vision-Language Models

Yi Ding, Ruqi Zhang

机构 * Department of Computer Science, Purdue University, USA(计算机科学系,普渡大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Published at NeurIPS 2025, 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11261 2025-10-24 cs.AI cs.CL 62%

Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework

Yunpu Zhao, Rui Zhang, Junbin Xiao, Changxin Ke, Ruibo Hou, Yifan Hao, Ling Li

机构 * School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学) State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences(处理器国家重点实验室,中国科学院计算技术研究所) Department of Computer Science, National University of Singapore(计算机科学系,新加坡国立大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences(软件智能研究中心,中国科学院软件研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Journal ref Neurocomputing, Volume 659, 2026, 131217

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19678 2025-10-23 cs.CV cs.AI 62%

I Spy With My Model's Eye: Visual Search as a Behavioural Test for MLLMs

John Burden, Jonathan Prunty, Ben Slater, Matthieu Tehenan, Greg Davis, Lucy Cheke

机构 * Leverhulme Centre for the Future of Intelligence, University of Cambridge(未来智能研究中心、剑桥大学) Department of Engineering, University of Cambridge(工程系、剑桥大学) Department of Psychology, University of Cambridge(心理学系、剑桥大学) Department of Computer Science, University of Cambridge(计算机科学系、剑桥大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19001 2025-10-23 cs.CV cs.AI cs.RO 62%

Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts

Seungjun Yu, Junsung Park, Youngsun Lim, Hyunjung Shim

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17800 2025-10-22 cs.CV cs.CL cs.LG 62%

Glyph: Scaling Context Windows via Visual-Text Compression

Jiale Cheng, Yusen Liu, Xinyu Zhang, Yulin Fei, Wenyi Hong, Ruiliang Lyu, Weihan Wang, Zhe Su, Xiaotao Gu, Xiao Liu, Yushi Bai, Jie Tang, Hongning Wang, Minlie Huang

机构 * The Conversational Artificial Intelligence (CoAI) Group, Tsinghua University(清华大学对话人工智能(CoAI)小组) Zhipu AI(智谱AI) The Knowledge Engineering Group (KEG), Tsinghua University(清华大学知识工程小组(KEG))

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17771 2025-10-21 cs.AI cs.CV 62%

Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs

Zhining Liu, Ziyi Chen, Hui Liu, Chen Luo, Xianfeng Tang, Suhang Wang, Joy Zeng, Zhenwei Dai, Zhan Shi, Tianxin Wei, Benoit Dumoulin, Hanghang Tong

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊) Penn State University(宾夕法尼亚州立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 21 pages, 10 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17651 2025-10-21 cs.CV cs.AI cs.LG 62%

Frugal Federated Learning for Violence Detection: A Comparison of LoRA-Tuned VLMs and Personalized CNNs

Sébastien Thuau, Siba Haidar, Ayush Bajracharya, Rachid Chelouah

机构 * esieaLab(esiea实验室) ESIEA(ESIEA学院) ETIS Laboratory(ETIS实验室) CNRS(法国国家科学研究中心) UMR8051(UMR8051研究中心) University of CY Cergy(CY塞克大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 7 pages, 1 figure, FLTA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17405 2025-10-21 cs.CL cs.AI 62%

AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages

Mardiyyah Oduwole, Prince Mireku, Fatimo Adebanjo, Oluwatosin Olajide, Mahi Aminu Aliyu, Jekaterina Novikova

机构 * ML Collective Ashesi University(阿什esi大学) Abubakar Tafawa Balewa University(阿布巴克尔·塔法瓦·巴勒瓦大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16973 2025-10-21 cs.CV cs.AI physics.med-ph 62%

Foundation Models in Medical Image Analysis: A Systematic Review and Meta-Analysis

Praveenbalaji Rajendran, Mojtaba Safari, Wenfeng He, Mingzhe Hu, Shansong Wang, Jun Zhou, Xiaofeng Yang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15430 2025-10-21 cs.CV cs.AI 62%

Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models

Shuang Liang, Zhihao Xu, Jialing Tao, Hui Xue, Xiting Wang

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Withdrawn due to an accidental duplicate submission. This paper (arXiv:2510.15430) was unintentionally submitted as a new entry instead of a new version of our previous work (arXiv:2508.09201)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14807 2025-10-21 eess.IV cs.AI cs.CV 62%

FetalCLIP: A Visual-Language Foundation Model for Fetal Ultrasound Image Analysis

Fadillah Maani, Numan Saeed, Tausifa Saleem, Zaid Farooq, Hussain Alasmawi, Werner Diehl, Ameera Mohammad, Gareth Waring, Saudabi Valappi, Leanne Bricker, Mohammad Yaqub

机构 * Department of Computer Vision(计算机视觉系) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫比兹人工智能大学) Department of Machine Learning(机器学习系) Corniche Hospital, Abu Dhabi Health Services Company (SEHA)(阿布扎赫尔医院,阿布扎赫健康服务公司(SEHA))

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05342 2025-10-20 cs.CV cs.AI 62%

Refer to Any Segmentation Mask Group With Vision-Language Prompts

Shengcao Cao, Zijun Wei, Jason Kuen, Kangning Liu, Lingzhi Zhang, Jiuxiang Gu, HyunJoon Jung, Liang-Yan Gui, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Adobe

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17092 2025-10-20 cs.CV cs.CL 62%

Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI

Syed Abdul Gaffar Shakhadri, Kruthika KR, Kartik Basavaraj Angadi

机构 * SandLogic Technologies Pvt Ltd(沙德逻辑技术 Pvt Ltd)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14583 2025-10-17 cs.CV cs.CL 62%

Talking Points: Describing and Localizing Pixels

Matan Rusanovsky, Shimon Malnick, Shai Avidan

机构 * Tel Aviv University(特拉维夫大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14304 2025-10-17 cs.CV cs.AI 62%

Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding

Kyungryul Back, Seongbeom Park, Milim Kim, Mincheol Kwon, SangHyeok Lee, Hyunyoung Lee, Junhee Cho, Seunghyun Park, Jinkyu Kim

机构 * CSE, Korea University(韩国大学计算机科学与工程系) KT Corporation(KT公司) Soongsil University(顺成大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments EMNLP 2025 Findings; Project: https://github.com/KR-0822/TCD

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12931 2025-10-16 cs.CV cs.CL 62%

Unifying Vision-Language Latents for Zero-label Image Caption Enhancement

Sanghyun Byun, Jung Ick Guack, Mohanad Odema, Baisub Lee, Jacob Song, Woo Seong Chung

机构 * LG Electronics USA(LG电子美国公司)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted to PMLR and NeurIPS 2025 UniReps

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15298 2025-10-16 cs.CV cs.MM 62%

MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering

Xinqi Fan, Jingting Li, John See, Moi Hoon Yap, Wen-Huang Cheng, Xiaobai Li, Xiaopeng Hong, Su-Jing Wang, Adrian K. Davision

机构 * Department of Computing and Mathematics, Manchester Metropolitan University(计算与数学系,曼彻斯特 Metropolitan 大学) State Key Laboratory of Cognitive Science and Mental Health, Institute of Psychology, Chinese Academy of Sciences(认知科学与心理健康国家重点实验室,心理学研究所,中国科学院) Department of Psychology, University of the Chinese Academy of Sciences(心理学系,中国科学院大学) National Taiwan University(台湾大学) Zhejiang University(浙江大学) University of Oulu(奥卢大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Comments Micro-Expression Grand Challenge (MEGC) at ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00378 2025-10-15 cs.AI cs.CV 62%

CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding

Shixin Yi, Lin Shang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments The paper is not yet mature and needs further improvement

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10052 2025-10-14 cs.CV cs.AI 62%

Think Twice to See More: Iterative Visual Reasoning in Medical VLMs

Kaitao Chen, Shaohao Rui, Yankai Jiang, Jiamin Wu, Qihao Zheng, Chunfeng Song, Xiaosong Wang, Mu Zhou, Mianxin Liu

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) Rutgers University(罗格斯大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments 25 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11034 2025-10-13 cs.LG cs.AI cs.CL 62%

CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Models

Aneesh Komanduri, Karuna Bhaila, Xintao Wu

机构 * Department of Electrical Engineering and Computer Science University of Arkansas(电气工程与计算机科学系 奎萨克大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted to the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025 Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03355 2025-10-13 cs.LG cs.AI cs.CV 62%

Robustness in Both Domains: CLIP Needs a Robust Text Encoder

Elias Abad Rocamora, Christian Schlarmann, Naman Deep Singh, Yongtao Wu, Matthias Hein, Volkan Cevher

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03363 2025-10-09 cs.CV cs.AI eess.IV 62%

Unified Unsupervised Anomaly Detection via Matching Cost Filtering

Zhe Zhang, Mingxiu Cai, Gaochang Wu, Jing Zhang, Lingqiao Liu, Dacheng Tao, Tianyou Chai, Xiatian Zhu

机构 * State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, Shenyang, China(合成过程工业综合自动化国家重点实验室,东北大学,沈阳,中国) University of Surrey(Surrey大学) School of Computer Science, Wuhan University(武汉大学计算机学院) School of Computer Science, The University of Adelaide(阿德莱德大学计算机学院) College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) Surrey Institute for People-Centred Artificial Intelligence, and Centre for Vision, Speech and Signal Processing, University of Surrey(Surrey人本人工智能研究所,以及视觉、语音和信号处理中心,Surrey大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 63 pages (main paper and supplementary material), 39 figures, 58 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18269 2025-10-09 cs.CV cs.AI 62%

MAMS: Model-Agnostic Module Selection Framework for Video Captioning

Sangho Lee, Il Yong Chun, Hogun Park

机构 * Sangho Lee 1,2(Sangho Lee 教授) Il Yong Chun 1,3(Il Yong Chun 教授) Hogun Park 1(Hogun Park 教授)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to the AAAI 2025 Main Technical Track. This is an extended version of the original submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10610 2025-10-07 cs.CV cs.CL 62%

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

Zhaowei Wang, Wenhao Yu, Xiyu Ren, Jipeng Zhang, Yu Zhao, Rohit Saxena, Liang Cheng, Ginny Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman

机构 * CSE Department, HKUST(香港科技大学计算机科学与工程系) Tencent AI Seattle Lab(腾讯AI西雅图实验室) University of Edinburgh(爱丁堡大学) NVIDIA AI Technology Center (NVAITC), NVIDIA, Santa Clara, USA(英伟达圣克拉拉人工智能技术中心)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted as a spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03483 2025-10-07 cs.CV cs.AI 62%

DuPLUS: Dual-Prompt Vision-Language Framework for Universal Medical Image Segmentation and Prognosis

Numan Saeed, Tausifa Jan Saleem, Fadillah Maani, Muhammad Ridzuan, Hu Wang, Mohammad Yaqub

机构 * Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(计算机视觉系,Mohamed bin Zayed人工智能大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏