arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1565 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1565 篇

2406.00481 2025-06-03 cs.CV 79%

Efficient Open Set Single Image Test Time Adaptation of Vision Language Models

Manogna Sreenivas, Soma Biswas

专题命中 其他VLM :vision language model(title);vision-language model(abstract);分类 cs.CV

Comments Accepted at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18855 2025-05-27 cs.CV cs.CL 79%

Inference Compute-Optimal Video Vision Language Models

Peiqi Wang, ShengYun Peng, Xuewen Zhang, Hanchao Yu, Yibo Yang, Lifu Huang, Fujun Liu, Qifan Wang

机构 * MIT(麻省理工学院) Georgia Tech(佐治亚理工学院) Meta UC Davis(加州大学戴维斯分校)

专题命中 其他VLM :vision language model(title,abstract);分类 cs.CV

Comments Annual Meeting of the Association for Computational Linguistics (ACL), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10541 2025-05-26 cs.CV 79%

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis

Pengfei Wang, Guohai Xu, Weinong Wang, Junjie Yang, Jie Lou, Yunhua Xue

机构 * Pengfei Wang(王鹏飞) Guohai Xu(徐国海) Weinong Wang(王文龙) Junjie Yang(杨俊杰) Jie Lou(娄杰) Yunhua Xue(许云华)

专题命中 其他VLM :multimodal large language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11221 2025-05-19 cs.LG 79%

Sample Efficient Reinforcement Learning via Large Vision Language Model Distillation

Donghoon Lee, Tung M. Luu, Younghwan Lee, Chang D. Yoo

机构 * Robotics Program KAIST(韩国釜山科学技术院机器人计划) Electrical Engineering KAIST(韩国釜山科学技术院电子工程)

专题命中 其他VLM :vision language model(title);vision-language model(abstract);分类 cs.LG

Comments 5 pages, ICASSP 2025. The first two authors are equally contributed

Journal ref ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09425 2025-05-14 cs.CV cs.CL 79%

Vision-Language Models Do Not Understand Negation

Kumail Alhamoud, Shaden Alshammari, Yonglong Tian, Guohao Li, Philip Torr, Yoon Kim, Marzyeh Ghassemi

机构 * Institution1(机构1) Institution2(机构2)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

Comments CVPR 2025; project page: https://negbench.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19739 2025-04-29 cs.CV 79%

Contrastive Language-Image Learning with Augmented Textual Prompts for 3D/4D FER Using Vision-Language Model

Muzammil Behzad, Guoying Zhao

机构 * Information & Computer Science Department, King Fahd University of Petroleum & Minerals(国王法赫德石油与矿物大学信息与计算机科学系) Center for Machine Vision and Signal Analysis, University of Oulu(奥卢大学机器视觉与信号分析中心)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15271 2025-04-22 cs.CV 79%

Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models

Guo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang, Yicheng Liu, Lidong Lu, De-An Huang, Wonmin Byeon, Matthieu Le, Tuomas Rintamaki, Tyler Poon, Max Ehrlich, Tuomas Rintamaki, Tyler Poon, Tong Lu, Limin Wang, Bryan Catanzaro, Jan Kautz, Andrew Tao, Zhiding Yu, Guilin Liu

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13023 2025-04-18 cs.CL cs.CV 79%

ChatEXAONEPath: An Expert-level Multimodal Large Language Model for Histopathology Using Whole Slide Images

Sangwook Kim, Soonyoung Lee, Jongseong Jang

专题命中 其他VLM :multimodal large language model(title);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10049 2025-04-15 cs.CV cs.CL 79%

Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure

Théo Gigant, Camille Guinaudeau, Frédéric Dufaux

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07171 2025-04-03 cs.CV cs.CL 79%

BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature

Alejandro Lozano, Min Woo Sun, James Burgess, Liangyu Chen, Jeffrey J Nirschl, Jeffrey Gu, Ivan Lopez, Josiah Aklilu, Austin Wolfgang Katzer, Collin Chiu, Anita Rau, Xiaohan Wang, Yuhui Zhang, Alfred Seunghoon Song, Robert Tibshirani, Serena Yeung-Levy

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07790 2025-04-01 cs.CV 79%

Cropper: Vision-Language Model for Image Cropping through In-Context Learning

Seung Hyun Lee, Jijun Jiang, Yiran Xu, Zhuofang Li, Junjie Ke, Yinxiao Li, Junfeng He, Steven Hickson, Katie Datsenko, Sangpil Kim, Ming-Hsuan Yang, Irfan Essa, Feng Yang

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13817 2025-03-18 cs.CV 79%

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection

Le Yang, Ziwei Zheng, Boxu Chen, Zhengyu Zhao, Chenhao Lin, Chao Shen

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09248 2025-03-18 cs.CV 79%

Bayesian Test-Time Adaptation for Vision-Language Models

Lihua Zhou, Mao Ye, Shuaifeng Li, Nianxin Li, Xiatian Zhu, Lei Deng, Hongbin Liu, Zhen Lei

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04201 2025-03-07 cs.CL cs.AI 79%

Knowledge-Decoupled Synergetic Learning: An MLLM based Collaborative Approach to Few-shot Multimodal Dialogue Intention Recognition

Bin Chen, Yu Zhang, Hongfei Ye, Ziyi Huang, Hongyang Chen

专题命中 其他VLM :MLLM(title);multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02393 2025-03-05 cs.CV 79%

Vision-Language Model IP Protection via Prompt-based Learning

Lianyu Wang, Meng Wang, Huazhu Fu, Daoqiang Zhang

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18303 2025-03-03 cs.CV 79%

Efficient and Context-Aware Label Propagation for Zero-/Few-Shot Training-Free Adaptation of Vision-Language Model

Yushu Li, Yongyi Su, Adam Goodge, Kui Jia, Xun Xu

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08410 2025-02-26 cs.AI 79%

Specialized curricula for training vision-language models in retinal image analysis

Robbie Holland, Thomas R. P. Taylor, Christopher Holmes, Sophie Riedl, Julia Mai, Maria Patsiamanidi, Dimitra Mitsopoulou, Paul Hager, Philip Müller, Hendrik P. N. Scholl, Hrvoje Bogunović, Ursula Schmidt-Erfurth, Daniel Rueckert, Sobha Sivaprasad, Andrew J. Lotery, Martin J. Menten

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.AI

Comments Under review at npj Digital Medicine

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22108 2025-02-18 cs.CL cs.AI 79%

Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench

Zheyuan Liu, Guangyao Dou, Mengzhao Jia, Zhaoxuan Tan, Qingkai Zeng, Yongle Yuan, Meng Jiang

专题命中 其他VLM :multimodal large language model(title,abstract);分类 cs.AI

Comments NAACL Main 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00832 2025-02-10 cs.AI cs.CY 79%

Taking the Next Step with Generative Artificial Intelligence: The Transformative Role of Multimodal Large Language Models in Science Education

Arne Bewersdorff, Christian Hartmann, Marie Hornberger, Kathrin Seßler, Maria Bannert, Enkelejda Kasneci, Gjergji Kasneci, Xiaoming Zhai, Claudia Nerdel

专题命中 其他VLM :multimodal large language model(title,abstract);分类 cs.AI

Comments revised version 2. September 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00574 2025-02-05 cs.CV cs.MM 79%

EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model

Deng Li, Xin Liu, Bohao Xing, Baiqiang Xia, Yuan Zong, Bihan Wen, Heikki Kälviäinen

专题命中 其他VLM :MLLM(title);multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17665 2025-01-30 cs.RO cs.AI 79%

Planning with Vision-Language Models and a Use Case in Robot-Assisted Teaching

Xuzhe Dang, Lada Kudláčková, Stefan Edelkamp

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14148 2025-01-30 cs.CV 79%

SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation

Shuvendu Roy, Ali Etemad

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13851 2025-01-24 cs.LG 79%

Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes

Shiling Deng, Serge Belongie, Peter Ebert Christensen

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.LG

Comments 18 pages, 5 figures, 13 tables, GitHub repository: https://github.com/Seefreem/meme_text_retrieval_p1

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11231 2025-01-22 cs.CV 79%

KPL: Training-Free Medical Knowledge Mining of Vision-Language Models

Jiaxiang Liu, Tianxiang Hu, Jiawei Du, Ruiyuan Zhang, Joey Tianyi Zhou, Zuozhu Liu

专题命中 其他VLM :vision-language model(title);visual language model(abstract);分类 cs.CV

Comments AAAI(Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18216 2025-01-22 cs.CV cs.CL 79%

ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation

Mengyang Wu, Yuzhi Zhao, Jialun Cao, Mingjie Xu, Zhongming Jiang, Xuehui Wang, Qinbin Li, Guangneng Hu, Shengchao Qin, Chi-Wing Fu

专题命中 其他VLM :multimodal large language model(title,abstract);分类 cs.CV

Comments Accepted by the AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07819 2025-01-15 cs.CV 79%

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding

Haomiao Xiong, Yunzhi Zhuge, Jiawen Zhu, Lu Zhang, Huchuan Lu

专题命中 其他VLM :multimodal large language model(title);MLLM(abstract);分类 cs.CV

Comments Accepted to IEEE Transactions on Multimedia (TMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07802 2025-01-15 cs.AI physics.space-ph 79%

Visual Language Models as Operator Agents in the Space Domain

Alejandro Carrasco, Marco Nedungadi, Enrico M. Zucchelli, Amit Jain, Victor Rodriguez-Fernandez, Richard Linares

专题命中 其他VLM :visual language model(title);vision-language model(abstract);分类 cs.AI

Comments Updated version of the paper presented in 2025 AIAA SciTech. https://arc.aiaa.org/doi/10.2514/6.2025-1543

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04352 2025-01-09 cs.CV 79%

Online Gaussian Test-Time Adaptation of Vision-Language Models

Clément Fuchs, Maxime Zanella, Christophe De Vleeschouwer

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.21080 2024-12-31 cs.CV 79%

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Yifei Huang, Jilan Xu, Baoqi Pei, Yuping He, Guo Chen, Lijin Yang, Xinyuan Chen, Yaohui Wang, Zheng Nie, Jinyao Liu, Guoshun Fan, Dechen Lin, Fang Fang, Kunpeng Li, Chang Yuan, Yali Wang, Yu Qiao, Limin Wang

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08176 2024-12-12 cs.CV cs.MM 79%

TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning

Jingjing Xie, Yuxin Zhang, Jun Peng, Zhaohong Huang, Liujuan Cao

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

Comments Accepted by AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏