arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1566 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1566 篇

2505.07347 2025-05-13 cs.CV 57%

AI-Enabled Accurate Non-Invasive Assessment of Pulmonary Hypertension Progression via Multi-Modal Echocardiography

Jiewen Yang, Taoran Huang, Shangwei Ding, Xiaowei Xu, Qinhua Zhao, Yong Jiang, Jiarong Guo, Bin Pu, Jiexuan Zheng, Caojin Zhang, Hongwen Fei, Xiaomeng Li

机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子与计算机工程系) Guangdong Cardiovascular Institute, Guangdong Provincial People’s Hospital (Guangdong Academy of Medical Sciences), Southern Medical University(广东省心血管病研究所,广东省人民医院(广东省医学科学院)) Department of Ultrasound, The First Affiliated Hospital of Guangzhou Medical University(广州市第一人民医院超声科) Department of Pulmonary Circulation, Shanghai Pulmonary Hospital, Tongji University School of Medicine(上海 pulmonary 医院,同济大学医学院) Department of Echocardiography, Fuwai Hospital Chinese Academy of Medical Sciences(阜外医院中国医学科学院) Guangdong Provincial Key Laboratory of South China Structural Heart Disease(广东省南方结构性心脏病重点实验室) State Key Laboratory of Cardiovascular Disease, Department of Echocardiography, National Center for Cardiovascular Diseases, Fuwai Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(心血管疾病国家重点实验室,国家心血管病中心,阜外医院,中国医学科学院和北京协和医学院) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03611 2025-05-07 cs.CV 57%

Learning Unknown Spoof Prompts for Generalized Face Anti-Spoofing Using Only Real Face Images

Fangling Jiang, Qi Li, Weining Wang, Wei Shen, Bing Liu, Zhenan Sun

机构 * School of Computer Science, University of South China(南方科技大学计算机科学学院) New Laboratory of Pattern Recognition, MAIS, CASIA(模式识别新实验室,MAIS,CASIA) School of Artificial Intelligence, UCAS(人工智能学院,UCAS) The Laboratory of Cognition and Decision Intelligence for Complex Systems, CASIA(复杂系统认知与决策智能实验室,CASIA) OPPO AI Center(OPPO AI中心)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00569 2025-05-02 cs.CV 57%

AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis

Enmin Zhong, Carlos R. del-Blanco, Daniel Berjón, Fernando Jaureguizar, Narciso García

机构 * Grupo de Tratamiento de Imágenes (GTI), Information Processing and Telecommunications Center, ETSI Telecomunicación, Universidad Politécnica de Madrid(图像处理小组(GTI)、信息处理与电信中心、电信工程学院、马德里理工大学)

专题命中 其他VLM :visual language model(abstract);分类 cs.CV

Comments 6 pages, 3 figures,Accepted for the poster session at the CV4Animals workshop: Computer Vision for Animal Behavior Tracking and Modeling In conjunction with Computer Vision and Pattern Recognition 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07072 2025-04-30 cs.CL cs.CV 57%

Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation

Israfel Salazar, Manuel Fernández Burda, Shayekh Bin Islam, Arshia Soltani Moakhar, Shivalika Singh, Fabian Farestam, Angelika Romanou, Danylo Boiko, Dipika Khullar, Mike Zhang, Dominik Krzemiński, Jekaterina Novikova, Luísa Shimabucoro, Joseph Marvin Imperial, Rishabh Maheshwary, Sharad Duwal, Alfonso Amayuelas, Swati Rajwal, Jebish Purbey, Ahmed Ruby, Nicholas Popovič, Marek Suppa, Azmine Toushik Wasi, Ram Mohan Rao Kadiyala, Olga Tsymboi, Maksim Kostritsya, Bardia Soltani Moakhar, Gabriel da Costa Merlin, Otávio Ferracioli Coletti, Maral Jabbari Shiviari, MohammadAmin farahani fard, Silvia Fernandez, María Grandury, Dmitry Abulkhanov, Drishti Sharma, Andre Guarnier De Mitri, Leticia Bossatto Marchezi, Setayesh Heydari, Johan Obando-Ceron, Nazar Kohut, Beyza Ermis, Desmond Elliott, Enzo Ferrante, Sara Hooker, Marzieh Fadaee

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments v2: corrected the author list

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19742 2025-04-29 cs.CV 57%

EcoWikiRS: Learning Ecological Representation of Satellite Images from Weak Supervision with Species Observations and Wikipedia

Valerie Zermatten, Javiera Castillo-Navarro, Pallavi Jain, Devis Tuia, Diego Marcos

机构 * EPFL(瑞士联邦理工学院) CNAM(法国国家科学与技术研究中心) INRIA(法国国家信息与自动化研究所) CIHEAM-IAMM(CIHEAM- IAMM) Univ. of Montpellier(蒙彼利埃大学)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

Comments Accepted at EarthVision 2025 (CVPRW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18961 2025-04-29 cs.IR cs.AI 57%

Feature Fusion Revisited: Multimodal CTR Prediction for MMCTR Challenge

Junjie Zhou

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家实验室,南京大学,中国) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学,中国)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

Comments A technical report for the MMCTR Challenge held by EReL@MIR Workshop at WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18509 2025-04-28 cs.CV 57%

Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation

Shivam Duggal, Yushi Hu, Oscar Michel, Aniruddha Kembhavi, William T. Freeman, Noah A. Smith, Ranjay Krishna, Antonio Torralba, Ali Farhadi, Wei-Chiu Ma

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments CVPR 2025. Project page and codes: https://eval3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16433 2025-04-24 cs.CV 57%

FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing

Hariseetharam Gunduboina, Muhammad Haris Khan, Biplab Banerjee

机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔分校) Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13971 2025-04-22 cs.CY cs.AI cs.ET cs.NI 57%

The Future of Internet of Things and Multimodal Language Models in 6G Networks: Opportunities and Challenges

Abdelrahman Soliman

机构 * University of Guelph(圭尔夫大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13650 2025-04-21 cs.CV 57%

EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model

Sijing Li, Tianwei Lin, Lingshuai Lin, Wenqiao Zhang, Jiang Liu, Xiaoda Yang, Juncheng Li, Yucheng He, Xiaohui Song, Jun Xiao, Yueting Zhuang, Beng Chin Ooi

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13209 2025-04-21 cs.CR cs.AI 57%

On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks

Ting Bi, Chenghang Ye, Zheyu Yang, Ziyi Zhou, Cui Tang, Jun Zhang, Zui Tao, Kailong Wang, Liting Zhou, Yang Yang, Tianlong Yu

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11509 2025-04-18 cs.IR cs.CV 57%

PATFinger: Prompt-Adapted Transferable Fingerprinting against Unauthorized Multimodal Dataset Usage

Wenyi Zhang, Ju Jia, Xiaojun Jia, Yihao Huang, Xinfeng Li, Cong Wu, Lina Wang

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10471 2025-04-15 cs.CV cs.CL 57%

MIEB: Massive Image Embedding Benchmark

Chenghao Xiao, Isaac Chung, Imene Kerboua, Jamie Stirling, Xin Zhang, Márton Kardos, Roman Solomatin, Noura Al Moubayed, Kenneth Enevoldsen, Niklas Muennighoff

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09598 2025-04-15 cs.CV 57%

DualPrompt-MedCap: A Dual-Prompt Enhanced Approach for Medical Image Captioning

Yining Zhao, Ali Braytee, Mukesh Prasad

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments 11 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14911 2025-04-15 cs.CV 57%

Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology

Siyuan Yan, Ming Hu, Yiwen Jiang, Xieji Li, Hao Fei, Philipp Tschandl, Harald Kittler, Zongyuan Ge

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Our dataset and code will be publicly available at https://github.com/SiyuanYan1/Derm1M

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13878 2025-04-15 cs.HC cs.CV 57%

Eye Gaze as a Signal for Conveying User Attention in Contextual AI Systems

Ethan Wilson, Naveen Sendhilnathan, Charlie S. Burlingham, Yusuf Mansour, Robert Cavin, Sai Deep Tetali, Ajoy Savio Fernandes, Michael J. Proulx

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

Comments To appear in ETRA '25: Proceedings of the 2025 Symposium on Eye Tracking Research and Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00672 2025-04-15 cs.CV 57%

ExpertAF: Expert Actionable Feedback from Video

Kumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos, Kris Kitani, Kristen Grauman

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07954 2025-04-11 cs.CV cs.CL 57%

Perception-R1: Pioneering Perception Policy with Reinforcement Learning

En Yu, Kangheng Lin, Liang Zhao, Jisheng Yin, Yana Wei, Yuang Peng, Haoran Wei, Jianjian Sun, Chunrui Han, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Jingyu Wang, Wenbing Tao

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

Comments Github page: https://github.com/linkangheng/PR1

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07643 2025-04-11 cs.IR cs.CL cs.CV 57%

CollEX -- A Multimodal Agentic RAG System Enabling Interactive Exploration of Scientific Collections

Florian Schneider, Narges Baba Ahmadi, Niloufar Baba Ahmadi, Iris Vogel, Martin Semmann, Chris Biemann

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06272 2025-04-10 cs.IR cs.AI 57%

RAVEN: An Agentic Framework for Multimodal Entity Discovery from Large-Scale Video Collections

Kevin Dela Rosa

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

Comments Presented at AI Agent for Information Retrieval: Generating and Ranking (Agent4IR) @ AAAI 2025 [https://sites.google.com/view/ai4ir/aaai-2025]

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04221 2025-04-08 cs.CV 57%

Evaluating Graphical Perception with Multimodal LLMs

Rami Huu Nguyen, Kenichi Maeda, Mahsa Geshvadi, Daniel Haehn

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments 6 pages, 5 figures, 1 teaser, IEEE Pacific Visualization 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03153 2025-04-07 cs.LG 57%

MORAL: A Multimodal Reinforcement Learning Framework for Decision Making in Autonomous Laboratories

Natalie Tirabassi, Sathish A. P. Kumar, Sumit Jha, Arvind Ramanathan

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

Comments 9 pages, 14 figures and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01662 2025-04-03 eess.IV cs.CV 57%

BioAtt: Anatomical Prior Driven Low-Dose CT Denoising

Namhun Kim, UiHyun Cho

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18499 2025-04-01 cs.CV 57%

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

Pengfei Zhou, Xiaopeng Peng, Jiajun Song, Chuanhao Li, Zhaopan Xu, Yue Yang, Ziyao Guo, Hao Zhang, Yuqi Lin, Yefei He, Lirui Zhao, Shuo Liu, Tianhua Li, Yuxuan Xie, Xiaojun Chang, Yu Qiao, Wenqi Shao, Kaipeng Zhang

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments 53 pages, 19 figures, accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07622 2025-04-01 cs.CV 57%

Pretrain like Your Inference: Masked Tuning Improves Zero-Shot Composed Image Retrieval

Junyang Chen, Hanjiang Lai

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments accepted by ICME 2025, this is the full version of paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00493 2025-03-28 cs.CV cs.CL 57%

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Duo Zheng, Shijia Huang, Liwei Wang

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21130 2025-03-28 cs.HC cs.CV 57%

VideoMix: Aggregating How-To Videos for Task-Oriented Learning

Saelyne Yang, Anh Truong, Juho Kim, Dingzeyu Li

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments In Proceedings of the 30th International Conference on Intelligent User Interfaces (IUI '25) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20680 2025-03-27 cs.CV cs.CL 57%

Vision as LoRA

Han Wang, Yongjie Ye, Bingru Li, Yuxiang Nie, Jinghui Lu, Jingqun Tang, Yanjie Wang, Can Huang

专题命中 其他VLM :MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05186 2025-03-27 cs.CV 57%

Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions

Chan Hur, Jeong-hun Hong, Dong-hun Lee, Dabin Kang, Semin Myeong, Sang-hyo Park, Hyeyoung Park

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19910 2025-03-26 cs.CV cs.IR 57%

CoLLM: A Large Language Model for Composed Image Retrieval

Chuong Huynh, Jinyu Yang, Ashish Tawari, Mubarak Shah, Son Tran, Raffay Hamid, Trishul Chilimbi, Abhinav Shrivastava

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments CVPR 2025. Project page: https://collm-cvpr25.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏