arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7370 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7370 篇

2411.00412 2025-06-23 cs.LG cs.AI cs.CL 76%

Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation

Bohan Lyu, Yadi Cao, Duncan Watson-Parris, Leon Bergen, Taylor Berg-Kirkpatrick, Rose Yu

机构 * Tsinghua University(清华大学) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments 37 pages, 16 figures

Journal ref In Proceedings of the Forty-second International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20252 2025-05-21 cs.CV cs.AI 76%

LogicQA: Logical Anomaly Detection with Vision Language Model Generated Questions

Yejin Kwon, Daeun Moon, Youngje Oh, Hyunsoo Yoon

机构 * Department of Industrial Engineering, Yonsei University(工业工程系,延世大学)

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV、cs.AI

Comments Accepted Industry Track at ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16047 2025-04-23 cs.CV cs.AI 76%

Evaluating Vision Language Models (VLMs) for Radiology: A Comprehensive Analysis

Frank Li, Hari Trivedi, Bardia Khosravi, Theo Dapamede, Mohammadreza Chavoshi, Abdulhameed Dere, Rohan Satya Isaac, Aawez Mansuri, Janice Newsome, Saptarshi Purkayastha, Judy Gichoya

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16488 2025-03-26 cs.HC cs.CV cs.LG 76%

VocalEyes: Enhancing Environmental Perception for the Visually Impaired through Vision-Language Models and Distance-Aware Object Detection

Kunal Chavan, Keertan Balaji, Spoorti Barigidad, Samba Raju Chiluveru

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02917 2025-03-06 eess.IV cs.AI cs.CV 76%

Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models

Deval Mehta, Yiwen Jiang, Catherine L Jan, Mingguang He, Kshitij Jadhav, Zongyuan Ge

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

Comments Accepted to Information Processing in Medical Imaging (IPMI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11168 2025-02-18 cs.CV cs.AI 76%

Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding

Xin Gu, Yaojie Shen, Chenxi Luo, Tiejian Luo, Yan Huang, Yuewei Lin, Heng Fan, Libo Zhang

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03722 2025-01-08 cs.CV cs.AI 76%

Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein

Xiaotong Guo, Deqian Yang, Dan Wang, Haochen Zhao, Yuan Li, Zhilin Sui, Tao Zhou, Lijun Zhang, Yanda Meng

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

Comments 8 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23214 2024-11-01 cs.LG cs.AI 76%

Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval

Sheryl Hsu, Omar Khattab, Chelsea Finn, Archit Sharma

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03959 2024-10-08 cs.CL cs.AI cs.CV cs.GR 76%

Grounding Language in Multi-Perspective Referential Communication

Zineng Tang, Lingjun Mao, Alane Suhr

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

Comments Accepted to EMNLP2024 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18868 2024-09-30 cs.CL cs.AI cs.LG 76%

Individuation in Neural Models with and without Visual Grounding

Alexey Tikhonov, Lisa Bylinina, Ivan P. Yamshchikov

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01326 2024-09-04 cs.RO cs.AI cs.LG 76%

Grounding Language Models in Autonomous Loco-manipulation Tasks

Jin Wang, Nikos Tsagarakis

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments Summit to ICRA@40. arXiv admin note: substantial text overlap with arXiv:2406.14655

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02244 2024-08-06 cs.CV cs.AI 76%

Evaluating Vision-Language Models for Zero-Shot Detection, Classification, and Association of Motorcycles, Passengers, and Helmets

Lucas Choi, Ross Greer

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03310 2024-07-19 cs.AI cs.CV 76%

V-IRL: Grounding Virtual Intelligence in Real Life

Jihan Yang, Runyu Ding, Ellis Brown, Xiaojuan Qi, Saining Xie

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

Comments Project page: https://virl-platform.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01256 2024-06-04 cs.CV cs.AI 76%

Augmented Commonsense Knowledge for Remote Object Grounding

Bahram Mohammadi, Yicong Hong, Yuankai Qi, Qi Wu, Shirui Pan, Javen Qinfeng Shi

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19336 2024-03-29 cs.CV cs.AI 76%

IVLMap: Instance-Aware Visual Language Grounding for Consumer Robot Navigation

Jiacui Huang, Hongtao Zhang, Mingbo Zhao, Zhou Wu

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03936 2023-12-08 cs.CV cs.CY cs.LG cs.SI 76%

The Potential of Vision-Language Models for Content Moderation of Children's Videos

Syed Hammad Ahmed, Shengnan Hu, Gita Sukthankar

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.LG

Comments 5 pages, 1 figure. Accepted at IEEE ICMLA 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01656 2023-12-06 cs.IR cs.AI cs.CV cs.HC 76%

The Contemporary Art of Image Search: Iterative User Intent Expansion via Vision-Language Model

Yilin Ye, Qian Zhu, Shishi Xiao, Kang Zhang, Wei Zeng

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

Comments Accepted by The 2024 ACM SIGCHI Conference on Computer-Supported Cooperative Work & Social Computing (CSCW) (Proc. CSCW 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00729 2023-11-08 cs.CV cs.AI 76%

ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action Detection

Thinh Phan, Khoa Vo, Duy Le, Gianfranco Doretto, Donald Adjeroh, Ngan Le

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.05186 2023-08-22 cs.CV cs.AI 76%

Suspected Object Matters: Rethinking Model's Prediction for One-stage Visual Grounding

Yang Jiao, Zequn Jie, Jingjing Chen, Lin Ma, Yu-Gang Jiang

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

Comments Accepted to ACM MM 23

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10438 2023-05-19 cs.CL cs.AI cs.CV cs.MM 76%

IMAGINATOR: Pre-Trained Image+Text Joint Embeddings using Word-Level Grounding of Images

Varuna Krishna, S Suryavardan, Shreyash Mishra, Sathyanarayanan Ramamoorthy, Parth Patwa, Megha Chakraborty, Aman Chadha, Amitava Das, Amit Sheth

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07528 2023-05-15 cs.CV cs.AI 76%

WEDGE: A multi-weather autonomous driving dataset built from generative vision-language models

Aboli Marathe, Deva Ramanan, Rahee Walambe, Ketan Kotecha

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI

Comments Accepted in Vision Datasets Understanding at CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.10285 2022-08-10 cs.CV cs.CL cs.LG 76%

Grounding Visual Representations with Texts for Domain Generalization

Seonwoo Min, Nokyung Park, Siwon Kim, Seunghyun Park, Jinkyu Kim

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

Comments ECCV 2022; 25 pages (including Supplementary Materials); Updated related works

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.03892 2022-03-28 cs.CL cs.AI cs.LG 76%

Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation Models

Steven Y. Feng, Kevin Lu, Zhuofu Tao, Malihe Alikhani, Teruko Mitamura, Eduard Hovy, Varun Gangal

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments Accepted to AAAI 2022. Code at https://github.com/styfeng/VisCTG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.11805 2022-01-19 cs.AI cs.LG 76%

Neural-Symbolic Integration for Interactive Learning and Conceptual Grounding

Benedikt Wagner, Artur d'Avila Garcez

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments Corrected references

Journal ref 1st Workshop on Human and Machine Decisions (WHMD 2021), NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.08409 2021-11-17 cs.LG cs.CV 76%

Grounding Psychological Shape Space in Convolutional Neural Networks

Lucas Bechberger, Kai-Uwe Kühnberger

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

Comments accepted at CIFMA2021 (https://cifma.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.01210 2021-06-21 cs.CV cs.LG cs.RO 76%

Embodied Language Grounding with 3D Visual Feature Representations

Mihir Prabhudesai, Hsiao-Yu Fish Tung, Syed Ashar Javed, Maximilian Sieb, Adam W. Harley, Katerina Fragkiadaki

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

Journal ref Conference on Computer Vision and Pattern Recognition. 2020, pp. 2220-2229

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.05379 2020-10-13 cs.CL cs.CV cs.LG 76%

MAF: Multimodal Alignment Framework for Weakly-Supervised Phrase Grounding

Qinxin Wang, Hao Tan, Sheng Shen, Michael W. Mahoney, Zhewei Yao

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.00133 2018-12-07 cs.CL cs.AI cs.LG 76%

Grounding Language for Transfer in Deep Reinforcement Learning

Karthik Narasimhan, Regina Barzilay, Tommi Jaakkola

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments JAIR 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.01871 2018-10-05 cs.LG cs.AI cs.RO stat.ML 76%

Grounding the Experience of a Visual Field through Sensorimotor Contingencies

Alban Laflaquière

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments 23 pages, 7 figures, published in Neurocomputing

Journal ref Neurocomputing, Volume 268, 13 December 2017, Pages 142-152

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.01870 2018-10-05 cs.LG cs.AI cs.RO stat.ML 76%

Grounding Perception: A Developmental Approach to Sensorimotor Contingencies

Alban Laflaquière, Nikolas Hemion, Michaël Garcia Ortiz, Jean-Christophe Baillie

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

Comments 8 pages, 4 figures, workshop at IROS 2015 conference

详情

展开后加载摘要…

URL PDF HTML 收藏