arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7370 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7370 篇

2506.17375 2025-06-24 q-bio.NC cs.AI 74%

Challenges in Grounding Language in the Real World

Peter Lindes, Kaoutar Skiker

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments 14 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16201 2025-06-23 cs.RO cs.CV 74%

FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation

Sen Wang, Le Wang, Sanping Zhou, Jingyi Tian, Jiayi Li, Haowen Sun, Wei Tang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(国家人类-机器混合增强智能重点实验室) National Engineering Research Center for Visual Information and Applications(国家视觉信息与应用工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) Xi’an Jiaotong University(西安交通大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13282 2025-06-17 cs.CV 74%

Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling

Daichi Tanaka, Takumi Karasawa, Shu Takenouchi, Rei Kawakami

机构 * Institute Science of Tokyo(东京科学研究所)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01953 2025-06-17 cs.LG stat.ME 74%

The Landscape of Causal Discovery Data: Grounding Causal Discovery in Real-World Applications

Philippe Brouillard, Chandler Squires, Jonas Wahl, Konrad P. Kording, Karen Sachs, Alexandre Drouin, Dhanya Sridhar

机构 * Mila-Québec, Université de Montréal(蒙特利尔大学魁北克分校) Carnegie Mellon University(卡内基梅隆大学) Deutsches Forschungszentrum für künstliche Intelligenz (DFKI)(德国人工智能研究中心) University of Pennsylvania(宾夕法尼亚大学) Next Generation Analytics and Modulo Bio(下一代分析与Modulo Bio) ServiceNow Research(ServiceNow研究) Mila-Québec, Université Laval(魁北克蒙特利尔大学拉瓦尔分校)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG

Comments 39 pages, 8 figures; CLeaR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06622 2025-06-13 cs.CE cs.AI 74%

QuantMCP: Grounding Large Language Models in Verifiable Financial Reality

Yifan Zeng

机构 * Sun Yat-sen University, Guangzhou, China(中山大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18463 2025-05-30 cs.CV 74%

A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language Models

Shiho Noda, Atsuyuki Miyai, Qing Yu, Go Irie, Kiyoharu Aizawa

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments Accepted at ICIP2025 Dataset and Benchmark Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19761 2025-05-27 cs.AI 74%

Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning

Zican Hu, Wei Liu, Xiaoye Qu, Xiangyu Yue, Chunlin Chen, Zhi Wang, Yu Cheng

机构 * Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室) The Hong Kong University of Science and Technology(香港科学与技术大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Accepted by ICML 2025, 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07304 2025-04-11 cs.CL cs.AI 74%

PAYADOR: A Minimalist Approach to Grounding Language Models on Structured Data for Interactive Storytelling and Role-playing Games

Santiago Góngora, Luis Chiruzzo, Gonzalo Méndez, Pablo Gervás

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Presented at the 15th International Conference on Computational Creativity (ICCC'24)

Journal ref Proceedings of the Fifteenth International Conference on Computational Creativity (2024) 101-106

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13727 2025-04-02 cs.CL cs.AI 74%

LLM-Human Pipeline for Cultural Context Grounding of Conversations

Rajkumar Pujari, Dan Goldwasser

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Oral at NAACL 2025 Main conference. Albuquerque, USA. Apr 29 - May 4, 2025. 19 pages, 9 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05804 2025-04-01 cs.CV 74%

CASA: Class-Agnostic Shared Attributes in Vision-Language Models for Efficient Incremental Object Detection

Mingyi Guo, Yuyang Liu, Zhiyuan Yan, Zongying Lin, Peixi Peng, Yonghong Tian

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20188 2025-03-27 cs.CV 74%

Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector

Xiao Guo, Xiufeng Song, Yue Zhang, Xiaohong Liu, Xiaoming Liu

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments 8 figures; 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19240 2025-03-26 cs.CV cs.HC 74%

Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding

Hao Guo, Jianfei Zhu, Wei Fan, Chunzhi Yi, Feng Jiang

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19801 2025-03-18 cs.SE cs.AI cs.CL 74%

CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells

Atharva Naik, Marcus Alenius, Daniel Fried, Carolyn Rose

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01601 2025-03-04 cs.CV 74%

Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR

Muhammad Musab Ansari

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02089 2025-02-19 cs.CL cs.AI 74%

RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Jonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella, Quentin Carbonneaux, Taco Cohen, Gabriel Synnaeve

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Add repair model ablation, update related work

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00486 2025-02-04 cs.CV 74%

GEB+: A Benchmark for Generic Event Boundary Captioning, Grounding and Retrieval

Yuxuan Wang, Difei Gao, Licheng Yu, Stan Weixian Lei, Matt Feiszli, Mike Zheng Shou

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Updated in Jan. 2025, In Proceedings of the European Conference on Computer Vision 2022 [ECCV 2022], 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14035 2024-12-12 cs.CL cs.AI 74%

Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models

Sherzod Hakimov, Yerkezhan Abdullayeva, Kushal Koshti, Antonia Schmidt, Yan Weiser, Anne Beyer, David Schlangen

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

Comments Accepted at COLING 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09273 2024-11-15 cs.CL cs.AI 74%

Cross-Modal Consistency in Multimodal Large Language Models

Xiang Zhang, Senyu Li, Ning Shi, Bradley Hauer, Zijun Wu, Grzegorz Kondrak, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04892 2024-11-08 cs.CV 74%

In the Era of Prompt Learning with Vision-Language Models

Ankit Jha

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments ICVGIP 2024, Young Faculty Symposium

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03491 2024-11-07 cs.CV 74%

An Application-Agnostic Automatic Target Recognition System Using Vision Language Models

Anthony Palladino, Dana Gajewski, Abigail Aronica, Patryk Deptula, Alexander Hamme, Seiyoung C. Lee, Jeff Muri, Todd Nelling, Michael A. Riley, Brian Wong, Margaret Duff

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV

Comments Accepted to the Thirty-Seventh Annual Conference on Innovative Applications of Artificial Intelligence (IAAI-25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14509 2024-10-21 cs.CV 74%

CLIP-VAD: Exploiting Vision-Language Models for Voice Activity Detection

Andrea Appiani, Cigdem Beyan

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07540 2024-10-11 cs.CV 74%

CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection

Guankun Wang, Han Xiao, Huxin Gao, Renrui Zhang, Long Bai, Xiaoxiao Yang, Zhen Li, Hongsheng Li, Hongliang Ren

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05601 2024-10-10 cs.CV 74%

ReFIR: Grounding Large Restoration Models with Retrieval Augmentation

Hang Guo, Tao Dai, Zhihao Ouyang, Taolin Zhang, Yaohua Zha, Bin Chen, Shu-tao Xia

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05267 2024-10-08 cs.CL cs.CV 74%

Grounding Partially-Defined Events in Multimodal Data

Kate Sanders, Reno Kriz, David Etter, Hannah Recknor, Alexander Martin, Cameron Carpenter, Jingyang Lin, Benjamin Van Durme

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Preprint; 9 pages; 2024 EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12821 2024-10-07 cs.CL cs.LG 74%

Identifying Factual Inconsistencies in Summaries: Grounding LLM Inference via Task Taxonomy

Liyan Xu, Zhenlin Su, Mo Yu, Jin Xu, Jinho D. Choi, Jie Zhou, Fei Liu

专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG

Comments Accepted to EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14109 2024-09-27 cs.CV 74%

Vision-Language Models Assisted Unsupervised Video Anomaly Detection

Yalong Jiang, Liquan Mao

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07048 2024-09-12 cs.CV 74%

Pushing the Limits of Vision-Language Models in Remote Sensing without Human Annotations

Keumgang Cha, Donggeun Yu, Junghoon Seo

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments This study was primarily conducted during the latter half of 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11286 2024-08-23 cs.CV 74%

Video Emotion Open-vocabulary Recognition Based on Multimodal Large Language Model

Mengying Ge, Dongkai Tang, Mingyang Li

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05941 2024-08-13 cs.CR cs.AI 74%

Multimodal Large Language Models for Phishing Webpage Detection and Identification

Jehyun Lee, Peiyuan Lim, Bryan Hooi, Dinil Mon Divakaran

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.AI

Comments To appear in eCrime 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18695 2024-07-29 cs.CV cs.CL 74%

Grounding Language Models for Visual Entity Recognition

Zilin Xiao, Ming Gong, Paola Cascante-Bonilla, Xingyao Zhang, Jie Wu, Vicente Ordonez

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏