arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7360 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7360 篇

2309.12657 2024-10-28 cs.CV 79%

Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding

Jiazhen Wang, Bin Liu, Changtao Miao, Zhiwei Zhao, Wanyi Zhuang, Qi Chu, Nenghai Yu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments This work has been submitted to the IEEE for possible publication. Camera-ready version and supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.17379 2024-10-23 cs.RO cs.AI 79%

EMPOWER: Embodied Multi-role Open-vocabulary Planning with Online Grounding and Execution

Francesco Argenziano, Michele Brienza, Vincenzo Suriani, Daniele Nardi, Domenico D. Bloisi

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted at IROS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15615 2024-10-22 cs.CV 79%

Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding

Yang Liu, Daizong Liu, Wei Hu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14544 2024-10-21 cs.AI 79%

Computational Grounding of Responsibility Attribution and Anticipation in LTLf

Giuseppe De Giacomo, Emiliano Lorini, Timothy Parker, Gianmarco Parretti

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13598 2024-10-18 cs.CV 79%

Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding

Jongbhin Woo, Hyeonggon Ryu, Youngjoon Jang, Jae Won Cho, Joon Son Chung

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACMMM 24

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12813 2024-10-18 cs.MM cs.CV 79%

ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models

Mengxue Qu, Xiaodong Chen, Wu Liu, Alicia Li, Yao Zhao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12369 2024-10-17 cs.CV 79%

Context-Infused Visual Grounding for Art

Selina Khan, Nanne van Noord

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12225 2024-10-17 cs.CV 79%

Evaluating Cascaded Methods of Vision-Language Models for Zero-Shot Detection and Association of Hardhats for Increased Construction Safety

Lucas Choi, Ross Greer

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14494 2024-10-15 cs.CV 79%

Revisiting Few-Shot Object Detection with Vision-Language Models

Anish Madan, Neehar Peri, Shu Kong, Deva Ramanan

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments The first two authors contributed equally. This work has been accepted to the Neural Information Processing Systems (NeurIPS) 2024 Datasets & Benchmark Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16475 2024-10-11 cs.CV 79%

Enhancing HOI Detection with Contextual Cues from Large Vision-Language Models

Yu-Wei Zhan, Fan Liu, Xin Luo, Xin-Shun Xu, Liqiang Nie, Mohan Kankanhalli

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06108 2024-10-10 cs.AI 79%

ConceptAgent: LLM-Driven Precondition Grounding and Tree Search for Robust Task Planning and Execution

Corban Rivera, Grayson Byrd, William Paul, Tyler Feldman, Meghan Booker, Emma Holmes, David Handelman, Bethany Kemp, Andrew Badger, Aurora Schmidt, Krishna Murthy Jatavallabhula, Celso M de Melo, Lalithkumar Seenivasan, Mathias Unberath, Rama Chellappa

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04426 2024-10-08 cs.CV 79%

CoVLM: Leveraging Consensus from Vision-Language Models for Semi-supervised Multi-modal Fake News Detection

Devank, Jayateja Kalla, Soma Biswas

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted in ACCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06214 2024-10-08 cs.CV 79%

CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding

Eslam Abdelrahman, Mohamed Ayman, Mahmoud Ahmed, Habib Slim, Mohamed Elhoseiny

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03161 2024-10-07 cs.AI 79%

Adaptive Masking Enhances Visual Grounding

Sen Jia, Lei Li

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Code will be available at https://github.com/git-lenny/IMAGE

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10774 2024-10-02 cs.CL cs.AI 79%

MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents

Liyan Tang, Philippe Laban, Greg Durrett

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17330 2024-09-27 cs.CV 79%

VL4AD: Vision-Language Models Improve Pixel-wise Anomaly Detection

Liangyu Zhong, Joachim Sicking, Fabian Hüger, Hanno Gottschalk

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 27 pages, 9 figures, to be published in ECCV 2024 2nd Workshop on Vision-Centric Autonomous Driving (VCAD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19071 2024-09-18 cs.CL cs.AI 79%

EmPO: Emotion Grounding for Empathetic Response Generation through Preference Optimization

Ondrej Sotolar, Vojtech Formanek, Alok Debnath, Allison Lahnala, Charles Welch, Lucie FLek

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments v02, 8 pages long paper, EMNLP ACL style

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08251 2024-09-13 cs.CV 79%

Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding

Hongyu Li, Tianrui Hui, Zihan Ding, Jing Zhang, Bin Ma, Xiaoming Wei, Jizhong Han, Si Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14206 2024-09-13 cs.CV 79%

LLM4VG: Large Language Models Evaluation for Video Grounding

Wei Feng, Xin Wang, Hong Chen, Zeyang Zhang, Houlun Chen, Zihan Song, Yuwei Zhou, Yuekui Yang, Haiyang Wu, Wenwu Zhu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.04999 2024-09-10 cs.CV cs.MM 79%

Visual Grounding with Multi-modal Conditional Adaptation

Ruilin Yao, Shengwu Xiong, Yichen Zhao, Yi Rong

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2024 [Oral]

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09012 2024-09-10 cs.CV 79%

FashionLOGO: Prompting Multimodal Large Language Models for Fashion Logo Embeddings

Zhen Wang, Da Li, Yulin Su, Min Yang, Minghui Qiu, Walton Wang

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16341 2024-09-10 cs.CV 79%

Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding

Yuanhao Xiong, Long Zhao, Boqing Gong, Ming-Hsuan Yang, Florian Schroff, Ting Liu, Cho-Jui Hsieh, Liangzhe Yuan

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to ICLR2024, see https://openreview.net/forum?id=5dlfiJIXoh for more details

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13400 2024-09-06 cs.CV 79%

HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding

Linhui Xiao, Xiaoshan Yang, Fang Peng, Yaowei Wang, Changsheng Xu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2024. The project page: https://github.com/linhuixiao/HiVG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.03109 2024-09-06 cs.CV cs.CR 79%

FIDAVL: Fake Image Detection and Attribution using Vision-Language Model

Mamadou Keita, Wassim Hamidouche, Hessen Bougueffa Eutamene, Abdelmalik Taleb-Ahmed, Abdenour Hadid

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14491 2024-09-04 cs.CV 79%

PD-APE: A Parallel Decoding Framework with Adaptive Position Encoding for 3D Visual Grounding

Chenshu Hou, Liang Peng, Xiaopei Wu, Xiaofei He, Wenxiao Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14154 2024-09-04 cs.CL cs.CV cs.CY 79%

MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms

Yiqiao Jin, Minje Choi, Gaurav Verma, Jindong Wang, Srijan Kumar

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV

Comments In Proceedings of ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00556 2024-09-04 cs.CV 79%

FADE: Few-shot/zero-shot Anomaly Detection Engine using Large Vision-Language Model

Yuanwei Li, Elizaveta Ivanova, Martins Bruveris

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments 13 pages, 2 figures, Accepted for BMVC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00238 2024-09-04 cs.CL cs.CV 79%

Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data

Spencer Whitehead, Jacob Phillips, Sean Hendryx

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16314 2024-08-30 cs.CV 79%

ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding

Minghang Zheng, Jiahua Zhang, Qingchao Chen, Yuxin Peng, Yang Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01024 2024-08-22 cs.AI 79%

Semantic Skill Grounding for Embodied Instruction-Following in Cross-Domain Environments

Sangwoo Shin, Seunghyun Kim, Youngsoo Jang, Moontae Lee, Honguk Woo

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Findings of ACL-2024 Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏