arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7360 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7360 篇

2112.08879 2022-07-22 cs.CV cs.CL 79%

Bottom Up Top Down Detection Transformers for Language Grounding in Images and Point Clouds

Ayush Jain, Nikolaos Gkanatsios, Ishita Mediratta, Katerina Fragkiadaki

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments First two authors contributed equally | ECCV 2022 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.08386 2022-07-19 cs.CV 79%

Entity-enhanced Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding

Xuejing Liu, Liang Li, Shuhui Wang, Zheng-Jun Zha, Zechao Li, Qi Tian, Qingming Huang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 17 pages, 10 figures, accepted by TPAMI. arXiv admin note: text overlap with arXiv:1908.10568

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02756 2022-07-07 cs.CV 79%

STVGFormer: Spatio-Temporal Video Grounding with Static-Dynamic Cross-Modal Understanding

Zihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin, Tiancai Ye, Wei-Shi Zheng

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Technical report. The 1st place solution in the HC-STVG track of the 4th Person in Context Challenge(2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00744 2022-07-05 cs.CV 79%

Gaussian Kernel-based Cross Modal Network for Spatio-Temporal Video Grounding

Zeyu Xiong, Daizong Liu, Pan Zhou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ICIP2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09358 2022-06-28 cs.CV 79%

What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text Inputs

Tal Shaharabany, Yoad Tewel, Lior Wolf

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.08172 2022-06-17 cs.CV 79%

RefCrowd: Grounding the Target in Crowd with Referring Expressions

Heqian Qiu, Hongliang Li, Taijin Zhao, Lanxiao Wang, Qingbo Wu, Fanman Meng

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.06619 2022-06-15 cs.CV 79%

TransVG++: End-to-End Visual Grounding with Language Conditioned Vision Transformer

Jiajun Deng, Zhengyuan Yang, Daqing Liu, Tianlang Chen, Wengang Zhou, Yanyong Zhang, Houqiang Li, Wanli Ouyang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:2104.08541

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00272 2022-06-09 cs.CV 79%

Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning

Li Yang, Yan Xu, Chunfeng Yuan, Wei Liu, Bing Li, Weiming Hu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.06686 2022-06-07 cs.CV 79%

Unpaired Referring Expression Grounding via Bidirectional Cross-Modal Matching

Hengcan Shi, Munawar Hayat, Jianfei Cai

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.08013 2022-05-24 cs.CV 79%

End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video Grounding

Mengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang, Zhou Zhao, Jiaxu Miao, Wenqiao Zhang, Wenming Tan, Jin Wang, Peng Wang, Shiliang Pu, Fei Wu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.06919 2022-05-17 cs.HC cs.AI 79%

Grounding Explainability Within the Context of Global South in XAI

Deepa Singh, Michal Slupczynski, Ajit G. Pillai, Vinoth Pandian Sermuga Pandian

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 4 pages, Presented at CHI 2022 Workshop on Human-Centered Explainable AI (HCXAI): Beyond Opening the Black-Box of AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08619 2022-05-17 cs.CL cs.AI 79%

Call for Customized Conversation: Customized Conversation Grounding Persona and Knowledge

Yoonna Jang, Jungwoo Lim, Yuna Hur, Dongsuk Oh, Suhyune Son, Yeonsoo Lee, Donghoon Shin, Seungryong Kim, Heuiseok Lim

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted paper at the Thirty-Sixth AAAI Conference on Artificial Intelligence (AAAI-22)

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.10634 2022-05-06 cs.CV 79%

Augmented 2D-TAN: A Two-stage Approach for Human-centric Spatio-Temporal Video Grounding

Chaolei Tan, Zihang Lin, Jian-Fang Hu, Xiang Li, Wei-Shi Zheng

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Best Paper Award at the 3rd Person in Context (PIC) Challenge CVPR Workshop 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05499 2022-04-13 cs.CV 79%

Position-aware Location Regression Network for Temporal Video Grounding

Sunoh Kim, Kimin Yun, Jin Young Choi

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted in AVSS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06085 2022-04-12 cs.CV cs.CL 79%

On Pursuit of Designing Multi-modal Transformer for Video Grounding

Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, Yuexian Zou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by Conference on Empirical Methods in Natural Language Processing (EMNLP 2021, Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.04347 2022-04-12 cs.CL cs.AI 79%

On the Importance of Karaka Framework in Multi-modal Grounding

Sai Kiran Gorthi, Radhika Mamidi

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02174 2022-04-06 cs.CV cs.CL 79%

Multi-View Transformer for 3D Visual Grounding

Shijia Huang, Yilun Chen, Jiaya Jia, Liwei Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments cvpr2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15442 2022-03-30 cs.CV cs.MM 79%

Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual Grounding

Jiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang, Xuwu Wang, Ji Zhang, Liang He, Xin Lin

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.14940 2022-03-29 cs.CV 79%

Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model

Yu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi, Yue Gao, Guoqi Li

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13049 2022-03-29 cs.CV 79%

Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning

Juncheng Li, Junlin Xie, Long Qian, Linchao Zhu, Siliang Tang, Fei Wu, Yi Yang, Yueting Zhuang, Xin Eric Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09905 2022-03-21 cs.CV 79%

Learning Affordance Grounding from Exocentric Images

Hongchen Luo, Wei Zhai, Jing Zhang, Yang Cao, Dacheng Tao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments CVPR2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06822 2022-03-15 cs.CV cs.CL cs.RO 79%

Grounding Commands for Autonomous Vehicles via Layer Fusion with Region-specific Dynamic Layer Attention

Hou Pong Chan, Mingxi Guo, Cheng-Zhong Xu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Submitted to IROS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.04222 2022-03-15 cs.CV cs.MM 79%

Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite Graphs

Kaifeng Gao, Long Chen, Yulei Niu, Jian Shao, Jun Xiao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2022. Code is available at https://github.com/Dawn-LX/VidSGG-BIG. We also won the 1st place of Video Relation Understanding (VRU) Grand Challenge in ACM Multimedia 2021, with a simplified version of our model.(The code for object tracklets generation is available at https://github.com/Dawn-LX/VidVRD-tracklets)

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.04281 2022-03-15 cs.CV 79%

Visual Grounding with Transformers

Ye Du, Zehua Fu, Qingjie Liu, Yunhong Wang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 7 pagrs, 3 figures. Accepted by ICME'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.05243 2022-03-11 cs.CV cs.CL cs.MM 79%

A Closer Look at Debiased Temporal Sentence Grounding in Videos: Dataset, Metric, and Approach

Xiaohan Lan, Yitian Yuan, Xin Wang, Long Chen, Zhi Wang, Lin Ma, Wenwu Zhu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02966 2022-03-08 cs.CV cs.CL 79%

Exploring Optical-Flow-Guided Motion and Detection-Based Appearance for Temporal Sentence Grounding

Daizong Liu, Xiang Fang, Wei Hu, Pan Zhou

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments arXiv admin note: text overlap with arXiv:2201.00457

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13959 2022-03-01 cs.IR cs.CL cs.LG 79%

Semi-Structured Query Grounding for Document-Oriented Databases with Deep Retrieval and Its Application to Receipt and POI Matching

Geewook Kim, Wonseok Hwang, Minjoon Seo, Seunghyun Park

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

Comments To appear in AAAI-22 Workshop on Knowledge Discovery from Unstructured Data in Financial Services

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08541 2022-01-17 cs.CV 79%

TransVG: End-to-End Visual Grounding with Transformers

Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, Houqiang Li

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments This paper has been accepted by ICCV2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.02848 2022-01-14 cs.CV cs.IR 79%

Learning Sample Importance for Cross-Scenario Video Temporal Grounding

Peijun Bao, Yadong Mu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 7 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.00457 2022-01-04 cs.CV 79%

Exploring Motion and Appearance Information for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Pan Zhou, Yang Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by AAAI2022

详情

展开后加载摘要…

URL PDF HTML 收藏