arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7360 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7360 篇

2301.11100 2023-01-27 cs.CV cs.CY cs.HC 79%

Vision-Language Models Performing Zero-Shot Tasks Exhibit Gender-based Disparities

Melissa Hall, Laura Gustafson, Aaron Adcock, Ishan Misra, Candace Ross

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.03160 2023-01-10 cs.CV 79%

Towards Real-Time Panoptic Narrative Grounding by an End-to-End Grounding Network

Haowei Wang, Jiayi Ji, Yiyi Zhou, Yongjian Wu, Xiaoshuai Sun

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 9 pages, 5 figures, accepted by AAAI23

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.02401 2023-01-09 cs.CL cs.AI 79%

You Truly Understand What I Need: Intellectual and Friendly Dialogue Agents grounding Knowledge and Persona

Jungwoo Lim, Myunghoon Kang, Yuna Hur, Seungwon Jung, Jinsung Kim, Yoonna Jang, Dongyub Lee, Hyesung Ji, Donghoon Shin, Seungryong Kim, Heuiseok Lim

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted at Findings of EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00514 2023-01-03 cs.CV 79%

Rethinking the Video Sampling and Reasoning Strategies for Temporal Sentence Grounding

Jiahao Zhu, Daizong Liu, Pan Zhou, Xing Di, Yu Cheng, Song Yang, Wenzheng Xu, Zichuan Xu, Yao Wan, Lichao Sun, Zeyu Xiong

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by EMNLP Findings, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.14757 2023-01-02 cs.CV 79%

A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-language Model

Mengde Xu, Zheng Zhang, Fangyun Wei, Yutong Lin, Yue Cao, Han Hu, Xiang Bai

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.13163 2022-12-29 cs.CV 79%

MRTNet: Multi-Resolution Temporal Network for Video Sentence Grounding

Wei Ji, Long Chen, Yinwei Wei, Yiming Wu, Tat-Seng Chua

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00836 2022-12-05 cs.CV 79%

UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding

Dave Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner, Angel X. Chang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12006 2022-12-02 cs.AI 79%

Differentiable Fuzzy $\mathcal{ALC}$: A Neural-Symbolic Representation Language for Symbol Grounding

Xuan Wu, Xinhao Zhu, Yizheng Zhao, Xinyu Dai

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13306 2022-12-02 cs.CV 79%

Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding

Yang Jin, Yongzhi Li, Zehuan Yuan, Yadong Mu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 18 pages, 7 figures, Accepted by Neurips 2022, Spotlight Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15516 2022-12-01 cs.CV 79%

DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding

Shilong Liu, Yaoyuan Liang, Feng Li, Shijia Huang, Hao Zhang, Hang Su, Jun Zhu, Lei Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to AAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15521 2022-11-29 cs.CV cs.CL 79%

G^3: Geolocation via Guidebook Grounding

Grace Luo, Giscard Biamby, Trevor Darrell, Daniel Fried, Anna Rohrbach

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Findings of EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16265 2022-11-29 cs.CV 79%

SeqTR: A Simple yet Universal Network for Visual Grounding

Chaoyang Zhu, Yiyi Zhou, Yunhang Shen, Gen Luo, Xingjia Pan, Mingbao Lin, Chao Chen, Liujuan Cao, Xiaoshuai Sun, Rongrong Ji

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 21 pages, 8 figures

Journal ref European Conference on Computer Vision, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14241 2022-11-28 cs.CV 79%

Look Around and Refer: 2D Synthetic Semantics Knowledge Distillation for 3D Visual Grounding

Eslam Mohamed Bakr, Yasmeen Alsaedy, Mohamed Elhoseiny

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Journal ref NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11492 2022-11-22 cs.CV 79%

ClipCrop: Conditioned Cropping Driven by Vision-Language Model

Zhihang Zhong, Mingxi Cheng, Zhirong Wu, Yuhui Yuan, Yinqiang Zheng, Ji Li, Han Hu, Stephen Lin, Yoichi Sato, Imari Sato

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09646 2022-11-18 cs.CV 79%

Language Conditioned Spatial Relation Reasoning for 3D Object Grounding

Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid, Ivan Laptev

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted in NeurIPS 2022; Project website: https://cshizhe.github.io/projects/vil3dref.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07912 2022-11-16 cs.CV 79%

YORO -- Lightweight End to End Visual Grounding

Chih-Hui Ho, Srikar Appalaraju, Bhavan Jasani, R. Manmatha, Nuno Vasconcelos

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to ECCVW on International Challenge on Compositional and Multimodal Perception

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.01901 2022-11-15 cs.CV cs.CL 79%

Incremental Object Grounding Using Scene Graphs

John Seon Keun Yi, Yoonwoo Kim, Sonia Chernova

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13942 2022-10-26 cs.LG cs.CL cs.MA 79%

Entity Divider with Language Grounding in Multi-Agent Reinforcement Learning

Ziluo Ding, Wanpeng Zhang, Junpeng Yue, Xiangjun Wang, Tiejun Huang, Zongqing Lu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12977 2022-10-25 cs.CV 79%

Language-free Training for Zero-shot Video Grounding

Dahye Kim, Jungin Park, Jiyoung Lee, Seongheon Park, Kwanghoon Sohn

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to WACV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12658 2022-10-25 cs.CL cs.CV 79%

Extending Phrase Grounding with Pronouns in Visual Dialogues

Panzhong Lu, Xin Zhang, Meishan Zhang, Min Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06354 2022-10-13 cs.CL cs.AI cs.SD eess.AS 79%

Text-to-Audio Grounding Based Novel Metric for Evaluating Audio Caption Similarity

Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments 9 pages, 8 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.00277 2022-09-02 cs.CV cs.CL cs.SD eess.AS 79%

Video-Guided Curriculum Learning for Spoken Video Grounding

Yan Xia, Zhou Zhao, Shangwei Ye, Yang Zhao, Haoyuan Li, Yi Ren

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.01691 2022-08-17 cs.RO cs.CL cs.LG 79%

Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan, Andy Zeng

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

Comments See website at https://say-can.github.io/ V1. Initial Upload. V2. Added PaLM results. Added study about new capabilities (drawer manipulation, chain of thought prompting, multilingual instructions). Added an ablation study of language model size. Added an open-source version of \algname on a simulated tabletop environment. Improved readability

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.05647 2022-08-12 cs.CV cs.MM 79%

PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative Grounding

Zihan Ding, Zi-han Ding, Tianrui Hui, Junshi Huang, Xiaoming Wei, Xiaolin Wei, Si Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.14698 2022-08-08 cs.CV 79%

Can Shuffling Video Benefit Temporal Bias Problem: A Novel Training Framework for Temporal Grounding

Jiachang Hao, Haifeng Sun, Pengfei Ren, Jingyu Wang, Qi Qi, Jianxin Liao

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ECCV2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.10168 2022-08-05 cs.CV 79%

Explore-And-Match: Bridging Proposal-Based and Proposal-Free With Transformer for Sentence Grounding in Videos

Sangmin Woo, Jinyoung Park, Inyong Koo, Sumin Lee, Minki Jeong, Changick Kim

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Code: https://github.com/sangminwoo/Explore-And-Match

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.13457 2022-07-28 cs.CV 79%

Reducing the Vision and Language Bias for Temporal Sentence Grounding

Daizong Liu, Xiaoye Qu, Wei Hu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.13325 2022-07-28 cs.CV 79%

SiRi: A Simple Selective Retraining Mechanism for Transformer-based Visual Grounding

Mengxue Qu, Yu Wu, Wu Liu, Qiqi Gong, Xiaodan Liang, Olga Russakovsky, Yao Zhao, Yunchao Wei

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 21 pages (including Supplementary Materials); Accepted to ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04769 2022-07-26 cs.AI 79%

On the Foundations of Grounding in Answer Set Programming

Roland Kaminski, Torsten Schaub

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments unpublished draft

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.01551 2022-07-25 cs.CV 79%

D3Net: A Unified Speaker-Listener Architecture for 3D Dense Captioning and Visual Grounding

Dave Zhenyu Chen, Qirui Wu, Matthias Nießner, Angel X. Chang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Project website: https://daveredrum.github.io/D3Net/

详情

展开后加载摘要…

URL PDF HTML 收藏