arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7387 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7387 篇

2402.18169 2024-03-01 cs.CL 67%

MIKO: Multimodal Intention Knowledge Distillation from Large Language Models for Social-Media Commonsense Discovery

Feihong Lu, Weiqi Wang, Yangyifei Luo, Ziqin Zhu, Qingyun Sun, Baixuan Xu, Haochen Shi, Shiqi Gao, Qian Li, Yangqiu Song, Jianxin Li

专题命中 视觉定位与Grounding :multimodal large language model(abstract);MLLM(abstract)

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01107 2024-02-27 cs.CV cs.AI cs.LG stat.ML 67%

Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models

Hyeonho Jeong, Jong Chul Ye

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to ICLR 2024, Project Page: http://ground-a-video.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17429 2024-02-02 cs.CV cs.AI cs.CL cs.LG 67%

Commonsense for Zero-Shot Natural Language Video Localization

Meghana Holla, Ismini Lourentzou

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07435 2023-12-13 cs.CV cs.AI cs.CL cs.LG 67%

Cross-modal Contrastive Learning with Asymmetric Co-attention Network for Video Moment Retrieval

Love Panta, Prashant Shrestha, Brabeem Sapkota, Amrita Bhattarai, Suresh Manandhar, Anand Kumar Sah

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00855 2023-12-13 cs.RO cs.AI cs.CL cs.CV cs.LG 67%

Grounded Decoding: Guiding Text Generation with Grounded Models for Embodied Agents

Wenlong Huang, Fei Xia, Dhruv Shah, Danny Driess, Andy Zeng, Yao Lu, Pete Florence, Igor Mordatch, Sergey Levine, Karol Hausman, Brian Ichter

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05828 2023-11-10 cs.CV cs.AI cs.LG 67%

Adapting Contrastive Language-Image Pretrained (CLIP) Models for Out-of-Distribution Detection

Nikolas Adaloglou, Felix Michels, Tim Kaiser, Markus Kollmann

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments version_02

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04067 2023-11-08 cs.LG cs.AI cs.CV 67%

Multitask Multimodal Prompted Training for Interactive Embodied Task Completion

Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.15059 2023-10-24 cs.RO cs.AI cs.CV cs.LG 67%

Robot Skill Generalization via Keypoint Integrated Soft Actor-Critic Gaussian Mixture Models

Iman Nematollahi, Kirill Yankov, Wolfram Burgard, Tim Welschehold

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at the International Symposium on Experimental Robotics (ISER) 2023. Videos at http://kis-gmm.cs.uni-freiburg.de/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.14108 2023-10-24 cs.LG cs.AI cs.CV 67%

CLIP meets Model Zoo Experts: Pseudo-Supervision for Visual Enhancement

Mohammadreza Salehi, Mehrdad Farajtabar, Maxwell Horton, Fartash Faghri, Hadi Pouransari, Raviteja Vemulapalli, Oncel Tuzel, Ali Farhadi, Mohammad Rastegari, Sachin Mehta

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.13372 2023-10-17 cs.CV cs.AI cs.LG 67%

Localizing Moments in Long Video Via Multimodal Guidance

Wayner Barrios, Mattia Soldan, Alberto Mario Ceballos-Arroyo, Fabian Caba Heilbron, Bernard Ghanem

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13952 2023-09-26 cs.CV cs.AI cs.CL cs.LG 67%

VidChapters-7M: Video Chapters at Scale

Antoine Yang, Arsha Nagrani, Ivan Laptev, Josef Sivic, Cordelia Schmid

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at NeurIPS 2023 Track on Datasets and Benchmarks; Project Webpage: https://antoyang.github.io/vidchapters.html ; 31 pages; 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04088 2023-09-08 cs.AI cs.CL cs.CV cs.LG cs.RO 67%

LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models

Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M. Sadler, Wei-Lun Chao, Yu Su

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 14 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.08347 2023-07-20 cs.CV cs.AI cs.LG 67%

M-FLAG: Medical Vision-Language Pre-training with Frozen Language Models and Latent Space Geometry Optimization

Che Liu, Sibo Cheng, Chen Chen, Mengyun Qiao, Weitong Zhang, Anand Shah, Wenjia Bai, Rossella Arcucci

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted by MICCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.03727 2023-06-07 cs.CV cs.AI cs.LG cs.RO 67%

Towards Visual Foundational Models of Physical Scenes

Chethan Parameshwara, Alessandro Achille, Matthew Trager, Xiaolong Li, Jiawei Mo, Matthew Trager, Ashwin Swaminathan, CJ Taylor, Dheera Venkatraman, Xiaohan Fei, Stefano Soatto

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments TLDR: Physical scenes are equivalence classes of sufficient statistics, and can be inferred uniquely by any agent measuring the same finite data; We formalize and implement an approach to representation learning that overturns "naive realism" in favor of an analytical approach of Russell and Koenderink. NeRFs cannot capture the physical scenes, but combined with Diffusion Models they can

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05665 2023-06-01 cs.CV cs.AI cs.LG cs.MM 67%

ImageBind: One Embedding Space To Bind Them All

Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, Ishan Misra

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments CVPR 2023 (Highlighted Paper). Website: https://imagebind.metademolab.com/ Code/Models: https://github.com/facebookresearch/ImageBind

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.07093 2023-04-18 cs.CV cs.AI cs.CL cs.GR cs.LG 67%

GLIGEN: Open-Set Grounded Text-to-Image Generation

Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, Yong Jae Lee

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14905 2023-03-28 cs.CV cs.AI cs.CL cs.LG cs.MM 67%

Multi-Modal Few-Shot Temporal Action Detection

Sauradip Nag, Mengmeng Xu, Xiatian Zhu, Juan-Manuel Perez-Rua, Bernard Ghanem, Yi-Zhe Song, Tao Xiang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.08729 2022-12-20 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 67%

Distribution-aware Goal Prediction and Conformant Model-based Planning for Safe Autonomous Driving

Jonathan Francis, Bingqing Chen, Weiran Yao, Eric Nyberg, Jean Oh

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted: 1st Workshop on Safe Learning for Autonomous Driving, at the International Conference on Machine Learning (ICML 2022); Best Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.05836 2022-10-13 cs.CV cs.AI cs.CL cs.LG cs.MM 67%

GLIPv2: Unifying Localization and Vision-Language Understanding

Haotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen, Liunian Harold Li, Xiyang Dai, Lijuan Wang, Lu Yuan, Jenq-Neng Hwang, Jianfeng Gao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments NeurIPS 2022; updated with reviewers' comments addressed; Code is released at https://github.com/microsoft/GLIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.08425 2022-09-20 cs.LG cs.AI cs.CV 67%

Introspective Learning : A Two-Stage Approach for Inference in Neural Networks

Mohit Prabhushankar, Ghassan AlRegib

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03857 2022-06-20 cs.CV cs.AI cs.CL cs.LG cs.MM 67%

Grounded Language-Image Pre-training

Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, Kai-Wei Chang, Jianfeng Gao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments CVPR 2022; updated visualizations; fixed hyper-parameters in Appendix C.1

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04363 2022-06-09 cs.CV cs.AI cs.LG 67%

Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image Captioning

Chia-Wen Kuo, Zsolt Kira

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments paper accepted in CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04725 2022-05-13 cs.CV cs.AI cs.LG 67%

Weakly-supervised segmentation of referring expressions

Robin Strudel, Ivan Laptev, Cordelia Schmid

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.07207 2022-03-09 cs.LG cs.AI cs.CL cs.CV cs.RO 67%

Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

Wenlong Huang, Pieter Abbeel, Deepak Pathak, Igor Mordatch

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Project website at https://huangwl18.github.io/language-planner

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.07074 2021-12-06 cs.LG cs.AI cs.CL cs.CV 67%

Memotion Analysis through the Lens of Joint Embedding

Nethra Gunti, Sathyanarayanan Ramamoorthy, Parth Patwa, Amitava Das

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted as Student Abstract at AAAI-22

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04927 2021-11-05 cs.CV cs.AI cs.CL cs.LG 67%

Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion

Alessandro Suglia, Qiaozi Gao, Jesse Thomason, Govind Thattai, Gaurav Sukhatme

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at Novel Ideas in Learning-to-Learn through Interaction (NILLI) workshop @ EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.02681 2021-10-20 cs.CL cs.AI cs.CV cs.LG 67%

VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer

Zineng Tang, Jaemin Cho, Hao Tan, Mohit Bansal

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments NeurIPS 2021 (19 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.05769 2021-10-13 cs.CV cs.AI cs.LG cs.MA 67%

Interpretation of Emergent Communication in Heterogeneous Collaborative Embodied Agents

Shivansh Patel, Saim Wani, Unnat Jain, Alexander Schwing, Svetlana Lazebnik, Manolis Savva, Angel X. Chang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Project page: https://shivanshpatel35.github.io/comon/ ; the first three authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.14447 2021-06-29 cs.CV cs.AI cs.LG 67%

Feature Combination Meets Attention: Baidu Soccer Embeddings and Transformer based Temporal Detection

Xin Zhou, Le Kang, Zhiyu Cheng, Bo He, Jingyu Xin

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Tech Report. Authors Xin Zhou, Le Kang, and Zhiyu Cheng made equal contributions

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.15454 2021-01-01 cs.CL cs.AI cs.CV cs.LG eess.AS 67%

Text-Free Image-to-Speech Synthesis Using Learned Segmental Units

Wei-Ning Hsu, David Harwath, Christopher Song, James Glass

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏