arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7348 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7348 篇

2103.09720 2021-04-01 cs.CV cs.AI 81%

Few-Shot Visual Grounding for Natural Human-Robot Interaction

Giorgos Tziafas, Hamidreza Kasaei

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments 6 pages, 4 figures, ICARSC2021 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.01449 2021-03-11 cs.CV cs.AI cs.CL cs.MM 81%

Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding

Long Chen, Wenbo Ma, Jun Xiao, Hanwang Zhang, Shih-Fu Chang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments Camera ready version at AAAI 2021, Codes are available at: https://github.com/ChopinSharp/ref-nms

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.07185 2021-01-26 cs.AI cs.LG stat.ML 81%

Grounding Language to Autonomously-Acquired Skills via Goal Generation

Ahmed Akakzia, Cédric Colas, Pierre-Yves Oudeyer, Mohamed Chetouani, Olivier Sigaud

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

Comments Published at ICLR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.05208 2021-01-14 cs.CV cs.AI cs.CL 81%

Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps Grounding

Dexin Wang, Deyi Xiong

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.09920 2020-08-07 cs.CV cs.CL cs.LG stat.ML 81%

Contrastive Learning for Weakly Supervised Phrase Grounding

Tanmay Gupta, Arash Vahdat, Gal Chechik, Xiaodong Yang, Jan Kautz, Derek Hoiem

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.LG

Comments ECCV 2020 (spotlight paper), Project page: http://tanmaygupta.info/info-ground

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.01470 2020-07-13 cs.LG cs.AI cs.MA cs.NE stat.ML 81%

Options as responses: Grounding behavioural hierarchies in multi-agent RL

Alexander Sasha Vezhnevets, Yuhuai Wu, Remi Leblond, Joel Z. Leibo

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

Comments First two authors contributed equally

Journal ref International Conference on Machine Learning 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.04304 2020-07-09 cs.CL cs.AI cs.LG cs.RO 81%

Unsupervised Online Grounding of Natural Language during Human-Robot Interactions

Oliver Roesler

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

Comments 11 pages, 6 figures, 3 tables; Published in Proceedings of the Second Grand Challenge and Workshop on Multimodal Language (Challenge-HML) in the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.07043 2020-06-15 cs.LG cs.AI cs.CL stat.ML 81%

Language-Conditioned Goal Generation: a New Approach to Language Grounding for RL

Cédric Colas, Ahmed Akakzia, Pierre-Yves Oudeyer, Mohamed Chetouani, Olivier Sigaud

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.03776 2020-06-09 cs.CV cs.CL cs.LG 81%

MAGNet: Multi-Region Attention-Assisted Grounding of Natural Language Queries at Phrase Level

Amar Shrestha, Krittaphat Pugdeethosapol, Haowen Fang, Qinru Qiu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.LG

Comments Submitted to The 2020 European Conference on Computer Vision (ECCV 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.05078 2020-03-27 cs.CV cs.CL cs.LG 81%

Visual Grounding in Video for Unsupervised Word Translation

Gunnar A. Sigurdsson, Jean-Baptiste Alayrac, Aida Nematzadeh, Lucas Smaira, Mateusz Malinowski, João Carreira, Phil Blunsom, Andrew Zisserman

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.LG

Comments CVPR 2020

Journal ref CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.02133 2019-11-07 cs.CV cs.CL cs.LG 81%

Contextual Grounding of Natural Language Entities in Images

Farley Lai, Ning Xie, Derek Doran, Asim Kadav

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.LG

Comments Accepted to NeurIPS 2019 workshop on Visually Grounded Interaction and Language (ViGIL)

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.12354 2019-10-29 cs.CL cs.AI cs.LG 81%

Task-Oriented Language Grounding for Language Input with Multiple Sub-Goals of Non-Linear Order

Vladislav Kurenkov, Bulat Maksudov, Adil Khan

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.06315 2019-10-15 cs.CV cs.CL cs.LG 81%

Dynamic Attention Networks for Task Oriented Grounding

Soumik Dasgupta, Badri N. Patro, Vinay P. Namboodiri

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.LG

Comments Accepted ICCV 2019 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.11683 2019-05-31 cs.CV cs.CL cs.LG eess.IV 81%

Multi-level Multimodal Common Semantic Space for Image-Phrase Grounding

Hassan Akbari, Svebor Karaman, Surabhi Bhargava, Brian Chen, Carl Vondrick, Shih-Fu Chang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.LG

Comments Accepted in CVPR 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.11277 2019-03-28 cs.LG cs.AI 81%

Towards Stable Symbol Grounding with Zero-Suppressed State AutoEncoder

Masataro Asai, Hiroshi Kajino

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

Comments Accepted in 29th International Conference of Automated Planning and Scheduling (ICAPS-2019), Planning and Learning track

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.08093 2019-03-28 cs.AI cs.LG 81%

Unsupervised Grounding of Plannable First-Order Logic Representation from Images

Masataro Asai

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

Comments Accepted in 29th International Conference of Automated Planning and Scheduling (ICAPS-2019), Planning and Learning track

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.08329 2018-09-06 cs.AI cs.CL cs.LG cs.RO 81%

Guided Feature Transformation (GFT): A Neural Language Grounding Module for Embodied Agents

Haonan Yu, Xiaochen Lian, Haichao Zhang, Wei Xu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

Comments CoRL 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.07230 2018-01-10 cs.LG cs.AI cs.CL cs.RO 81%

Gated-Attention Architectures for Task-Oriented Language Grounding

Devendra Singh Chaplot, Kanthashree Mysore Sathyendra, Rama Kumar Pasumarthi, Dheeraj Rajagopal, Ruslan Salakhutdinov

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

Comments To appear in AAAI-18

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.05293 2017-09-18 cs.RO cs.AI cs.CV 81%

Commonsense Scene Semantics for Cognitive Robotics: Towards Grounding Embodied Visuo-Locomotive Interactions

Jakob Suchan, Mehul Bhatt

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments to appear in: ICCV 2017 Workshop - Vision in Practice on Autonomous Robots (ViPAR), International Conference on Computer Vision (ICCV), Venice, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.03745 2017-02-21 cs.CV cs.CL cs.LG 81%

Grounding of Textual Phrases in Images by Reconstruction

Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, Bernt Schiele

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.LG

Comments published at ECCV 2016 (oral); updated to final version

详情

展开后加载摘要…

URL PDF HTML 收藏
1508.03276 2015-08-14 cs.AI cs.CL cs.CV cs.HC 81%

Talking about the Moving Image: A Declarative Model for Image Schema Based Embodied Perception Grounding and Language Generation

Jakob Suchan, Mehul Bhatt, Harshita Jhavar

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

Comments 19 pages. Unpublished report

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15650 2026-07-13 cs.CV 版本更新 80%

AffordanceSAM: Segment Anything Once More in Affordance Grounding

可负担性语义分割模型:在可负担性基础上再次分割任何事物

Dengyang Jiang, Zanyi Wang, Hengzhuang Li, Sizhe Dang, Teli Ma, Wei Wei, Guang Dai, Lei Zhang, Harry Yang, Mengmeng Wang

机构 * SGIT AI Lab(SGIT人工智能实验室) NWPU(西北工业大学) HKUST(香港科技大学) XJTU(西安交通大学) HUST(华中科技大学) ZJUT(浙江工业大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 研究聚焦全监督可负担性基础,提出AffordanceSAM,通过设计适应模块和标注数据集,以三阶段训练方式扩展SAM泛化能力,在AGD20K基准上达最优性能,展现强大泛化能力。

Comments [ACM MM 2026] SAM Meets Affordance Grounding

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02049 2026-07-03 cs.CL cs.AI cs.CY 新提交 80%

SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses

SPLIT:英语和乌克兰语LLM回应中的跨语言共情与文化根基

Anna Chorna

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 提出SPLIT基准测试,评估LLM在英语和乌克兰语危机情境下的共情准确性、语言自然性和文化根基,发现模型在乌克兰语中表现下降,且人类与AI评估者一致性弱。

Comments 19 pages, 5 figures, 3 tables. Benchmark paper introducing SPLIT for evaluating empathy, linguistic naturalness, and cultural grounding in English and Ukrainian LLM responses

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05199 2026-04-29 cs.CV 80%

DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework

DEGround:一种有效的基于自身视角的3D视觉定位基线框架

Yani Zhang, Dongming Wu, Hao Shi, Yingfei Liu, Tiancai Wang, Xingping Dong

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) MMLab, The Chinese University of Hong Kong(香港中文大学MMLab) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出DEGround框架,通过统一的框架实现检测与定位的物体级共享,引入两个任务特定模块提升细粒度指令定位性能,实验表明在多个基准上表现最佳,尤其在EmbodiedScan数据集上精度提升显著。

Comments 1st place on EmbodiedScan visual grounding

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02662 2026-02-02 cs.LG 80%

Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning

通过在线强化学习将大语言模型接地于交互环境

Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, Pierre-Yves Oudeyer

机构 * Inria (Flowers)(Inria(Flowers)) University of Bordeaux(波尔多大学) Hugging Face Univ Angers, LERIA, SFR MATHSTIC(昂热大学,LERIA,SFR MATHSTIC) Sorbonne Université(索邦大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

AI总结 本文提出GLAM方法,通过在线强化学习将大语言模型接地于交互环境,以提升样本效率和泛化能力,并探讨在线学习的影响。

Comments The associated code can be found at https://github.com/flowersteam/Grounding_LLMs_with_online_RL. This is an extended version of the paper published at ICML 2023: https://proceedings.mlr.press/v202/carta23a

Journal ref PMLR 202 (2023):3676-3713

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04897 2025-10-29 cs.CV 80%

From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes

Tianxu Wang, Zhuofan Zhang, Ziyu Zhu, Yue Fan, Jing Xiong, Pengxiang Li, Xiaojian Ma, Qing Li

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) Tsinghua University(清华大学) Peking University(北京大学) Beijing Institute of Technology(北京理工大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV;multimodal large language model(comments)

Comments Update v3 of the NeurIPS 2025 Datasets and Benchmarks paper (v2), including additional evaluations of state-of-the-art multimodal large language models. Project page: https://anywhere-3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16900 2025-02-18 cs.RO cs.AI cs.CL cs.HC 80%

A Roadmap for Embodied and Social Grounding in LLMs

Sara Incao, Carlo Mazzola, Giulia Belgiovine, Alessandra Sciutti

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments Accepted Version of a conference paper presented at Robophilosophy Conference 2024

Journal ref Incao, S., Mazzola, C., Belgiovine, G., Sciutti, A., 2025, A Roadmap for Embodied and Social Grounding in LLMs. In J. Seibt, P. Fazekas, & O. S. Quick (Eds.), Social Robots with AI: Prospects, Risks, and Responsible Methods, IOS Press

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11380 2024-08-22 cs.RO cs.AI cs.SY eess.SY 80%

Reflex-Based Open-Vocabulary Navigation without Prior Knowledge Using Omnidirectional Camera and Multiple Vision-Language Models

Kento Kawaharazuka, Yoshiki Obinata, Naoaki Kanazawa, Naoto Tsukamoto, Kei Okada, Masayuki Inaba

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.AI;VLM(comments)

Comments Accepted at Advanced Robotics, website - https://haraduka.github.io/omnidirectional-vlm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00848 2024-06-04 cs.CV 80%

Eating Smart: Advancing Health Informatics with the Grounding DINO based Dietary Assistant App

Abdelilah Nossair, Hamza El Housni

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments The work presented in this paper was part of the proceedings for the First International Conference on Artificial Intelligence (ICATA 2024)

Journal ref Eating Smart: Advancing Health Informatics with the Grounding DINO-based Dietary Assistant App, International Journal of Scientific and Innovative Studies, June 2024, Volume 3, Number 3, Pages 26-34, Available online at IJSRIS

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12198 2024-03-26 cs.CV 80%

Mask Grounding for Referring Image Segmentation

Yong Xien Chng, Henry Zheng, Yizeng Han, Xuchong Qiu, Gao Huang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted by CVPR2024; Project page: https://yxchng.github.io/projects/mask-grounding

详情

展开后加载摘要…

URL PDF HTML 收藏