arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7348 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7348 篇

2604.17019 2026-04-21 cs.AI 82%

Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents

Mini-BEHAVIOR-Gran:揭示指令粒度对语言引导的具身智能体的影响

Sukai Huang, Chenyuan Zhang, Fucai Ke, Zhixi Cai, Gholamreza Haffari, Lizhen Qu, Hamid Rezatofighi

机构 * Faculty of Information Technology, Monash University(信息技术学院,莫纳什大学)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);分类 cs.AI

AI总结 本文提出Mini-BEHAVIOR-Gran基准,通过多级指令变体研究指令粒度对具身智能体性能的影响,发现粒度与性能呈非单调U型关系,粗粒度性能反弹与浅层 grounding 有关。

Comments 23 pages, Keywords: Language Grounding, Language Granularity, Instruction Following Agent, Width-based Planning Research Area: Multimodality and Language Grounding to Vision, Robotics and Beyond Research Area Keywords: vision language navigation, multimodality, neurosymbolic approaches

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28055 2026-07-31 cs.IR 新提交 82%

VIG-RL: Learning to Search and Insert for Verified Image Grounding

VIG-RL:学习搜索与插入以实现可验证的图像定位

Qinhan Yu, Jun Guang, Chong Chen, Wentao Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract)

AI总结 本文针对现有检索增强框架无法动态推理视觉证据插入时机与位置的问题,提出自主智能体框架VIG-RL,将相关工作流建模为主动决策过程,经强化学习优化后实现SOTA性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11018 2026-07-14 cs.RO 新提交 82%

Whole-Body Semantic-to-Actuation Grounding of Elephant-Inspired Soft-Trunk Motion via Lightweight Flow Matching

通过轻量级流匹配实现大象启发式软躯干运动的全身语义到驱动的基础

Tingcong Liu, Tongshun Chen, Siyi Ma, Yuhao Wang, Aye Phyu Phyu Aung, Ibrahim Alsarraj, J. Senthilnath, Bo An, Ke Wu

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract)

AI总结 针对近距离人机交互中类似躯干机器人关联开放词汇响应难的问题,提出基于轻量级流匹配的全身语义到驱动基础框架,将多模态大语言模型响应转换为元组,参数化轨迹并采样运动,实验显示其提升关联正确性、减少推理时间,还提升了人机交互满意度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20873 2026-06-23 cs.CL 新提交 82%

SciLens: Multi-modal Scientific Claim Verification with Agentic Entailment and Grounding

SciLens: 多模态科学声明验证的智能蕴含与归因框架

Yueming Wang, Tianshi Zheng, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See

机构 * The Hong Kong University of Science and Technology(香港科技大学) Hong Kong Baptist University(香港浸会大学) NVIDIA AI Technology Center(英伟达人工智能技术中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)

AI总结 提出SciLens框架,通过将声明分解为原子命题并归因到表格/图表证据,实现多模态科学声明验证,在SciClaimEval上达到79.2%宏F1和63.1%配对准确率。

Comments KDD 2026 SciSoc Agents & LLMs (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15427 2026-06-16 cs.LG cs.AI cs.CV 新提交 82%

Post-Launch Capability Expansion of Vision-Language Models via Prompting for On-Orbit Spacecraft Inspection

通过提示实现视觉语言模型发射后能力扩展用于在轨航天器检测

Nicholas A. Welsh, Lennon J. Shikhman, Monty Nehru Attazs, Seemanthini K. Putane, Van Minh Nguyen, Ryan T. White

机构 * Florida Institute of Technology(佛罗里达理工学院) University of Florida(佛罗里达大学)

专题命中 视觉定位与Grounding :vision-language model(title);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 研究利用提示驱动的视觉语言模型在轨扩展语义能力,无需修改权重即可通过自然语言提示检测新航天器部件,在129张图像上零样本实例分割达到0.385 mAP@0.5。

Comments 5 pages, 1 figure, 2 tables. Equal contribution by Nicholas A. Welsh and Lennon Shikhman. Published in the CVPR2026 Workshop on AI4Space

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18439 2026-06-02 cs.CL 82%

Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation

基于视觉线索检测手语翻译中的幻觉:是依据视觉信息还是猜测?

Yasser Hamidullah, Koel Dutta Chowdhury, Yusser Al Ghussin, Shakib Yazdani, Cennet Oguz, Josef van Genabith, Cristina España-Bonet

机构 * German Research Center for Artificial Intelligence (DFKI GmbH)(德国人工智能研究中心(DFKI GmbH)) Saarland Informatics Campus(萨尔兰州信息学校园) Barcelona Supercomputing Center (BSC-CNS)(巴塞罗那超级计算中心(BSC-CNS))

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)

AI总结 针对手语翻译中模型依赖语言先验而非视觉输入导致幻觉的问题,提出一种基于特征敏感性和反事实信号的令牌级可靠性度量,用于量化视觉信息利用程度,并在两个基准上验证其预测幻觉率、跨数据集泛化及与文本信号结合提升风险评估的效果。

Comments Published at ICLR2026 Code available at \url{https://github.com/yhamidullah/hallucination-slt}

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24972 2026-04-29 cs.CL 82%

Dynamic Decision Learning: Test-Time Evolution for Abnormality Grounding in Rare Diseases

动态决策学习:罕见疾病异常定位的测试时间演化

Jun Li, Mingxuan Liu, Jiazhen Pan, Che Liu, Wenjia Bai, Cosmin I. Bercea, Julia A. Schnabel

机构 * Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Imperial College London(帝国理工学院伦敦分校) University of Trento(特伦托大学) King's College London(伦敦国王学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)

AI总结 本文提出动态决策学习框架,通过优化指令和视觉扰动下的预测整合,提升冻结大视觉语言模型在罕见疾病异常定位中的表现,实验显示其在罕见疾病案例中mAP@75提升达105%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16342 2026-04-21 cs.HC 82%

SAGE: Sensor-Augmented Grounding Engine for LLM-Powered Sleep Care Agent

SAGE:基于传感器的地面引擎用于LLM驱动的睡眠护理代理

Hansoo Lee, Yoonjae Cho, Sonya S. Kwak, Rafael A. Calvo

专题命中 视觉定位与Grounding :grounding(title,abstract)

AI总结 SAGE通过整合传感器数据,解决睡眠护理中数据与行动之间的鸿沟问题,提升个性化和信任度。

Comments Accepted to the Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA '26). 6 pages

Journal ref Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25537 2026-03-27 cs.CL 82%

Humans vs Vision-Language Models: A Unified Measure of Narrative Coherence

人类与视觉-语言模型:叙事连贯性的一种统一衡量方法

Nikolai Ilinykh, Hyewon Jang, Shalom Lappin, Asad Sayeed, Sharid Loáiciga

机构 * University of Gothenburg(哥德堡大学) Queen Mary University of London(伦敦玛丽女王大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract)

AI总结 研究通过比较人类写作与视觉语言模型生成的叙事,评估视觉基础故事中的叙事连贯性,发现模型在连贯性方面与人类存在系统性差异。

Comments 9 pages of content, 1 page of appendices, 9 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20193 2026-03-23 cs.CV cs.AI cs.LG 82%

From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering

从遮罩到像素与意义:一种新的分类、基准和度量标准用于VLM图像篡改

Xinyi Shang, Yi Tang, Jiacheng Cui, Ahmed Elhagry, Salwa K. Al Khatib, Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Jing-Hao Xue, Hao Li, Salman Khan, Zhiqiang Shen

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) University College London(伦敦大学学院)

专题命中 视觉定位与Grounding :VLM(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出了一种基于像素和意义的VLM图像篡改检测新方法,引入了新的分类体系和基准,提出了基于像素的度量标准,并评估了现有模型在微编辑和非遮罩篡改上的表现。

Comments Code and data at: https://github.com/VILA-Lab/PIXAR (Accepted in CVPR 2026 Findings, but not opted in)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18210 2026-03-20 cs.RO 82%

GoalVLM: VLM-driven Object Goal Navigation for Multi-Agent System

GoalVLM: 多智能体系统中基于视觉语言模型的目标导航

MoniJesu James, Amir Atef Habel, Aleksey Fedoseev, Dzmitry Tsetserokou

机构 * Intelligent Space Robotics Laboratory(智能空间机器人实验室) Center for Digital Engineering(数字工程中心) Skolkovo Institute of Science and Technology(斯克尔科夫科学与技术研究所)

专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract)

AI总结 本文提出GoalVLM,一种基于视觉语言模型的多智能体零样本开放词汇目标导航框架,通过整合SAM3和SpaceOM实现语义优先级探索,无需重新训练即可完成复杂目标导航任务。

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15134 2026-03-17 cs.RO 82%

Confusion-Aware In-Context-Learning for Vision-Language Models in Robotic Manipulation

面向机器人操作的上下文学习:考虑混淆的视觉语言模型

Yayun He, Zuheng Kang, Botao Zhao, Zhouyin Wu, Junqing Peng, Jianzong Wang

机构 * Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司) Shenzhen Bao'an Middle School(深圳宝安中学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract)

AI总结 本文提出CAICL方法,通过混淆定位与分析提升视觉语言模型在机器人操作中处理混淆场景的能力,实验显示其在VIMA-Bench上成功率达85.5%。

Comments Accepted by the 29th International Conference on Computer Supported Cooperative Work in Design (CSCWD 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03101 2026-03-04 cs.LG cs.AI cs.CV 82%

ALARM: Automated MLLM-Based Anomaly Detection in Complex-EnviRonment Monitoring with Uncertainty Quantification

ALARM: 基于多模态大语言模型的复杂环境监控异常检测与不确定性量化

Congjing Zhang, Feng Lin, Xinyi Zhao, Pei Guo, Wei Li, Lin Chen, Chaoyue Zhao, Shuai Huang

机构 * Department of Industrial and Systems Engineering, University of Washington(华盛顿大学工业与系统工程系) Wyze Labs, Inc.(Wyze实验室)

专题命中 视觉定位与Grounding :MLLM(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 ALARM通过结合不确定性量化与多模态大语言模型,实现了复杂环境中的异常检测,展现出在不同领域中的高准确性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16072 2026-02-20 cs.RO 82%

I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models

I-FailSense:面向通用机器人故障检测的视觉-语言模型

Clemence Grislain, Hamed Rahimi, Olivier Sigaud, Mohamed Chetouani

机构 * ISIR, Sorbonne Université, CNRS(ISIR,索邦大学,国家科学研究中心)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract)

AI总结 I-FailSense通过构建专门用于检测语义错位故障的数据集,提出了一种开源视觉-语言模型框架,有效提升机器人故障检测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16793 2026-02-05 cs.RO 82%

PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence

PhysBrain: 人眼视角数据作为视觉语言模型到物理智能的桥梁

Xiaopeng Lin, Shijie Lian, Bin Yu, Ruoqi Yang, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Yurun Jin, Yukun Shi, Jiyan He, Cong Huang, Bojun Cheng, Kai Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Zhongguancun Academy(中关村学院) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院) Harbin Institute of Technology(哈尔滨工业大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 视觉定位与Grounding :vision language model(title,abstract);grounding(abstract)

AI总结 PhysBrain通过将人类眼动视频转化为多级具身监督,提升机器人在眼动感知和长期规划中的能力,实现从人类视角到机器人控制的有效迁移。

Comments 21 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17507 2026-01-27 cs.RO 82%

MetaWorld: Skill Transfer and Composition in a Hierarchical World Model for Grounding High-Level Instructions

MetaWorld: 一个用于地面指令基础的分层世界模型中的技能迁移与组合

Yutong Shen, Hangxu Liu, Kailin Pei, Ruizhe Xia, Tongtong Feng

机构 * Beijing University of Technology(北京理工大学) Fudan University(复旦大学) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract)

AI总结 MetaWorld通过整合语义规划与物理控制,利用专家策略迁移提升人形机器人在定位-操作任务中的性能。

Comments 8 pages, 4 figures, Submitted to ICLR 2026 World Model Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20876 2026-01-13 cs.RO 82%

Proprioception Enhances Vision Language Model in Generating Captions and Subtask Segmentations for Robot Task

本体感知增强视觉语言模型在为机器人任务生成描述和子任务分割中的应用

Kanata Suzuki, Shota Shimizu, Tetsuya Ogata

机构 * Faculty of Science and Engineering, Waseda University(工学部,早稻田大学) Artificial Intelligence Laboratory, Fujitsu Limited(Fujitsu 人工智能实验室) National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究院)

专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract)

AI总结 本研究通过引入本体感知数据,提升视觉语言模型在机器人任务描述和子任务分割中的性能,以增强机器人模仿学习效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27680 2025-12-02 cs.CV cs.AI cs.LG 82%

PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting

PETAR:基于掩码感知的视觉-语言建模的局部发现生成用于PET自动报告

Danyal Maqbool, Changhee Lee, Zachary Huemann, Samuel D. Church, Matthew E. Larson, Scott B. Perlman, Tomas A. Romero, Joshua D. Warner, Meghan Lubner, Xin Tie, Jameson Merkow, Junjie Hu, Steve Y. Cho, Tyler J. Bradshaw

机构 * University of Wisconsin–Madison Department of Computer Sciences(威斯康星大学麦迪逊分校计算机科学系) University of Wisconsin–Madison Department Radiology(威斯康星大学麦迪逊分校放射学系) Microsoft(微软公司)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 PETAR通过引入PETARSeg-11K数据集和PETAR-4B模型,实现基于掩码感知的3D PET自动报告生成,提升医学影像分析的精度与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19315 2025-11-25 cs.RO 82%

Rethinking Intermediate Representation for VLM-based Robot Manipulation

重新思考基于VLM的机器人操作中的中间表示

Weiliang Tang, Jialin Gao, Jia-Hui Pan, Gang Wang, Li Erran Li, Yunhui Liu, Mingyu Ding, Pheng-Ann Heng, Chi-Wing Fu

机构 * CUHK(香港中文大学) Amazon(亚马逊) UNC(北卡罗来纳大学教堂山分校)

专题命中 视觉定位与Grounding :VLM(title,abstract);vision-language model(abstract)

AI总结 本文提出SEAM表示方法,通过分解中间表示为词汇和语法,提升VLM在机器人操作中的可理解和通用性,结合检索增强的少样本学习策略实现高效操作,并在动作通用性和VLM可理解性上展示出优于主流方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17358 2025-11-24 cs.CL 82%

Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding

不要学习,而是依托:自然语言推理与视觉依托的案例

Daniil Ignatev, Ayman Santeer, Albert Gatt, Denis Paperno

机构 * Utrecht University(乌特勒支大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract)

AI总结 本文提出了一种基于视觉依托的零样本自然语言推理方法,通过生成视觉表示并比较与假设的相似度,实现高精度推理,展示了对文本偏见的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17359 2025-11-04 cs.IR 82%

MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval

Tianyuan Li, Lei Wang, Ahtamjan Ahmat, Yating Yang, Bo Ma, Rui Dong, Bangju Han

专题命中 视觉定位与Grounding :MLLM(title);grounding(abstract);multimodal large language model(abstract)

Comments We plan to revise the methodology and update the experimental analysis before resubmission

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21445 2025-10-27 cs.CL cs.AI cs.CV cs.LG 82%

REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring

Thanh Cong Ho, Farah Kharrat, Abderrazek Abid, Fakhri Karray

机构 * 2 Department of Electrical Computer Engineering University of Waterloo, Waterloo, ON, Canada N2L 3G1 Email 3 College of Computer Information Sciences Prince Sultan University Email

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Journal ref 2024 IEEE International Symposium on Medical Measurements and Applications (MeMeA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16924 2025-10-21 cs.CL 82%

Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?

Zhihui Yang, Yupei Wang, Kaijie Mo, Zhe Zhao, Renfen Hu

机构 * Beijing Normal University(北京师范大学) Tencent AI Lab(腾讯AI实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)

Comments Accepted to EMNLP 2025 (Findings). This version corrects a redundant sentence in the Results section that appeared in the camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11302 2025-10-21 cs.CV cs.AI cs.LG 82%

When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models

Samer Al-Hamadani

机构 * Automated Manufacturing Department(自动化制造部门) Al-Khwarizmi College of Engineering(阿尔·卡瓦尔米工程学院) University of Baghdad(巴格达大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments 30 pages, 12 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16455 2025-10-21 cs.CL 82%

RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning

Deyi Ji, Yuekui Yang, Haiyang Wu, Shaoping Ma, Tianrun Chen, Lanyun Zhu

机构 * Tencent(腾讯公司) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Zhejiang University(浙江大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract)

Comments ACL 2025 (Oral, Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13572 2025-09-18 cs.RO 82%

Using Visual Language Models to Control Bionic Hands: Assessment of Object Perception and Grasp Inference

Ozan Karaali, Hossam Farag, Strahinja Dosen, Cedomir Stefanovic

机构 * Department of Electronic Systems, Aalborg University, Denmark(电子系统系,奥胡斯大学) Department of Health Science and Technology, Aalborg University, Denmark(健康科学与技术系,奥胡斯大学)

专题命中 视觉定位与Grounding :visual language model(title);vision language model(abstract);VLM(abstract)

Comments ICAT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08142 2025-09-11 eess.SP 82%

Privacy Preserving Semantic Communications Using Vision Language Models: A Segmentation and Generation Approach

Haoran Chang, Mingzhe Chen, Huaxia Wang, Qianqian Zhang

专题命中 视觉定位与Grounding :vision language model(title);vision-language model(abstract);VLM(abstract)

Comments 6 pages, 6 figures, Accepted at IEEE MILCOM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00669 2025-08-14 cs.LG cs.AI cs.CV cs.RO 82%

Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding

Duc Cao-Dinh, Khai Le-Duc, Anh Dao, Bach Phan Tat, Chris Ngo, Duy M. H. Nguyen, Nguyen X. Khanh, Thanh Nguyen-Tang

机构 * Hanyang University(汉阳大学) University of Toronto(多伦多大学) University Health Network(大学健康网络) Knovel Engineering Lab(Knovel工程实验室) Michigan State University(密歇根州立大学) KU Leuven(鲁汶大学) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心) Max Planck Research School for Intelligent Systems (IMPRS-IS)(马克斯·普朗克智能系统研究学校) University of Stuttgart(斯图加特大学) UC Berkeley(伯克利大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments Preprint, 51 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05021 2025-08-08 cs.RO 82%

MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding

Weifan Zhang, Tingguang Li, Yuzhen Liu

机构 * Tencent Robotics X(腾讯机器人X) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);visual language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01723 2025-08-05 cs.RO 82%

OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping

Danyang Li, Zenghui Yang, Guangpeng Qi, Songtao Pang, Guangyong Shang, Qiang Ma, Zheng Yang

机构 * School of Software, Tsinghua University(清华大学软件学院) School of computer science and engineering, Central South University(中南大学计算机科学与工程学院) Inspur Yunzhou Industrial Internet Co., Ltd(Inspur Yunzhou工业互联网有限公司) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract)

Comments ACM MM '25

详情

展开后加载摘要…

URL PDF HTML 收藏