COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
Comments CVPR 2025
专题命中 视觉定位与Grounding :vision-language model(title);vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Comments preprint
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments Accepted by CVPR 2025
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.LG
Comments Accepted by CVPR 2025
专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments ACL 2023 Outstanding Paper
专题命中 视觉定位与Grounding :grounding(title);vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV、cs.LG
专题命中 视觉定位与Grounding :visual language model(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI
Comments accepted by SafeGenAI workshop of NeurIPS 2024
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.LG
Comments Accepted for publication at WACV 2025
专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(abstract);分类 cs.CV、cs.LG
Comments 7 pages, 3 figures
专题命中 视觉定位与Grounding :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments Under peer review
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments ACL 2024 (Findings)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract);分类 cs.CV、cs.LG
Comments ICLR 2024 Spotlight
专题命中 视觉定位与Grounding :multimodal large language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments CVPR 2024 camera-ready version, code is available at https://github.com/RenShuhuai-Andy/TimeChat
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
Comments Preprint. 48 pages, 22 figures, 10 tables
专题命中 视觉定位与Grounding :grounding(title,abstract);LLaVA(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
专题命中 视觉定位与Grounding :grounding(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI
Comments In CVPR 2023
专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI
Comments Camera ready for Findings of EMNLP 2021
专题命中 视觉定位与Grounding :grounding(title,abstract);visual reasoning(abstract);分类 cs.AI、cs.LG
Comments Code available at https://github.com/SeverTopan/SATNet
视觉定位:一项综述
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
AI总结 本综述梳理视觉定位的发展与背景,总结近年进展与新挑战,定义规范研究设置,介绍相关数据集与应用,提出未来方向,是该领域最全面的综述,适合不同阶段研究者。
Comments Accepted by TPAMI 2025. We keep tracing related works at https://github.com/linhuixiao/Awesome-Visual-Grounding, article publication page: https://ieeexplore.ieee.org/abstract/document/11235566
Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 3, pp. 2749-2771, March 2026
MM-Conv: 一种多模态数据集和基准,用于上下文感知的3D对话中指代解析
机构 * KTH Royal Institute of Technology(皇家理工学院)
专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);分类 cs.CV;VLM(comments)
AI总结 本文提出了一种多模态数据集和基准,用于在动态3D环境中实现上下文感知的指代解析,通过引入包含6.7小时第一人称VR交互的同步语音、动作、注视和3D场景几何数据的基准,以及一个两阶段的指代解析流水线,改进了对话中的指代解析性能。
Comments Extended version of the paper published at LREC 2026 (Palma de Mallorca, Spain), with expanded VLM baselines and inter-annotator agreement analysis
Journal ref Proceedings of the 15th Language Resources and Evaluation Conference (LREC 2026), Palma de Mallorca, Spain
FARM: 使用关系空间记忆找到任何物体
机构 * UC Berkeley(加州大学伯克利分校) ; Stanford University(斯坦福大学)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);grounding(abstract)
AI总结 提出FARM系统,通过实时构建包含几何、视觉语言描述和视角证据的开放词汇物体级记忆,并利用VLM解析查询和显式空间约束,在44k语言查询中Recall@5和Recall@10分别提升164%和224%,Accuracy@1提升35%。
CaptionFormer:时空对象的统一分割、跟踪与描述
机构 * Inria, École Normale Supérieure, CNRS, PSL Research University(法国国家科学研究中心、巴黎高等师范学院、国家科学研究中心、巴黎综合理工研究所) ; Google DeepMind(谷歌DeepMind)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 提出 CaptionFormer 模型,通过利用 VLM 生成合成描述并扩展数据集,实现视频中对象轨迹的联合检测、分割、跟踪与描述,在三个基准上达到最优。
Comments 17 pages, 10 figures
SceneAligner: 在真实场景中实现基于3D的平面定位
机构 * Cornell University(康奈尔大学) ; Kempner Institute, Harvard University(哈佛大学 Kempner 院)
专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 本文提出了一种在真实场景中实现基于3D重建的平面定位方法,通过将任务 grounding 在场景的重建3D表示中,解决了现有方法在大规模建筑和栅格化平面图中应用受限的问题。
Comments Project Page: https://Cornell-VAILab.github.io/SceneAligner