Modular Framework for Visuomotor Language Grounding
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
Comments Published in LREC 2020. Publication URL https://www.aclweb.org/anthology/2020.lrec-1.268/; Dataset DOI https://doi.org/10.25835/0017546
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
Comments ACL 2020
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
Comments IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019
专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG
Comments NeurIPS 2019 Workshop ViGIL : Visually Grounded Interaction and Language
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
Comments Accepted to EMNLP 2019
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
Comments AAAI-19 Workshop on Games and Simulations for Artificial Intelligence
专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
Comments Accepted to ECCV 2018
Journal ref European Conference on Computer Vision (ECCV), 2018
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV
Comments CVPR 2017
专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG
Comments 7 pages, 4 figures, ICDL-Epirob 2015 conference
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
Journal ref Physica D 42: 335-346
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI
Journal ref Journal Of Artificial Intelligence Research, Volume 15, pages 31-90, 2001
从语料到协同进化的能力:面向通用图像生成的以能力为中心的数据设计
机构 * Alibaba Group(阿里巴巴集团)
专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn);分类 cs.CV、cs.AI
AI总结 该研究提出能力驱动的数据基础设施,构建三类监督引擎,整理大规模图像语料,训练多模态扩散模型,在CPI-Bench等评估中展现良好生成与迁移能力。
Comments 19 pages, 10 figures
扫视、审视与思考:将视频异常检测从无训练推进到智能体推理
机构 * Beijing Jiaotong University(北京交通大学) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 该研究针对视频异常检测的“何时-何事”脱节问题,提出无训练框架GtS及工具增强型智能体方法,扩展基准并引入联合评估指标,实现了更高的异常检测准确性与推理速度。
Comments 34 pages, 8 figures, 8 tables. Journal extension of our AAAI 2026 paper (arXiv:2507.21507)
SNR-Edit: 基于结构的噪声校正用于无反向的流式编辑
机构 * State Key lab of CAD\&CG, Zhejiang University(CAD与CG国家重点实验室,浙江大学) ; UniTTEC Co. Ltd.(UniTTEC有限公司) ; Research center, Ningbo Meidong Container Terminal Co.,Ltd.(宁波梅东集装箱码头有限公司研究中心)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);分类 cs.CV、cs.AI
AI总结 SNR-Edit通过自适应噪声控制实现无反向的流式编辑,提升潜在轨迹的结构保真度,同时保持高效率。
通过模态差距感知自蒸馏从符号状态学习视觉空间规划
机构 * Tsinghua University(清华大学) ; The Hong Kong University of Science and Technology(香港科技大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI
AI总结 提出MGSD两阶段框架,通过冷启动接地和特权教师蒸馏弥合视觉与符号规划之间的模态差距,在视觉规划基准上显著提升性能。
Comments 18 pages, preprint
面向智能体图像编辑的领域 grounded 候选选择:以阴影去除为例
机构 * Stony Brook University(石溪大学) ; UNC Charlotte(北卡罗来纳大学夏洛特分校)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI
AI总结 该研究以阴影去除为例,提出领域 grounded 的智能体候选选择流程,结合物理原理约束商用视觉-语言模型的生成,在 ShadowRemovalRefine 基准上使 CDD 降低至少 47%,证明经典低级视觉先验仍具实用价值。
动态目标掩码作为视觉目标条件强化学习的目标表示
机构 * University of Alberta(阿尔伯塔大学) ; McGill University(麦吉尔大学) ; University of Ottawa(渥太华大学) ; AMII ; CIFAR Canada AI Chair(CIFAR加拿大人工智能主席) ; Vector Institute(向量研究所)
专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn);分类 cs.CV、cs.LG
AI总结 该研究针对视觉目标条件强化学习,提出动态目标掩码作为目标表示,结合Detic等预训练检测器生成掩码,在机械臂和仿真导航任务中实现高成功率与高效学习,支持仿真到现实迁移。
为正确的边界框分配信用:结构化视觉感知的边际贡献分配
机构 * Amap, Alibaba Group(高德,阿里巴巴集团)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 针对多模态大语言模型结构化视觉感知中响应级监督的粒度不匹配问题,提出MCR-GRPO框架,通过边界框级边际贡献分配优化,在多个基准上取得了优于现有GRPO基线的最先进性能。
PokeGym: 一种以视觉驱动的长周期基准用于视觉-语言模型
机构 * SIAS, UESTC(电子科技大学深圳高等研究院) ; CFAR/IHPC A*STAR(新加坡科技研究局高性能计算研究所)
专题命中 视觉定位与Grounding :VLM(abstract_cn);grounding(abstract_cn);分类 cs.CV、cs.AI
AI总结 PokeGym通过复杂的3D开放世界游戏评估视觉-语言模型在长周期任务中的表现,揭示了物理死锁恢复是当前模型的主要瓶颈,并发现高级模型存在元认知分歧。
Comments Tech report
SciFigPlag-Bench:面向来源感知的科学图片抄袭检测基准
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.LG
AI总结 本文提出SciFigPlag-Bench基准,用于科学图片的来源感知抄袭检测,含四类任务,实验发现视觉语言模型在细粒度来源推理等方面仍有挑战。
Comments 30 pages, 18 figures
看清实际存在的东西:用于在视觉语言模型中对代理视觉证据进行反事实评估的PriVE-Bench和PriVE-Tools
机构 * The University of Manchester(曼彻斯特大学) ; The University of Melbourne(墨尔本大学) ; University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; University of Edinburgh(爱丁堡大学) ; University of Southern California(南加州大学) ; Fudan University(复旦大学) ; University of Newcastle(纽卡斯尔大学)
专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI
AI总结 本文针对视觉语言模型常依先验而非图像回答问题的现象,引入PriVE-Bench和PriVE-Tools。通过配对图像及工具衍生证据评估模型,对比不同输入,发现视觉证据工具在特定情况有帮助,但不能解决所有模型依先验回答的问题。
Groc-PO:用于真实多模态大语言模型的基于上下文的偏好优化
机构 * University of Science and Technology of China(中国科学技术大学) ; Xiaomi Corporation(小米公司) ; State Key Laboratory of Communication Content Cognition, People’s Daily Online(人民日报社传播内容认知国家重点实验室)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 研究针对多模态大语言模型的不真实问题,提出基于上下文的偏好优化框架Groc-PO,构建相关数据集,通过多阶段偏好样本捕捉基础上下文,加强上下文相关推理,减轻跨阶段错误传播,提升模型性能。
Comments Accepted by ACM-MM 2026
纠正注意力分散引起的视觉模糊以减少幻觉:算法与理论
机构 * National University of Singapore(新加坡国立大学) ; University of Science and Technology of China(中国科学技术大学)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 本文揭示多模态大语言模型中的物体幻觉与类人注意力分散现象相关,并提出一种无需额外训练的注意力聚焦方法(AFIP)通过跨头注意力增强和动态历史注意力强化来纠正视觉模糊,从而减少幻觉。
Journal ref ICML2026
ProCal:开放词汇目标检测的推理时提议校准
机构 * Korea University(高丽大学)
专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);分类 cs.CV、cs.AI
AI总结 提出ProCal方法,通过结合定位感知前景分数和背景感知抑制分数计算提议先验,在推理时校准分类分数,提升开放词汇目标检测中未见类别的定位质量。
CPG-PAD: 概念引导提示的呈现攻击检测
机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室) ; China Mobile Financial Technology Co., Ltd.(中移动金融科技有限公司) ; Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新研究院人工智能与机器人创新中心) ; School of Computer Science and Engineering, the Faculty of Innovation Engineering, Macau University of Science and Technology(澳门科技大学创新工程学院计算机科学与工程学院)
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);分类 cs.CV、cs.AI
AI总结 提出概念引导提示的呈现攻击检测框架,通过可解释AI发现攻击相关视觉概念并注入提示空间,提升跨域泛化能力,在多个基准数据集上达到最优。
Comments Accepted by IEEE Transactions on Information Forensics & Security (TIFS)