arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1267 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 340 篇

2510.14538 2026-08-07 cs.AI cs.LG 版本更新 81%

Symbol Grounding in Neuro-Symbolic AI: A Gentle Introduction to Reasoning Shortcuts

神经符号AI中的符号接地:推理快捷方式的入门介绍

Emanuele Marconato, Samuele Bortolotti, Emile van Krieken, Paolo Morettin, Elena Umili, Antonio Vergari, Efthymia Tsamoura, Andrea Passerini, Stefano Teso

机构 * University of Trento(特伦托大学) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) Sapienza University of Rome(罗马大学) University of Edinburgh(爱丁堡大学) Huawei Labs(华为实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI、cs.LG

AI总结 本文探讨神经符号AI中推理快捷方式的问题,分析其成因与影响,并提供解决方法与策略,以提升模型的可靠性和可信度。

Comments Published on JAIR (Integration of Logical Constraints in Deep Learning special track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23876 2026-07-08 cs.CV cs.AI 版本更新 81%

Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance

基于信息基础引导的视觉自回归采样再思考

Ky Dan Nguyen, Hoang Lam Tran, Anh-Dung Dinh, Daochang Liu, Weidong Cai, Xiuying Wang, Chang Xu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 研究基于下一尺度预测的自回归模型在图像生成中因信息不一致致引导信号分散问题,提出信息基础引导(IGG)框架,通过动态加权将引导锚定到语义重要令牌,在相关任务中生成更优图像,有效纠正基于AR的方法。

Comments Accepted to The Forty-Third International Conference on Machine Learning (ICML 2026); 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06294 2026-06-23 cs.CV cs.AI 版本更新 81%

Towards One-to-Many Temporal Grounding

面向一对多时间定位

Qi Xu, Yue Tan, Shihao Chen, Jiahao Meng, Anna Wang, Shunping Ji, Hao Fei, Jason Li

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 针对一对多时间定位(OMTG)任务,提出包含基准、数据集和奖励函数的系统解决方案,显著提升多段视频定位性能。

Comments Accepted to ICML'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15650 2026-07-13 cs.CV 版本更新 80%

AffordanceSAM: Segment Anything Once More in Affordance Grounding

可负担性语义分割模型:在可负担性基础上再次分割任何事物

Dengyang Jiang, Zanyi Wang, Hengzhuang Li, Sizhe Dang, Teli Ma, Wei Wei, Guang Dai, Lei Zhang, Harry Yang, Mengmeng Wang

机构 * SGIT AI Lab(SGIT人工智能实验室) NWPU(西北工业大学) HKUST(香港科技大学) XJTU(西安交通大学) HUST(华中科技大学) ZJUT(浙江工业大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 研究聚焦全监督可负担性基础,提出AffordanceSAM,通过设计适应模块和标注数据集,以三阶段训练方式扩展SAM泛化能力,在AGD20K基准上达最优性能,展现强大泛化能力。

Comments [ACM MM 2026] SAM Meets Affordance Grounding

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23733 2026-08-13 cs.CL 版本更新 80%

Multimodal QUD: Inquisitive Questions from Scientific Figures

多模态QUD:来自科学图表的探究性问题

Yating Wu, William Rudman, Venkata S Govindarajan, Alexandros G. Dimakis, Junyi Jessy Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Ithaca College(伊萨卡学院) UC Berkeley, BespokeLabs.ai(伯克利大学,BespokeLabs.ai)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract_cn);VLM(abstract_cn)

AI总结 本文提出多模态QUD数据集,通过结合图表与文本上下文生成探究性问题,提升多模态推理能力,实现更高质量的视觉 grounding。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07343 2026-06-16 cs.CV cs.AI cs.LG cs.RO 版本更新 80%

Seeing Roads Through Words: A Language-Guided Framework for RGB-T Driving Scene Segmentation

通过文字看道路:一种语言引导的RGB-T驾驶场景分割框架

Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy

机构 * National University of Singapore(新加坡国立大学) University of Technology Sydney(悉尼科技大学)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 提出CLARITY框架,利用视觉语言模型先验动态调整RGB-T融合策略,并引入暗目标语义保留和层次化解码器,在MFNet数据集上达到62.3% mIoU和77.5% mAcc的新SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11839 2026-08-13 cs.LG 版本更新 79%

Grounding Large Language Models as Generalizable Policies in Network Control

大语言模型作为网络优化的通用策略

Duo Wu, Linjia Kang, Zhimin Wang, Fangxin Wang, Wei Zhang, Chongbo Sun, Xuefeng Tao, Wei Yang, Le Zhang, Wenwu Zhu, Peng Cui, Zhi Wang

机构 * Bytedance(字节跳动) Shenzhen International Graduate School(深圳国际研究生院) Tsinghua University(清华大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Department of Computer Science and Technology(计算机科学与技术系)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

AI总结 本文提出Trailblazer框架,利用大语言模型实现跨任务和环境的通用网络策略,通过网络对齐和策略协作机制提升效率与泛化能力。

Comments Arxiv version. Official version has been submitted to IEEE Transactions on Mobile Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17162 2026-08-12 cs.AI q-bio.GN 版本更新 79%

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

JEPA-DNA:通过联合嵌入预测架构夯实基因组基础模型

Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit

机构 * Applied AI Architecture, NVIDIA, Israel(NVIDIA应用人工智能架构,以色列) Worldwide Field Ops, NVIDIA, Israel(NVIDIA全球现场运营,以色列) Developer Programs, NVIDIA, Israel(NVIDIA开发者计划,以色列) Cancer Research Center and Wohl Institute of Translational Medicine, Sheba Medical Center, Tel Hashomer, Israel(癌症研究中心和Wohl转化医学研究所,Sheba医疗中心,Tel Hashomer,以色列) Windreich Department of AI and Human Health, Icahn School of Medicine at Mount Sinai, New York, USA(AI与人类健康风reich部门,Mount Sinai医学中心,纽约,美国)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 提出JEPA-DNA框架,将联合嵌入预测架构与生成式目标结合,通过潜在空间监督全局序列嵌入,实现从令牌恢复到语义对齐的转变,在17项基因组基准任务上提升线性探测和零样本性能,达到新最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12602 2026-08-11 cs.CV 版本更新 79%

Decouple and Reason: Anatomically Guided Two-Stage Voxel-Level Grounding of Free-Text Findings in 3D Chest CT

解耦与推理:3D胸部CT中自由文本发现的解剖学引导两阶段体素级定位

Kwang-Hyun Uhm, Inhwa Son, Sung-Jea Ko

机构 * Department of Artificial Intelligence, Gachon University(韩国加图立大学人工智能系) MEDAI

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 研究3D胸部CT中自由文本发现的体素级定位难题,提出解耦框架,分病变分割和文本-体积推理两阶段,利用解剖学引导解决空间模糊性,在基准测试中取得领先,证明解耦是处理该复杂性的有效范式。

Comments Accepted to MICCAI 2026 (Spotlight Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04935 2026-08-10 cs.CV 版本更新 79%

Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detection

释放视觉-语言模型在通用人工智能生成图像检测中的潜力

Weihan Cai, Hao Tan, Zichang Tan, Jun Wan, Xinping Gao

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 该研究针对AI生成图像检测,发现视觉-语言模型PE比DINOv3更具潜力,提出语义原型校准(SPC)方法得到PE-SPC,在多基准测试中达到新的最先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13621 2026-08-07 cs.CV 版本更新 79%

Visual Intention Grounding for Egocentric Assistants

面向第一人称视角助手的视觉意图定位

Pengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li, Arjun Akula, Angela Yao

机构 * National University of Singapore(新加坡国立大学) Google DeepMind(谷歌DeepMind)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出首个第一人称视觉意图定位数据集EgoIntention,及推理至定位(RoG)指令微调方法,解决多模态模型在第一人称视角下的意图定位问题,实现了统一的视觉定位能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04383 2026-07-30 cs.SD cs.AI 版本更新 79%

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

Auto-AEG:用于开放词汇音频事件定位的可扩展数据构建

Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang, Dongjie Fu, Boyun Zhang, Yongbo He, Tao Jin

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 研究开放词汇音频事件定位任务,因数据稀缺受限。提出Auto-AEG可扩展管道,通过自动数据构建和模型微调构建监督,结合合成音频与伪标签训练,提升模型性能。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24118 2026-07-28 cs.CV 版本更新 79%

An LMM for Precisely Grounding Elements in Documents

一种用于精确文档元素定位的大型多模态模型

Yijian Lu, Chuangxin Zhao, Kai Sun, Lei Hou, Ji Qi, Juanzi Li

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 提出PreciseDoc,一种通过合成数据与强化学习联合训练实现文档元素精确定位的大型多模态模型,显著提升文档理解与视觉定位精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28371 2026-07-28 cs.CY cs.AI 版本更新 79%

Coherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure

无需基础的协调,协调的无需成功:可观测性与认知失败

Camilo Chacón Sartori

机构 * Institut Català de Nanociència i Nanotecnologia (ICN2)(加泰罗尼亚纳米科学与纳米技术研究所(ICN2)) Artificial Intelligence Research Institute (IIIA-CSIC)(人工智能研究所(IIIA-CSIC))

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本文探讨了大型语言模型在低可观测性和高可观测性领域中协调与基础的分离现象,提出认知三角模型以评估人工智能代理的认知能力。

Comments Error found in some results. I need to correct it before re-uploading

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10230 2026-07-24 cs.LG cs.SD eess.AS 版本更新 79%

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

帧级内部工具使用用于音频语言模型中的时序定位

Joseph An, Phillip Keung, Jiaqi Wang, Orevaoghene Ahia, Noah A. Smith

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

AI总结 本文提出帧级内部工具使用方法,通过二元分类器和IHP损失提升音频语言模型的时序定位能力,实现50倍加速和鲁棒长度泛化。

Comments To appear in COLM 2026. Refer to https://github.com/inkitori/taudio/ for the codebase

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21744 2026-07-22 cs.SE cs.AI q-bio.BM 版本更新 79%

Agentic AI-assisted coding offers a unique opportunity to instill epistemic grounding during software development

智能AI辅助编码为软件开发过程中注入认知基础提供了独特契机

Magnus Palmblad, Jared M. Ragland, Benjamin A. Neely

机构 * Center for Proteomics and Metabolomics, Leiden University Medical Center(蛋白质组学与代谢组学中心,莱顿大学医学中心) National Institute of Standards and Technology - Charleston(国家标准与技术研究院-查尔斯顿)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 研究以蛋白质组学为例,提出社区管理的领域范围认知基础文档用于智能AI辅助编码。核心方法是文档编码硬约束和约定参数。主要贡献是助力非领域专家生成优质代码等,增强各方信心并让领域专家参与定制软件开发。

Comments Letter, 12 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13586 2026-07-21 cs.CV 版本更新 79%

UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets

UniPhysGen:用于可模拟3D资产的统一物理基础

Xian Li, Rong Wei, Lujie Yang, Haolin Huang, Junyuan Fang, Siliang Tang, Jun Xiao, Rui Tang, Juncheng Li

机构 * Zhejiang University(浙江大学) Manycore Tech Inc.(众核科技公司) University of Electronic Science and Technology of China(电子科技大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 研究针对现有3D资产缺乏统一物理语义问题,提出UniPhys框架及UniPhysGen模型,通过联合推理关节语义和固有物理属性,减轻几何捷径偏差,经实验验证其性能先进,生成资产可用于机器人模拟环境。

Comments 39 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09760 2026-07-17 cs.CV cs.RO eess.IV 版本更新 79%

PanoAffordanceNet: Towards Holistic Affordance Grounding in 360° Indoor Environments

PanoAffordanceNet:迈向360°室内环境中的整体 affordance 地标

Guoliang Zhu, Wanjun Jia, Caoyang Shao, Yuheng Zhang, Zhiyong Li, Kailun Yang

机构 * School of Artificial Intelligence and Robotics, Hunan University, China(湖南大学人工智能与机器人学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 PanoAffordanceNet通过DASM和OSDH解决360°室内环境中的整体affordance地标问题,整合多级约束抑制语义漂移,构建360-AGD数据集,提升具身智能的场景感知能力。

Comments Accepted to IEEE/RSJ IROS 2026. The source code and benchmark dataset will be made publicly available at https://github.com/GL-ZHU925/PanoAffordanceNet

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22830 2026-07-16 cs.CV cs.RO 版本更新 79%

A Comparative Evaluation of Large Vision-Language Models for 2D Object Detection under SOTIF Conditions

大型视觉-语言模型在SOTIF条件下2D目标检测的比较评估

Ji Zhou, Yilin Ding, Yongqi Zhao, Jiachen Xu, Dong Bi, Johannes Betz, Arno Eichberger

机构 * Institute of Automotive Engineering, Graz University of Technology(汽车工程研究所,格拉茨技术大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 本文评估了大型视觉-语言模型在SOTIF条件下2D目标检测的性能,发现其在复杂场景中召回率显著优于传统方法,但几何精度仍有一定优势。

Comments 8 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27243 2026-07-07 cs.CV 版本更新 79%

Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models

检索头能看见图像吗?长上下文视觉语言模型中的多模态检索头

Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao, Yiming Du, Haobo Li, Xiyu Ren, Ginny Wong, Simon See, Lishu Luo, Haodong Duan, Pasquale Minervini, Yangqiu Song

机构 * HKUST(香港科技大学) University of Edinburgh(爱丁堡大学) CUHK(香港中文大学) NVAITC, NVIDIA, Santa Clara, USA(NVIDIA Santa Clara 分公司) Tsinghua University(清华大学)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV

AI总结 本文提出一种多模态检索头检测方法,发现视觉语言模型中仅有4.4-10.2%的注意力头贡献了50%的正检索分数,这些头对长上下文推理至关重要,且可直接用于文档检索提升性能。

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30169 2026-07-07 cs.CY cs.AI cs.MA 版本更新 79%

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms

分离性身份:语言模型代理缺乏声誉机制的基础

Botao Amber Hu, Helena Rong, Max Van Kleek

机构 * University of Oxford(牛津大学) New York University Shanghai(纽约大学上海分校)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本文指出语言模型代理因本体上的分离性(模块可替换、身份流动)而无法满足声誉机制所需的身份持续性、行为可预测性和制裁敏感性,从而提出转向基于可观察性、事前、构成性、协议的行为约束。

Comments Accepted by FaccT 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16809 2026-07-07 cs.RO cs.AI 版本更新 79%

CABTO: Context-Aware Behavior Tree Grounding for Robot Manipulation

CABTO:面向机器人操作的上下文感知行为树接地

Yishuai Cai, Xinglin Chen, Yunxin Mao, Kun Hu, Yaodong Yang, Yuanpei Chen, Wenjing Yang, Ji Wang, Minglong Li

机构 * National University of Defense Technology(国防科技大学) PsiBot Peking University(北京大学) PKU-Psibot Lab(北京大学-PsiBot实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本文提出CABTO框架,通过预训练大模型和上下文反馈,自动构建完整且一致的行为树系统,解决机器人操作中行为树接地的复杂性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12009 2026-07-02 cs.CV 版本更新 79%

Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale

Affogato: 基于大规模自动数据生成的开词汇可操作区域定位

Junha Lee, Eunha Park, Chunghyun Park, Dahyun Kang, Minsu Cho

机构 * Pohang University of Science and Engineering (POSTECH)(浦项科技大学) SqueezeBits RLWLRD

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 提出Affogato框架,通过自动化流程生成750K 3D可操作区域热力图与自然语言查询对,解决数据瓶颈,实现开词汇可操作区域定位,并验证其跨架构迁移能力。

Comments ECCV 2026, Project page: https://junha-l.github.io/affogato/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26567 2026-07-01 cs.CV 版本更新 79%

AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision

AirZoo: 一种用于地面几何3D视觉的统一大规模数据集

Xiaoya Cheng, Rouwan Wu, Xinyi Liu, Zeyu Cui, Yan Liu, Na Zhao, Yu Liu, Maojun Zhang, Shen Yan

机构 * National University of Defense Technology(国防科技大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 本文提出AirZoo数据集,通过可扩展生成流程、全面场景多样性和丰富的几何标注,解决无人机感知中复杂视角变换和环境条件的挑战,验证其在三维重建和图像检索中的有效性。

Comments ECCV 2026. Project page: https://nudt-sawlab.github.io/AirZoo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24561 2026-07-01 cs.CV 版本更新 79%

RGBT-GroundBench: Visual Grounding Beyond RGB in Complex Real-World Scenarios

RGBT-GroundBench: 超越RGB的复杂真实世界场景中的视觉定位

Tianyi Zhao, Jiawen Xi, Linhui Xiao, Junnan Li, Xue Yang, Maoxun Yuan, Xingxing Wei

机构 * Institute of Artificial Intelligence, State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学人工智能研究院、虚拟现实技术与系统国家重点实验室) Pengcheng Laboratory(鹏城实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 提出首个大规模RGB-热红外视觉定位基准RGBT-GroundBench,包含4万+图像和3级细粒度标注,统一评估协议支持RGB/热红外/融合输入,揭示低照度下性能显著下降,并引入参考基线RGBT-VGNet。

Comments 40pages, 9figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05663 2026-06-29 cs.CV 版本更新 79%

Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding

保持证据链:面向视频时序定位的无训练令牌剪枝的语义证据分配

Jiaqi Li, Shuntian Zheng, Yixian Shen, Jia-Hong Huang, Xiaoman Lu, Minzhe Ni, Yu Guan

机构 * University of Warwick(沃里克大学) University of Amsterdam(阿姆斯特丹大学) Amazon AGI(亚马逊人工智能研究院)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 提出SemVID无训练剪枝框架,通过证据保留和连接强度原则选择对象、运动和上下文令牌,在仅保留12.5%视觉令牌时保持95.4% mIoU,实现5.8倍预填充加速。

Comments Project at https://jiaqili404.github.io/SemVID

Journal ref The 19th European Conference on Computer Vision (ECCV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01085 2026-06-25 cs.CV cs.CL 版本更新 79%

Generalised Medical Phrase Grounding

通用医学短语定位

Wenjun Zhang, Shekhar S. Chandra, Aaron Nicolson

机构 * The University of Queensland(昆士兰大学) Australian e-Health Research Centre, CSIRO Health and Biosecurity(澳大利亚电子健康研究中心,CSIRO健康与生物安全)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 针对现有医学短语定位方法无法处理多区域、非诊断性及不可定位短语的问题,提出通用医学短语定位任务及MedGrounder模型,采用两阶段训练策略,在零样本迁移和多区域/不可定位短语上优于基线方法。

Comments Accepted by IEEE Transactions on Medical Imaging

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03579 2026-06-23 cs.RO cs.LG 版本更新 79%

Intent-Handover: Grounding Language in Human-Usage Regions for Trustworthy Robot-to-Human Handovers

Intent-Handover:在人类使用区域中接地语言以实现可信的机器人到人交接

Hanxin Zhang, Abdulqader Dhafer, Hongbiao Dong, Zhou Daniel Hao

机构 * DANiLab, University of Leicester(莱斯特大学DANiLab) University of Leicester(莱斯特大学) School of Computing and Mathematical Sciences, University of Leicester(莱斯特大学计算与数学科学学院) School of Metallurgy and Materials, University of Birmingham(伯明翰大学冶金与材料学院)

专题命中 视觉定位与Grounding :grounding(title);vision-language model(abstract);分类 cs.LG

AI总结 提出Intent-Handover方法,通过视觉语言模型识别目标物体和人类使用区域,结合抓取优化和人体姿态跟踪,实现安全、可信的机器人到人物体交接。

Comments Accepted at the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07415 2026-06-10 cs.CV cs.CL 版本更新 79%

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring

ChartREG++:面向多样化指代线索和多目标指代的图表指代表达式定位基准与改进

Tianhao Niu, Ziyu Han, Xuan Dong, Qingfu Zhu, Wanxiang Che

机构 * Research Center for Social Computing and Interactive Robotics(社会计算与交互机器人研究中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 针对现有图表指代表达式定位基准的局限,提出支持多种定位形式、多目标指代、多样化线索和图表类型的基准,并利用代码驱动合成流水线生成像素级实例掩码,训练实例分割模型集成到多模态定位框架,显著提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00898 2026-08-11 cs.CL cs.DL 版本更新 79%

Citation Grounding Measures the Oracle: Graph Coverage Determines Reported LLM Hallucination Rates in Law

引用溯源:通过法律引用图检测和减少LLM引用幻觉

Volodymyr Ovcharov

机构 * LEX AI LLC

专题命中 视觉定位与Grounding :grounding(title,abstract)

AI总结 提出引用溯源(CG)指标,利用乌克兰法院判决的引用图(1.008亿判决,5.02亿边)检测LLM法律引用幻觉,并通过CG-DPO方法(基于真实判决构建偏好对)减少幻觉,在100个法律查询上CG为0.791-0.873,幻觉率13-21%。

Comments 21 pages, 4 figures, 5 tables. Substantially revised: title, framing and several v1 results changed. Adds a coverage sweep and a separability analysis; corrects the DPO configuration, the density-accuracy correlation and the qualitative examples. Code and data: https://huggingface.co/datasets/overthelex/citation-grounding-eval

详情

展开后加载摘要…

URL PDF HTML 收藏