arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 2506 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 736 篇

2606.28273 2026-06-29 cs.CL 新提交 89%

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models

视觉默认,先验覆盖:视觉-语言模型中感知-知识冲突的因果机制

Niclas Lietzow, Danielle Bitterman, Carsten Eickhoff, William Rudman, Michal Golovanevsky

机构 * University of Tübingen(图宾根大学) Harvard University(哈佛大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title,abstract);grounding(abstract)

AI总结 通过激活修补和消融实验,发现VLM中视觉默认激活,而先验知识依赖少量因果注意力头(2.5-4.8%),形成不对称因果结构。

Comments 14 pages, 11 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12781 2026-08-14 cs.CV 新提交 89%

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

超越正确性:混合思维多模态大语言模型(MLLM)的响应行为基准测试与对齐

Xinming Wang, Weinong Wang, Hongming Yang, Yansong Lin, Zheng Ruan, Shangpin Peng, Qiming Peng, Nan Qiao, Fengyuan Lu, Guoqing Ma, Marito Li, Songyang Zhang, Saiyong Yang, Han Hu, Yonglong Tian, Xu-Yao Zhang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Large Language Model Department, Tencent(腾讯大语言模型部) University of Electronic Science and Technology of China(电子科技大学) Hong Kong University of Science and Technology(香港科技大学) Zhongguancun Academy(中关村学院)

专题命中 视觉定位与Grounding :MLLM(title_cn,summary_cn);grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 该研究针对混合思维 MLLM 的思维与非思维模式响应错位问题,构建 PatternEval 基准并开发 PatternRL 方法,可减轻跨模式错位且任务性能损失极小。

Comments 8 tables and 6figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03631 2026-08-05 cs.CV 新提交 89%

SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification

SEER:用于受控空间关系分类的自 grounding 证据接口

Feixiang Liu, Likun Wang, Qiang Qiu, Hui Xu, Huawei Shen, Xueqi Cheng

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);grounding(title_cn,abstract);分类 cs.CV

AI总结 本文提出针对冻结VLM的无训练推理时证据接口SEER,通过构建查询特定证据缓解空间关系分类错误,在多数据集上取得显著性能提升,证明该干预措施的有效性。

Comments 23 pages total, 2 figures. Code: https://github.com/SouthWinter/SEER

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15517 2026-07-20 cs.CV cs.AI 新提交 89%

SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification

SLAPBench:用于四指SLAP指纹验证的多模态大语言模型基准测试

Bibesh Pyakurel, M. G. Sarwar Murshed

机构 * University of Wisconsin–Green Bay(威斯康星大学格林湾分校)

专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);multimodal large language model(title,abstract);分类 cs.CV、cs.AI

AI总结 研究四指SLAP指纹验证,介绍SLAPBench基准,评估多个MLLM在不同提示下的表现,发现提示控制崩溃,模型能力控制歧视,建立了特定于SLAP的MLLM基线,揭示了模型在指纹验证中的能力差距和公平性问题。

Comments 19 pages, 6 figures, 2 tables. Includes appendix with supporting figures and per-subgroup fairness detail. Code and data: https://github.com/bibeshpyakurel/SLAPBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09691 2026-08-11 cs.CV 新提交 89%

Diffuse the object, keep its label: curating detector training data from a few unlabeled photographs via VLM-built 3D vegetation scenes

扩散目标,保留其标签:通过VLM构建的3D植被场景从少量未标记照片中整理检测器训练数据

Mario Malizia, Marnix Enting, Rob Haelterman, Ken Hasselmann

机构 * Royal Military Academy(皇家军事学院) KU Leuven(鲁汶大学) Flanders Make(佛兰德制造研究院)

专题命中 视觉定位与Grounding :VLM(title,title_cn);分类 cs.CV

AI总结 本研究针对植被中小物体标记图像稀缺导致检测器跨站点泛化差的问题,通过VLM构建3D植被场景合成标记训练图像,实现无监督站点适应,在排雷基准上性能优于传统跨站点标签复用。

Comments Accepted at the Curated Data for Efficient Learning (CDEL) Workshop @ ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01258 2026-08-04 cs.CV 新提交 89%

A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2 & Nano Banana2

多模态大语言模型生成图像检测的基准数据集:GPT Image2与Nano Banana2

Zirui Zhang, Yinbo Yu, Donghai Guan, Chunwei Tian, Daoqiang Zhang, Qi Zhu

机构 * College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)

专题命中 视觉定位与Grounding :MLLM(title,summary_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 本文构建了含三种生成协议的MLLM生成图像检测基准数据集,评估现有检测器性能并提出SAP-DSP基线框架,验证了其检测稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03595 2026-07-07 cs.CV cs.AI cs.RO 新提交 88%

Token-Based Affordance Grounding with Large Vision-Language Models

基于令牌的大型视觉语言模型的可供性基础

Seung Il Lee, Qinqian Lei, Daguang Xu, Dong Yang, Robby T. Tan, Yixin Chen, Bo Wang

机构 * University of Mississippi(密西西比大学) National University of Singapore(新加坡国立大学) NVIDIA(英伟达)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);grounding(title,abstract);分类 cs.CV、cs.AI

AI总结 研究旨在解决现有视觉语言模型在行动定位上的问题,提出TokAG零样本可供性基础框架,利用令牌级语义空间信号定位相关区域,引入空间感知令牌选择机制,在多基准测试中优于先前方法。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00020 2026-07-02 cs.RO 新提交 88%

EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories

EmbodimentSemantic:面向具身操作轨迹的空间场景图数据集与视觉语言模型基准

Hassan Jaber, Refinath S N, Luca Cagliero, Christopher E. Mower, Haitham Bou-Ammar

机构 * Politecnico di Torino(都灵理工大学) University College London(伦敦大学学院)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(title);grounding(abstract)

AI总结 提出EmbodimentSemantic数据集与基准,通过场景图三元组评估VLM在具身操作中的空间关系理解,发现现有模型在深度感知和视角依赖关系上存在不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03647 2026-07-07 cs.CV 新提交 88%

Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs

医学视觉语言模型真的能“看”吗?用于视觉依赖型医学视觉语言模型的反事实基础框架和硬负对比训练

Anas Zafar, Leema Krishna Murali, Siddhant Bharadwaj, Ashish Vashist, Jia Wu

机构 * The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心) Eisai Inc.(卫材株式会社) IISc, Bangalore(印度科学研究所班加罗尔分校) Cohere Labs Community(Cohere实验室社区)

专题命中 视觉定位与Grounding :vision language model(title,abstract);grounding(title,abstract);分类 cs.CV

AI总结 探讨医学视觉语言模型是依据视觉证据推理还是利用文本捷径,引入反事实评估框架和对比检索增强学习方法,提升模型视觉依赖能力并揭示跨域诊断差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02470 2026-08-04 cs.CV cs.AI 新提交 88%

Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment

基于专用分割模型实现具身智能视觉语言模型(Agentic VLMs)的细粒度车辆损伤评估

Vishwajeet Shivaji Hogale, Anjali Pai, Nitya Ravi

机构 * Northeastern University(东北大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 针对VLMs空间定位不可靠问题,本文提出TinyDamage架构,将空间定位委托给专用多任务分割模型,集成至7节点LangGraph智能体流程,在车辆损伤评估中大幅降低报告虚构率。

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31148 2026-07-01 cs.CV cs.AI cs.CL 新提交 88%

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding

PruneGround: 用于3D视觉定位的即插即用空间剪枝框架

Duc Cao Dinh, Khai Le-Duc, Florent Draye, Chris Ngo, Terry Jingchen Zhang, Bernhard Schölkopf, Zhijing Jin

机构 * Knovel Engineering Lab(Knovel工程实验室) University of Toronto(多伦多大学) Vector Institute(向量研究所) ELLIS Institute(ELLIS研究所) Max Planck Institute(马克斯·普朗克研究所)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision language model(abstract);分类 cs.CV、cs.AI

AI总结 提出PruneGround框架,通过语言引导的空间剪枝、多视角描述重构和LLM定位器,在剪枝区域高效定位目标,在多个基准上取得最优结果。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21595 2026-07-24 cs.CV cs.AI cs.LG 新提交 88%

3D-Aware VLMs with Implicit and Explicit Geometries

具有隐式和显式几何的3D感知视觉语言模型

Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao, Ran Xu, Shijian Lu, Gongjie Zhang

机构 * Nanyang Technological University(南洋理工大学) DAMO Academy, Alibaba Group(阿里巴巴达摩院) HuPan Lab(湖畔实验室)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 研究针对多数2D视觉输入的VLM处理3D任务困难的问题,提出VLM-IE3D框架,通过引入隐式和显式几何令牌及3D感知适配器,融合几何表示与视觉线索,在多种3D任务中表现优异。

Comments Accepted by ECCV 2026, Open Sourced

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09789 2026-08-11 cs.CV 新提交 87%

ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection

ADOPD:面向多模态大语言模型(MLLM)的工业异常检测的参考特权在线策略蒸馏

Jingtai He, Shiyuan Meng, Wenchao Meng, Qinmin Yang

机构 * Zhejiang University(浙江大学)

专题命中 视觉定位与Grounding :MLLM(title,title_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 本研究针对工业异常检测问题,提出ADOPD参考特权在线策略蒸馏框架,通过参考感知教师监督仅查询的学生模型,在MMAD基准零样本推理下获77.31%平均准确率,优于Qwen3-VL-4B主干及其一样本设置。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07427 2026-08-10 cs.AI cs.PF 新提交 87%

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy

一图胜千词:视觉语言模型如何在提升准确率的同时降低AI能源成本

Bhavika Jalli, Nikhil Korati Prasanna, Jayanta Choudhury

专题命中 视觉定位与Grounding :vision language model(title);VLM(abstract,abstract_cn);vision-language model(abstract);grounding(abstract)

AI总结 该研究提出将时间序列编码为二维图的视觉语言模型,可大幅减少输入词元、降低推理能耗,同时在电信异常检测等任务中提升准确率,解决了LLM处理多维度KPI数据的低效问题。

Comments Accepted at the 14th European Conference on Renewable Energy Systems (ECRES), July 7--9, 2026, London, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04533 2026-08-06 cs.CV 新提交 87%

EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation

EgoAfford:基于自我中心指称分割的面向任务的可供性定位

Xinyuan Guan, Feifan Chen, Xinyu Zhan, Fu-Cheng Zhang, Cewu Lu, Lixin Yang

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 针对部件级可供性定位难以适配复杂多步桌面任务的问题,提出基准EgoAfford及参考模型EgoLens,验证了下一步推理与动作角色条件部件定位的互补挑战,为感知与规划联合研究提供基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04385 2026-08-06 cs.CV 新提交 87%

ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination

ReGround:通过自诊断与视觉重检验恢复多步推理中的视觉接地

Lei Peng, Shuai Lv, Wei Hu

机构 * University of Science and Technology of China(中国科学技术大学) School of Artificial Intelligence and Data Science(人工智能与数据科学学院) State Key Laboratory of Precision and Intelligent Chemistry(精准与智能化学国家重点实验室)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 ReGround是无需架构修改或外部工具的两阶段框架,通过自诊断与视觉重检验解决VLMs多步推理中的视觉接地丢失问题,在八个基准上获一致增益且推理开销适度。

Comments Accepted to ACM Multimedia 2026 (MM '26). 8 pages main text, 4 figures, plus appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00877 2026-08-04 cs.LG 新提交 87%

GeoArbiter: Verifiability-Guided Grounding for Remote-Sensing Multimodal LLMs

GeoArbiter:面向遥感多模态大语言模型的可验证性导向 grounding 方法

Xuechen Li

机构 * University of Minnesota, Twin Cities(明尼苏达大学双城分校)

专题命中 视觉定位与Grounding :grounding(title,title_cn);multimodal large language model(abstract);分类 cs.LG

AI总结 该研究针对遥感多模态大语言模型的事实断言问题,提出无需训练的 GeoArbiter 流程,通过仅注入图像无法验证的地理事实,有效降低了模型的断言级幻觉并提升了对冲突记录的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00235 2026-08-04 cs.CV 新提交 87%

Attention-Steered Vision-Language Models for Sign Language Translation

用于手语翻译的注意力引导视觉-语言模型

Meibo Hu, Guohao Sun, Annemarie D. Ross, Sheng Li, Zhiqiang Tao

机构 * Rochester Institute of Technology(罗切斯特理工学院) University of Virginia(弗吉尼亚大学) National Technical Institute for the Deaf(国家聋人技术学院)

专题命中 视觉定位与Grounding :vision-language model(title,abstract);VLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV

AI总结 针对视觉-语言模型在手语翻译中时空视觉定位差的问题,提出AttnSign框架,通过空间注意力监督和RL运动节奏引导提升性能,在How2Sign和OpenASL基准上表现优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17884 2026-07-21 cs.AI 新提交 87%

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding

ST-Veto:通过泰勒预测和视觉基础实现用于扩散多模态语言模型的时空令牌否决

Keuntae Kim, Beomseok Lee, Hyunwoo Kim, Yong Suk Choi

专题命中 视觉定位与Grounding :grounding(title);VLM(abstract,abstract_cn);vision language model(abstract);multimodal large language model(abstract)

AI总结 研究针对扩散多模态大语言模型推理不足问题,提出无需训练的ST-Veto方法,利用二阶泰勒预测和图像注意力质量否决不稳定及弱基础令牌,与更安全候选交换,在多基准测试中优于其他方法,提升准确率且无额外成本。

Comments ICML 2026 - main

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05222 2026-07-07 cs.CV 新提交 87%

A Multimodal Reasoning Typology for Grounding Chart-Image Coherence in Science Communication

面向科学传播中图表-图像连贯性锚定的多模态推理类型学

Avina Nakarmi, Sohom Sen, Xun Song, Sreyashi Samaddar, Aritra Dasgupta

机构 * New Jersey Institute of Technology(新泽西理工学院) Brooklyn College(布鲁克林学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 本研究提出R1-R5多模态推理类型学,刻画科学文献中图表、图像与文本的协同推理缺口,可系统识别连贯性锚定状态,缩小不同人群解读科学结论的认知差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01050 2026-07-07 cs.CV 新提交 87%

GeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process Supervision

GeoSearcher: 基于锚点引导的渐进推理遥感视觉定位与过程监督

Dianyu Wang, Peirong Zhang, Xuyang Li, Xiaoxuan Liu, Lei Wang

机构 * Key Laboratory of Target Cognition and Application Technology (TCAT), Chinese Academy of Sciences(中国科学院目标认知与应用技术重点实验室) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 提出GeoSearcher,通过锚点引导的渐进推理和过程监督,将遥感视觉定位转化为两阶段过程,解决小目标定位和复杂查询的挑战。

Comments 14 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02020 2026-07-03 cs.AI 新提交 87%

Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails

持续多模态学习中的隐藏遗忘:当准确率幸存但基础失效时

Qianyu Chen, Canran Xiao, Runxuan Tang

机构 * Nanyang Technological University(南洋理工大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)

专题命中 视觉定位与Grounding :grounding(title,abstract);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI

AI总结 针对持续多模态学习中模型答案准确但证据使用方式改变的问题,提出无重放依赖约束框架RCL,通过冻结旧模型、反事实干预估计证据依赖并联合优化,有效降低隐藏遗忘。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27084 2026-06-26 cs.CV eess.IV 新提交 87%

Pseudo-Text-Conditioned 3D Grounding DINO for Organ Localization in Abdominal CT

伪文本条件3D Grounding DINO用于腹部CT器官定位

Siqi Chen, Han Gong, Keyi Hou, Jingxuan Yang, Sheethal Bhat, Andreas Maier

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(埃尔朗根-纽伦堡大学)

专题命中 视觉定位与Grounding :grounding(title,title_cn);分类 cs.CV

AI总结 提出CT-3GDINO,一种轻量级3D检测器,通过冻结伪文本类令牌替代真实文本编码器,实现腹部CT中五个器官的定位,在193个体积上达到0.5830 mAP。

Comments 24 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20946 2026-06-23 cs.CL cs.CV 新提交 87%

Scaling Diverse Language Generation for 3D Visual Grounding

面向3D视觉定位的多样化语言生成扩展

Austin T. Wang, Dongchen Yang, Angel X. Chang

机构 * Simon Fraser University(西蒙菲莎大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(summary_cn,abstract_cn);分类 cs.CV

AI总结 提出ViGiL3D++方法,通过场景图约束采样与LLM语言生成结合,生成多样化视觉定位查询,提升3DVG模型泛化能力并揭示VLM局限性。

Comments 39 pages, 14 figures, 16 tables. Project Page: https://3dlg-hcvc.github.io/vigil3dpp

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16124 2026-06-16 cs.CV 新提交 87%

Training-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos

面向遥感图像与视频的无训练开放词汇视觉定位

Ke Li, Di Wang, Yongshan Zhu, Ting Wang, Weiping Ni, Tao Lei, Quan Wang, Xinbo Gao

机构 * School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院) Interdisciplinary Institute of Artificial Intelligence, Xidian University(西安电子科技大学跨学科人工智能研究院) School of Artificial Intelligence, Xidian University(西安电子科技大学人工智能学院) Northwest Institute of Nuclear Technology(西北核技术研究所) School of Physics and Information Engineering, Fuzhou University(福州大学物理与信息工程学院)

专题命中 视觉定位与Grounding :grounding(title,abstract);VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 提出无训练框架RSVG-ZeroOV,利用冻结的通用基础模型通过概览-聚焦-演化范式实现零样本开放词汇遥感视觉定位,并扩展至视频时空定位,在多个基准上超越现有零样本方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16122 2026-06-16 cs.AI 新提交 87%

Thinking with Visual Grounding

视觉锚定思维

Junkai Zhang, Yihe Deng, Kai-Wei Chang, Wei Wang

机构 * University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

专题命中 视觉定位与Grounding :grounding(title,abstract);vision-language model(abstract);VLM(abstract_cn);visual reasoning(abstract)

AI总结 提出视觉锚定思维方法,让视觉语言模型在推理时交替生成自然语言和视觉锚点(点或框),并通过合成数据管道和锚定感知强化学习训练,在计数和空间推理任务上显著提升性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24730 2026-07-28 cs.CV cs.AI 新提交 87%

KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability

KANEx:将柯尔莫哥洛夫-阿诺德网络的可解释性转化为医学可解释性

Krithi Shailya, Ananya Lakshmi Ravi, Venkatanathan K. V., Sowmya S. Sundaram, Gokul S. Krishnan, Aditi Anand, Balaraman Ravindran

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯分校) Vanderbilt University School of Medicine(范德堡大学医学院)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 研究针对医学应用中视觉模型黑箱问题,提出KANEx框架,利用柯尔莫哥洛夫-阿诺德网络的符号透明度为VLM推理奠基,设计KAN-Map热图生成方法,经实验验证该方法能提升语义相似度、视觉定位及推理质量,为可信医学人工智能发展助力。

Comments MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21579 2026-06-23 cs.CV cs.AI 新提交 87%

The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection

VLM在零样本程序性错误检测中的惊人有效性

Serdar Ozsoy, Lars Doorenbos, Federico Spurio, Gianpiero Francesca, Juergen Gall

机构 * University of Bonn(波恩大学) Toyota Motor Europe(丰田欧洲公司) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习和人工智能研究所)

专题命中 视觉定位与Grounding :VLM(title_cn,summary_cn);分类 cs.CV、cs.AI

AI总结 提出ZeProM框架,利用单个预训练VLM同时解决零样本程序性错误检测和时间动作分割,在EgoPER和CaptainCook4D基准上接近或超越全监督方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07117 2026-08-10 cs.CV 新提交 86%

Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report Generation

超越流畅性:脊柱MRI报告生成的临床基准与异常增强基线

Bruno Palau, Franziska Vogt, Daria Laslo, Haobo Li, Ender Konukoglu, Maria Monzon, Catherine R. Jutzeler

机构 * ETH Zurich(苏黎世联邦理工学院) Swiss Institute of Bioinformatics (SIB)(瑞士生物信息学研究所)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 该研究针对放射学报告的耗时与阅片差异问题,构建腰椎MRI的VLM临床基准,提出半监督U-Net++生成异常热图增强VLM的框架,提升诊断可靠性与可解释性。

Comments Maria Monzon and Catherine R. Jutzeler contributed equally as shared last authors. Accepted at the CV4Clinic Workshop, CVPR 2026

Journal ref Proc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2026, pp. 6759-6770

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00920 2026-07-03 cs.CV 新提交 86%

GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images

GMO-E$^2$DIT:面向电商图像的接地多操作编辑

Zipeng Guo, Xiaoan Liu, Lichen Ma, Cheng Wang, Yu He, Xiaolong Fu, Jingling Fu, Xinyuan Shan, Shaojie Guo, Luohang Liu, Junshi Huang, Yan Li

机构 * Wuhan University(武汉大学) Xi’an Jiaotong University(西安交通大学) Sun Yat-sen University(中山大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);vision-language model(abstract);grounding(abstract);分类 cs.CV

AI总结 提出GMO-E$^2$DIT框架,通过VLM代理构建区域接地编辑议程,结合掩码条件图像编辑器和反思循环,实现多操作、可审计的电商图像编辑,在指令准确性和编辑保真度上超越现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏