arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7348 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7348 篇

2305.18924 2023-08-29 cs.AI cs.LO cs.PL 80%

Bottom-Up Grounding in the Probabilistic Logic Programming System Fusemate

Peter Baumgartner, Elena Tartaglia

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

Comments This is an extended version of the ICLP 2023 paper at ICLP2023:4654" target="_blank" rel="noopener">https://cgi.cse.unsw.edu.au/~eptcs/paper.cgi?ICLP2023:4654. It also includes an improvement to the grounding algorithm in Section 3

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01821 2023-05-30 cs.CV 80%

Toward Explainable and Fine-Grained 3D Grounding through Referring Textual Phrases

Zhihao Yuan, Xu Yan, Zhuo Li, Xuhao Li, Yao Guo, Shuguang Cui, Zhen Li

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments New dataset for 3D visual grounding is available at https://yanx27.github.io/phraserefer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02687 2022-07-07 cs.CV 80%

Team PKU-WICT-MIPL PIC Makeup Temporal Video Grounding Challenge 2022 Technical Report

Minghang Zheng, Dejie Yang, Zhongjie Ye, Ting Lei, Yuxin Peng, Yang Liu

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments 2st Place in PIC Makeup Temporal Video Grounding (MTVG) Challenge in ACM-MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.12346 2021-03-24 cs.CV 80%

Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos

Sijie Song, Xudong Lin, Jiaying Liu, Zongming Guo, Shih-Fu Chang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

Comments Accepted to CVPR2021. The project page is at https://sijiesong.github.io/co-grounding

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23733 2026-08-13 cs.CL 版本更新 80%

Multimodal QUD: Inquisitive Questions from Scientific Figures

多模态QUD:来自科学图表的探究性问题

Yating Wu, William Rudman, Venkata S Govindarajan, Alexandros G. Dimakis, Junyi Jessy Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Ithaca College(伊萨卡学院) UC Berkeley, BespokeLabs.ai(伯克利大学,BespokeLabs.ai)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract_cn);VLM(abstract_cn)

AI总结 本文提出多模态QUD数据集,通过结合图表与文本上下文生成探究性问题,提升多模态推理能力,实现更高质量的视觉 grounding。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20913 2026-06-23 cs.CV cs.AI cs.LG 新提交 80%

PROTON: Prototype-Based Test-Time Online OOD Detection for Medical VLMs

PROTON: 基于原型的测试时在线OOD检测方法用于医学视觉语言模型

Abhijit Das, Nichula Wasalathilaka, Yifan Lu, Adinath Dukre, Dwarikanath Mahapatra, Shadab Khan, Imran Razzak

机构 * MBZUAI(穆罕默德·本·扎耶德人工智能大学) University of Peradeniya(佩拉德尼亚大学) Khalifa University(哈利法大学) ADIA Lab(阿布扎比投资局实验室) MedOS

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 针对医学视觉语言模型在部署时难以检测分布外输入的问题,提出PROTON方法,通过在线原型库和自适应融合原型距离与最大概念匹配得分,无需修改模型或训练数据,在多个OOD场景下提升检测性能。

Journal ref 29th International Conference on Medical Image Computing and Computer Assisted Intervention 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18738 2026-06-18 cs.SD 新提交 80%

GRIDEX: Grid-Grounded Forensic Explanations for Deepfake Spectrogram Analysis

GRIDEX:基于网格的深度伪造频谱图取证解释

Thi Ngan Ha Do, Tingmin Wu, Alsharif Abuadbba, Kristen Moore

机构 * CSIRO(澳大利亚联邦科学与工业研究组织)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);grounding(abstract)

AI总结 提出GRIDEX框架,通过两阶段学习(SFT+GRPO)定位频谱图异常区域并生成结构化取证解释,提升伪造检测的可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07343 2026-06-16 cs.CV cs.AI cs.LG cs.RO 版本更新 80%

Seeing Roads Through Words: A Language-Guided Framework for RGB-T Driving Scene Segmentation

通过文字看道路:一种语言引导的RGB-T驾驶场景分割框架

Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy

机构 * National University of Singapore(新加坡国立大学) University of Technology Sydney(悉尼科技大学)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 提出CLARITY框架,利用视觉语言模型先验动态调整RGB-T融合策略,并引入暗目标语义保留和层次化解码器,在MFNet数据集上达到62.3% mIoU和77.5% mAcc的新SOTA。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13870 2026-06-15 cs.CV cs.AI cs.LG 新提交 80%

Mirage Probes: How Vision Models Fake Visual Understanding

幻象探针:视觉模型如何伪造视觉理解

Daniel Ben-Levi, Judah Goldfeder, Weiliang Zhao, Raz Lapid, Amit LeVi, Allen G. Roush, Ravid Shwartz-Ziv, Hod Lipson

机构 * Columbia University(哥伦比亚大学) Intuit Technion(以色列理工学院) Thoughtworks New York University(纽约大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract_cn);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 提出幻象探针框架,通过对比探针揭示视觉语言模型在无图像时也能回答问题的两种幻象行为:文本偏见和虚假图像,并证明后者需要表征级干预。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07723 2026-06-09 cs.RO 新提交 80%

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

VoLo: 面向开放词汇长时程操控的物理编排器

Siyi Chen, Hugo Hadfield, Alex Zook, Mikaela Angelina Uy, Chan Hee Song, Erwin Coumans, Xuning Yang, Faisal Ladhak, Qing Qu, Stan Birchfield, Jonathan Tremblay, Valts Blukis

机构 * NVIDIA(英伟达) University of Michigan(密歇根大学)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract)

AI总结 提出VoLoAgent,利用VLM将VLA/WAM作为可中断工具进行物理编排,实现开放词汇长时程操控,并在新基准RoboVoLo上显著优于现有系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06061 2026-06-05 cs.RO 80%

A Conversational Framework for Human-Robot Collaborative Manipulation with Distributed Generative AI models

基于分布式生成式AI模型的人机协作操作对话框架

Arash Ghasemzadeh Kakroudi, Roel Pieters

机构 * Automation Technology and Mechanical Engineering, Tampere University(自动化技术与机械工程,塔尔库大学)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);grounding(abstract)

AI总结 提出一个分布式对话框架,集成语言和视觉语言模型与ROS 2执行栈,实现从自由形式用户命令生成结构化操作请求,并通过视觉基础将图像空间目标转换为机器人框架目标,实验验证了端到端任务可靠性和延迟。

Comments Accepted to the 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026). The final published version will appear under the title "A Distributed Conversational Framework for Human-Robot Collaborative Manipulation Using Local LLMs and VLMs"

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00985 2026-06-02 cs.RO 80%

Make Your VLA More Robust Without More Data By Interleaving Motion Planning

通过交错运动规划使您的VLA更鲁棒而无需更多数据

Dan BW Choe, Sundhar Vinodh Sangeetha, Samuel Coogan, Shreyas Kousik

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 视觉定位与Grounding :VLM(summary_cn,abstract)

AI总结 提出MPVI框架,将基于模型的运动规划与视觉-语言-动作模型交错结合,通过VLM完成检查和本体感受触发实现可靠切换,无需额外训练即可提升长时域移动操作任务的鲁棒性,在BEHAVIOR-1K基准上任务进度提升113%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30957 2026-06-01 cs.RO 80%

RDGen: Demonstration Generation for High-Quality Robot Learning via Reinforcement Learning

RDGen: 通过强化学习生成高质量机器人学习的演示

Zijian Zhu, Menglin Zou, Zhuang Li, Yaojie Tu, Xinhai Sun

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract,abstract_cn)

AI总结 提出RDGen框架,利用从仿真到真实的强化学习策略生成高质量机器人演示轨迹,用于训练视觉-语言-动作模型,相比人工遥操作产生更平滑轨迹并提升下游性能。

Comments 13 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04343 2026-05-19 cs.IR 80%

The Personalization Paradox: Semantic Loss vs. Reasoning Gains in Agentic AI Q&A

个性化悖论:语义损失与推理增益在代理AI问答中的权衡

Satyajit Movidi, Stephen Russell

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract)

AI总结 本文研究了个性化对系统性能的影响,通过比较不同配置发现个性化虽提升推理和 grounding 能力,但导致语义相似度下降,揭示了现有 LLM 评估方法的不足。

Journal ref Cloud Computing and Data Science 2026 May 18;7(2):290-313

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10032 2026-05-13 cs.CL 80%

PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning

PlantMarkerBench: 一个多物种证据导向植物标记推理基准

Sajib Acharjee Dip, Song Li, Liqing Zhang

机构 * Department of Computer Science, Virginia Tech(弗吉尼亚理工学院计算机科学系) School of Plant and Environmental Sciences, Virginia Tech(弗吉尼亚理工学院植物与环境科学学院) Fralin Biomedical Research Institute, Virginia Tech(弗吉尼亚理工学院弗拉林生物医学研究学院) FBRI Cancer Research Center, Washington, DC(华盛顿特区FBRI癌症研究中心)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract)

AI总结 本文提出PlantMarkerBench,通过整合文献检索与生物 grounding,构建了包含4种植物的基准,评估文献支持的植物标记证据解释,发现模型在功能、间接和弱支持证据上表现欠佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09802 2026-05-12 cs.CV cs.AI cs.LG 80%

CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection

CrossVL: 用于跨视角视觉语言检测的复杂度感知特征路由与配对课程学习

Zhipeng Liu, Chunbo Luo

机构 * Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出CrossVL框架,结合复杂度感知路径聚合和配对课程学习,提升跨视角视觉语言模型的检测性能,实验显示其在MAVREC数据集上提升了aerial mAP并缩小了地面与空中视角的性能差距。

Comments Accepted to CVPR 2026. Code available at https://github.com/1nyourlife/Crossvl_cvpr2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27600 2026-05-01 cs.IR 80%

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG

净化多模态检索:用于RAG的片段级证据选择

Xihang Wang, Zihan Wang, Chengkai Huang, Cao Liu, Ke Zeng, Quan Z. Sheng, Lina Yao

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);multimodal large language model(abstract)

AI总结 本文提出FES-RAG框架,通过片段级证据选择提升多模态检索效果,减少噪声干扰,实验显示在M2RAG基准上性能提升27%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25914 2026-04-29 cs.CL 80%

DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios

DV-World:在现实场景中评估数据可视化代理的基准测试

Jinxiang Meng, Shaoping Huang, Fangyu Lei, Jingyu Guo, Haoxiang Liu, Jiahao Su, Sihan Wang, Yao Wang, Enrui Wang, Ye Yang, Hongze Chai, Jinming Lv, Anbang Yu, Huangjing Zhang, Yitong Zhang, Yiming Huang, Zeyao Ma, Shizhu He, Jun Zhao, Kang Liu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) National University of Singapore(新加坡国立大学) Renmin University of China(中国人民大学)

专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn);MLLM(abstract,abstract_cn)

AI总结 DV-World通过260个任务评估数据可视化代理在现实专业生命周期中的能力,涵盖表格操作、视觉进化和意图对齐,采用混合评估框架揭示现有模型在复杂数据可视化挑战中的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25323 2026-04-29 cs.RO 80%

ANCHOR: A Physically Grounded Closed-Loop Framework for Robust Home-Service Mobile Manipulation

ANCHOR:一种基于物理的闭环框架,用于鲁棒的家庭服务移动操作

Jinhao Jiang, Shengyu Fang, Sibo Zuo, Yujie Tang, Yirui Li

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract)

AI总结 ANCHOR通过物理 grounding 和结构化故障处理,提升家庭服务机器人在动态环境中的任务成功率和扰动恢复能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12387 2026-04-15 q-bio.GN 80%

oxo-call: Documentation-grounded Skill Augmentation for Accurate Bioinformatics Command-line Generation with Large Language Models

oxo-call:基于文档的技能增强用于准确的生物信息学命令行生成与大型语言模型

Yun Peng, Yujun Sun, Jia Ding, Bin Yan, Zhangyu Wang, Chunyang Wang, Chenyang Shu, Jian-Guo Zhou, Shixiang Wang

专题命中 视觉定位与Grounding :grounding(summary_cn,abstract)

AI总结 oxo-call通过文档优先 grounding 和精选技能增强策略,提升生物信息学命令行生成的准确性,提供150+内置技能和可扩展的流程引擎。

Comments 19 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17651 2025-10-21 cs.CV cs.AI cs.LG 80%

Frugal Federated Learning for Violence Detection: A Comparison of LoRA-Tuned VLMs and Personalized CNNs

Sébastien Thuau, Siba Haidar, Ayush Bajracharya, Rachid Chelouah

机构 * esieaLab(esiea实验室) ESIEA(ESIEA学院) ETIS Laboratory(ETIS实验室) CNRS(法国国家科学研究中心) UMR8051(UMR8051研究中心) University of CY Cergy(CY塞克大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 7 pages, 1 figure, FLTA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11616 2025-08-18 cs.CV cs.AI cs.CL cs.LG 80%

Controlling Multimodal LLMs via Reward-guided Decoding

Oscar Mañas, Pierluca D'Oro, Koustuv Sinha, Adriana Romero-Soriano, Michal Drozdzal, Aishwarya Agrawal

机构 * Mila - Quebec AI Institute(魁北克AI研究院) Université de Montréal(蒙特利尔大学) McGill University(麦吉尔大学) Meta FAIR Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Published at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23573 2025-04-01 cs.CV cs.AI cs.LG 80%

DASH: Detection and Assessment of Systematic Hallucinations of VLMs

Maximilian Augustin, Yannic Neuhaus, Matthias Hein

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03151 2025-01-07 cs.AI cs.CV cs.LG 80%

Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches

Alhassan Mumuni, Fuseini Mumuni

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15419 2026-08-18 cs.CV 新提交 79%

ArtLang: Structured Language-to-Kinematics Grounding for Articulated 3D Actuation

ArtLang:面向铰接3D驱动的结构化语言到运动学的绑定

Sylvia Yuan, Dan Wang, Ravi Ramamoorthi, Xinrui Cui

机构 * University of California San Diego(加利福尼亚大学圣迭戈分校) University of North Texas(北得克萨斯大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 研究针对铰接物体控制的语义匿名问题,提出ArtLang框架,通过语义-运动学铰接图等技术实现开放词汇语言控制,经多类实验验证了可靠的语言绑定与连续铰接控制能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15382 2026-08-18 cs.AI cs.IR 新提交 79%

Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot

将医疗大语言模型(LLM)基于因果知识图谱:框架、指标与心血管试点研究

Ummara Mumtaz, Aimen Noor, Awais Ahmed

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.AI

AI总结 本研究提出以因果知识图谱为核心的医疗LLM评估框架,在心血管试点中验证其有效性,发现集成条件C4在因果推理相关指标上表现最优,未基于图的C1原始干预准确性最高但缺乏因果与证据基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14228 2026-08-17 cs.LG 新提交 79%

AutoSchema: Live Schema Grounding for Agentic Text-to-Sparql over Heterogeneous Knowledge Graphs

AutoSchema:面向异构知识图谱的智能体文本转SPARQL的实时模式接地

Yiming Zhang, Koji Tsuda

机构 * The University of Tokyo(东京大学) National Institute for Materials Science(国立材料科学研究所) RIKEN Center for Advanced Intelligence Project(理化学研究所高级智能项目中心)

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.LG

AI总结 提出无需训练的AutoSchema框架,用于异构知识图谱的智能体文本转SPARQL的实时模式接地,在多项生物医学KGQA等任务中优于TogoMCP,可支持未记录RDF图谱。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12748 2026-08-14 cs.CV 新提交 79%

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

扩展表示多样性:用于视觉定位的调制注意力与重构正则化

Junyi Hu, Tian Bai, Fengyi Wu, Yian Huang, Wei Wen, Zaoli Li, Junli Lin, Xingchen Li, Zhenming Peng, Yi Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 针对指代表达理解模型跨数据集泛化能力有限的问题,提出含mACH与JEPA辅助流的架构及Objects365-Caption数据,实现了强泛化与具竞争力的REC性能。

Comments 21 pages, 10 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12746 2026-08-14 cs.CV cs.CL 新提交 79%

Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors

双流跨锚校正:接地长文本描述与对象级锚点的域限制

LingKai Bu

专题命中 视觉定位与Grounding :grounding(title);multimodal large language model(abstract);分类 cs.CV

AI总结 针对多模态大语言模型的长文本描述对象幻觉问题,本文提出双流跨锚校正方法,通过耦合感知流与认知流提升精度,在长文本场景下实现最优性能,且存在域条件性限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12683 2026-08-14 cs.RO cs.CV 新提交 79%

FUSE: Active Functional Affordance Grounding through Adaptive Semantic-Geometric Evidence Acquisition

FUSE:通过自适应语义-几何证据获取实现主动功能可供性接地

Zhou Chen, Sathyanarayanan N. Aakur

专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV

AI总结 该研究提出主动功能可供性接地任务,构建FUSE框架结合不确定性驱动探索与摊销规划器,基于Habitat基准验证其在非神示接地中性能最优且计算量降低1.33倍。

Comments Under review. 15 Pages. 9 tables, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏