arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7370 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7370 篇

2603.05732 2026-03-09 cs.CV 74%

From Phase Grounding to Intelligent Surgical Narratives

从相位接地到智能手术叙述

Ethan Peterson, Huixin Zhan

机构 * New Mexico Institute of Mining and Technology(新墨西哥矿业技术研究所)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 本文提出基于CLIP的多模态框架,通过自动创建手术时间线和叙述,减少手动标注的需要。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00732 2026-03-03 cs.RO cs.CV 74%

UniHM: Unified Dexterous Hand Manipulation with Vision Language Model

UniHM: 一体化的视觉语言模型用于统一的灵巧手操作

Zhenhao Zhang, Jiaxin Liu, Ye Shi, Jingya Wang

机构 * ShanghaiTech University(上海科技大学) InstAdapt

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV

AI总结 UniHM通过统一的视觉语言模型实现灵巧手操作,利用开放词汇指令提升泛化能力和物理可行性。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11244 2026-02-13 cs.CV 74%

Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models

压力测试揭示视频-语言模型中脆弱的时间与视觉基础

Sethuraman T, Savya Khosla, Aditi Tiwari, Vidya Ganesh, Rakshana Jayaprakash, Aditya Jain, Vignesh Srinivasakumar, Onkar Kishor Susladkar, Srinidhi Sunkara, Aditya Shanmugham, Rakesh Vaideeswaran, Abbaas Alif Mohamed Nishar, Simon Jenni, Derek Hoiem

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Adobe Research(Adobe研究) Qualtrics iManage NVIDIA Amazon Science(亚马逊科学) Capital One

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 REVEAL{}通过五个压力测试揭示视频-语言模型在时间与视觉基础方面的脆弱性,展示其在处理视频内容、时间序列和运动方面的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21769 2026-02-12 cs.CV 74%

H2OFlow: Grounding Human-Object Affordances with 3D Generative Models and Dense Diffused Flows

H2OFlow: 通过3D生成模型和密集扩散流接地人类-物体 affordances

Harry Zhang, Luca Carlone

机构 * MIT(麻省理工学院)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 H2OFlow通过3D生成模型和密集扩散流学习人类-物体交互的3D affordances,无需人工标注,有效泛化至现实物体。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16038 2026-01-23 cs.AI 74%

Grounding Large Language Models in Reaction Knowledge Graphs for Synthesis Retrieval

将反应知识图谱接地于大语言模型以实现合成检索

Olga Bunkova, Lorenzo Di Fruscia, Sophia Rupprecht, Artur M. Schweidtmann, Marcel J. T. Reinders, Jana M. Weber

机构 * Department of Intelligent Systems, Delft University of Technology(智能系统系,代尔夫特理工大学) Department of Chemical Engineering, Delft University of Technology(化学工程系,代尔夫特理工大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 本研究通过将反应路径检索转化为图查询生成问题,利用对齐示例的单样本提示提升大语言模型在合成规划中的检索准确性。

Comments Accepted at ML4Molecules 2025 (ELLIS UnConference workshop), Copenhagen, Denmark, December 2, 2025. Workshop page: https://moleculediscovery.github.io/workshop2025/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15533 2026-01-23 cs.AI 74%

From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models

从生成引擎到可操作模拟器:世界模型中物理基础的必要性

Zhikang Chen, Tingting Zhu

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 本文提出将世界模型重新定义为可操作模拟器,强调因果结构和约束意识,以提升医疗决策中的反事实推理和长期预见能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14052 2026-01-21 cs.CV 74%

Vision Also You Need: Navigating Out-of-Distribution Detection with Multimodal Large Language Model

视觉也需要:利用多模态大语言模型进行分布外检测导航

Haoran Xu, Yanlin Liu, Zizhao Tong, Jiaze Li, Kexue Fu, Yuyang Zhang, Longxiang Gao, Shuaiguang Li, Xingyu Li, Yanran Xu, Changwei Wang

机构 * Zhejiang University(浙江大学) Tsinghua University(清华大学) University of Chinese Academy of Sciences(中国科学院大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(教育部计算电力网络与信息安全重点实验室,山东计算机科学中心(国家超算中心济南中心),齐鲁工业大学(山东科学院)) Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science(山东省计算电力互联网与服务计算重点实验室,山东省计算机科学基础研究中心) University of Electronic Science and Technology of China(电子科技大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) RWTH Aachen University(亚琛工业大学)

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.CV

AI总结 本文提出MM-OOD方法,利用多模态大语言模型的推理能力,通过多轮对话增强分布外检测,提升近远OOD任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06204 2026-01-19 cs.CV cs.MA 74%

Cascading multi-agent anomaly detection in surveillance systems via vision-language models and embedding-based classification

通过视觉-语言模型和基于嵌入的分类实现监视系统中的级联多代理异常检测

Tayyab Rehman, Giovanni De Gasperis, Aly Shmahell

机构 * University of L’Aquila(拉奎拉大学) SPEE S.R.L(SPEE公司)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

AI总结 本文提出一种级联多代理框架,结合视觉-语言模型和嵌入分类,实现高效且可解释的异常检测,减少延迟并提升监控系统性能。

Comments Author email changed, Acknowlegement changes

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07795 2026-01-13 cs.CV 74%

Vision-Language Model for Accurate Crater Detection

面向准确陨石坑检测的视觉-语言模型

Patrick Bauer, Marius Schwinning, Florian Renk, Andreas Weinmann, Hichem Snoussi

机构 * University of Technology of Troyes(图卢兹技术大学) Hochschule Darmstadt(达姆施塔特应用科学大学) GMV for European Space Agency(欧洲航天局GMV) European Space Agency(欧洲航天局) Technische Hochschule Würzburg-Schweinfurt(维尔茨堡-施韦因富特技术大学)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

AI总结 本文提出基于Vision Transformer的深度学习模型,用于在复杂月球成像条件下实现高精度陨石坑检测,通过低秩适应策略和组合损失函数提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02029 2026-01-06 cs.CV 74%

Leveraging 2D-VLM for Label-Free 3D Segmentation in Large-Scale Outdoor Scene Understanding

利用2D-VLM实现无标注3D分割以支持大规模户外场景理解

Toshihiko Nishimura, Hirofumi Abe, Kazuhiko Murasaki, Taiga Yoshida, Ryuichi Tanida

机构 * NTT Corporation(日本电报电话公司)

专题命中 视觉定位与Grounding :VLM(title);分类 cs.CV

AI总结 本文提出基于2D-VLM的无标注3D分割方法,通过虚拟相机投影和自然语言提示实现大规模户外场景的语义分割,支持开放词汇识别。

Comments 19

Journal ref 19th International Conference on Machine Vision Applications (MVA2025), IEICE Transactions on Information and Systems letter

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11792 2025-12-23 cs.AI 74%

Solver-Informed RL: Grounding Large Language Models for Authentic Optimization Modeling

Solver-Informed RL: 为真实优化建模奠定大语言模型基础

Yitian Chen, Jingfan Xia, Siyu Shao, Dongdong Ge, Yinyu Ye

机构 * Cardinal Operations, China(中国卡迪纳尔运营公司) Shanghai University of Finance and Economics(上海财经大学) The University of Hong Kong(香港大学) Antai School of Economics and Management, Shanghai Jiao Tong University(上海交通大学安泰经济管理学院) Department of Management Science and Engineering, Stanford University(斯坦福大学管理科学与工程系)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 SIRL通过强化学习与外部优化求解器结合,提升大语言模型在优化建模中的准确性与实用性。

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02967 2025-12-16 cs.CL cs.AI cs.IR 74%

Grounding Large Language Models in Clinical Evidence: A Retrieval-Augmented Generation System for Querying UK NICE Clinical Guidelines

将大语言模型 grounded 在临床证据中:一种用于查询英国 NICE 临床指南的检索增强生成系统

Matthew Lewis, Samuel Thio, Amy Roberts, Catherine Siju, Whoasif Mukit, Rebecca Kuruvilla, Zhangshu Joshua Jiang, Niko Möller-Grell, Aditya Borakati, Richard JB Dobson, Spiros Denaxas

机构 * Institute of Health Informatics, University College London(健康信息学研究所,伦敦大学学院) Department of Biostatistics and Health Informatics, King’s College London(生物统计学与健康信息学系,伦敦国王学院) EPSRC DRIVE-Health CDT, London, U.K.(EPSRC DRIVE-Health CDT,伦敦,英国) Imperial College Healthcare NHS Trust(帝国学院医疗 NHS 信托) Black Country Healthcare NHS Foundation Trust(黑斯廷斯地区医疗 NHS 基础信托) Feldon Lane Surgery, The Dudley Group NHS Foundation Trust(费尔顿车道诊所,德比集团 NHS 基础信托) The Cleveland Clinic, London, U.K.(克利夫兰诊所,伦敦,英国) CogStack Limited(CogStack 有限公司)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

AI总结 本文提出一种基于RAG的系统,用于查询英国NICE临床指南,通过检索增强生成技术提升回答的准确性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10195 2025-12-12 cs.CL cs.LG cs.MA 74%

AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding

AutoMedic: 一种用于具有医学数据集支撑的临床对话代理的自动化评估框架

Gyutaek Oh, Sangjoon Park, Byung-Hoon Kim

机构 * Yonsei University College of Medicine(延世大学医学院) Yonsei Institute for Digital Health(延世大学数字健康研究所) Yonsei University(延世大学) Institute of Behavioral Sciences in Medicine(医学行为科学研究所)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG

AI总结 AutoMedic是一种基于医学数据集的自动化评估框架,用于评估临床对话代理的多方面性能,包括准确性、效率、同理心和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06335 2025-12-10 cs.CV 74%

Harnessing Object Grounding for Time-Sensitive Video Understanding

利用物体接地提升时间敏感视频理解

Tz-Ying Wu, Sharath Nittur Sridhar, Subarna Tripathi

机构 * Intel(英特尔公司)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

AI总结 本文提出 GO-Tokenizer 以提升视频大型语言模型的时间敏感视频理解能力,通过实时编码紧凑的物体信息,提高模型性能并减少噪声影响。

Comments Accepted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22256 2025-12-01 cs.CV 74%

UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation

UMind-VL:一种通用的超声视觉-语言模型,用于统一的 grounded perception 和全面的 interpretation

Dengbo Chen, Ziwei Zhao, Kexin Zhang, Shishuang Zhao, Junjie Hou, Yaqian Wang, Nianxi Liao, Anlan Sun, Fei Gao, Jia Ding, Yuhang Liu, Dong Wang

机构 * Yizhun Medical AI Team(义诊医疗AI团队)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

AI总结 UMind-VL 是一种通用超声视觉-语言模型,通过统一的 grounded perception 和 comprehensive interpretation 实现对医学影像的高效理解和诊断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23258 2025-10-07 cs.CV 74%

OracleGS: Grounding Generative Priors for Sparse-View Gaussian Splatting

Atakan Topaloglu, Kunyi Li, Michael Niemeyer, Nassir Navab, A. Murat Tekalp, Federico Tombari

机构 * ETH Zürich(苏黎世联邦理工学院) Koç University(科卡大学) KUIS AI Center(KUIS人工智能中心) Technical University of Munich(慕尼黑技术大学) Google(谷歌) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Project page available at: https://atakan-topaloglu.github.io/oraclegs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02592 2025-10-06 cs.AI 74%

Multimodal Large Language Model Framework for Safe and Interpretable Grid-Integrated EVs

Jean Douglas Carvalho, Hugo Kenji, Ahmad Mohammad Saber, Glaucia Melo, Max Mauro Dias Santos, Deepa Kundur

专题命中 视觉定位与Grounding :multimodal large language model(title);分类 cs.AI

Comments This paper has been presented at the 2025 IEEE PES Conference on Innovative Smart Grid Technologies (ISGT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12145 2025-09-16 cs.CV 74%

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

Hyolim Kang, Yunsu Park, Youngbeom Yoo, Yeeun Choi, Seon Joo Kim

机构 * Yonsei University(延世大学)

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06564 2025-08-14 cs.CV 74%

Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC

Guanyu Hu, Dimitrios Kollias, Xinyu Yang

机构 * Xi'an Jiaotong University(西安交通大学) Queen Mary University of London(伦敦女王玛丽大学) Center for Multimodal AI(多模态人工智能中心) Digital Environment Research Institute(数字环境研究院)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments accepted for publication at ACM Multimedia (ACM MM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08761 2025-08-13 cs.CL cs.AI 74%

DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation

Stavros Doropoulos, Stavros Vologiannidis, Ioannis Magnisalis

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07466 2025-08-12 cs.AI 74%

Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs

Dom Huh, Prasant Mohapatra

机构 * UC Davis(加州大学戴维斯分校) University of South Florida(佛罗里达州立大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21649 2025-07-30 cs.CV 74%

The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM

Shibo Gao, Peipei Yang, Haiyang Guo, Yangyang Liu, Yi Chen, Shuai Li, Han Zhu, Jian Xu, Xu-Yao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 视觉定位与Grounding :MLLM(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18276 2025-07-25 cs.RO cs.CV 74%

Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding

Xiaojie Zhang, Yuanfei Wang, Ruihai Wu, Kunqi Xu, Yu Li, Liuyu Xiang, Hao Dong, Zhaofeng He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) School of Computer Science, Peking University(北京大学计算机学院) School of EECS, Peking University(北京大学电子工程学院)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02600 2025-07-22 cs.CV cs.RO eess.IV 74%

Resource-Efficient Affordance Grounding with Complementary Depth and Semantic Prompts

Yizhou Huang, Fan Yang, Guoliang Zhu, Gen Li, Hao Shi, Yukun Zuo, Wenrui Chen, Zhiyong Li, Kailun Yang

机构 * School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China(人工智能与机器人学院和机器人视觉感知与控制技术国家工程研究中心,湖南大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Accepted to IROS 2025. The source code will be made publicly available at https://github.com/DAWDSE/BiT-Align

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05336 2025-07-08 cs.CV 74%

VideoMolmo: Spatio-Temporal Grounding Meets Pointing

Ghazi Shazan Ahmad, Ahmed Heakl, Hanan Gani, Abdelrahman Shaker, Zhiqiang Shen, Fahad Shahbaz Khan, Salman Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Linköping University(林奈大学) Australian National University(澳大利亚国立大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16769 2025-07-03 cs.CV 74%

Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models

Muhammad Atta ur Rahman, Dooseop Choi, Seung-Ik Lee, KyoungWook Min

机构 * Artificial Intelligence Creative Research Lab, ETRI University of Science(人工智能创意研究实验室,ETRI大学) Field Robotics Research Section, ETRI University of Science(机器人领域研究部,ETRI大学) Artificial Intelligence Creative Research Lab, ETRI Daejeon, South Korea(人工智能创意研究实验室,ETRI大田,韩国)

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV

Comments Accepted at the 17th IEEE International Conference on Advanced Computational Intelligence (ICACI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07007 2025-07-02 cs.CV 74%

Grounding Creativity in Physics: A Brief Survey of Physical Priors in AIGC

Siwei Meng, Yawei Luo, Ping Liu

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Accepted by IJCAI 2025 Survey Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22442 2025-07-01 cs.LG 74%

Features-based embedding or Feature-grounding

Piotr Makarevich

专题命中 视觉定位与Grounding :grounding(title);分类 cs.LG

Comments 13 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05266 2025-06-25 cs.CV q-bio.NC 74%

Brain Mapping with Dense Features: Grounding Cortical Semantic Selectivity in Natural Images With Vision Transformers

Andrew F. Luo, Jacob Yeung, Rushikesh Zawar, Shaurya Dewan, Margaret M. Henderson, Leila Wehbe, Michael J. Tarr

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV

Comments Accepted at ICLR 2025, code: https://github.com/aluo-x/BrainSAIL

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15802 2025-06-24 cs.CV 74%

Visual Prompt Engineering for Vision Language Models in Radiology

Stefan Denner, Markus Bujotzek, Dimitrios Bounias, David Zimmerer, Raphael Stock, Klaus Maier-Hein

机构 * Division of Medical Image Computing, German Cancer Research Center, Heidelberg, Germany(德国癌症研究中心医学图像计算部) Faculty of Mathematics and Computer Science, Heidelberg University, Heidelberg, Germany(海德堡大学数学与计算机科学学院) Medical Faculty Heidelberg, University of Heidelberg, Heidelberg, Germany(海德堡大学医学学院)

专题命中 视觉定位与Grounding :vision language model(title);分类 cs.CV

Comments Accepted at ECCV 2024 Workshop on Emergent Visual Abilities and Limits of Foundation Models & Medical Imaging with Deep Learning 2025

详情

展开后加载摘要…

URL PDF HTML 收藏