arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7370 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7370 篇

1908.10751 2019-08-30 physics.comp-ph physics.geo-ph 78%

A full Stokes subgrid model for simulation of grounding line migration in ice sheets

Gong Cheng, Per Lötstedt, Lina von Sydow

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.06147 2019-07-30 cs.MM eess.IV 78%

Grounding Object Detections With Transcriptions

Yasufumi Moriya, Ramon Sanabria, Florian Metze, Gareth J. F. Jones

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.00347 2019-06-11 cs.CL 78%

Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation

Ronghang Hu, Daniel Fried, Anna Rohrbach, Dan Klein, Trevor Darrell, Kate Saenko

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments ACL 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.00430 2019-06-04 cs.RO physics.app-ph 78%

Effects of Different Hand-Grounding Locations on Haptic Performance With a Wearable Kinesthetic Haptic Device

Sajid Nisar, Melisa Orta Martinez, Takahiro Endo, Fumitoshi Matsuno, Allison M. Okamura

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 8 pages, 11 figures, 1 table

Journal ref IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 351-358, April 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.06966 2018-11-19 cs.RO 78%

Temporal Grounding Graphs for Language Understanding with Accrued Visual-Linguistic Context

Rohan Paul, Andrei Barbu, Sue Felshin, Boris Katz, Nicholas Roy

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Published in ICJAI 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.08266 2018-08-29 cs.CL 78%

A Visual Attention Grounding Neural Model for Multimodal Machine Translation

Mingyang Zhou, Runxiang Cheng, Yong Jae Lee, Zhou Yu

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.11162 2018-06-01 cs.PL cs.LO 78%

Constraint Answer Set Programming without Grounding

Joaquín Arias, Manuel Carro, Elmer Salazar, Kyle Marple, Gopal Gupta

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Paper presented at the 34nd International Conference on Logic Programming (ICLP 2018), Oxford, UK, July 14 to July 17, 2018 18 pages, LaTeX

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.01097 2017-12-05 cs.CL cs.RO 78%

Generalized Grounding Graphs: A Probabilistic Framework for Understanding Grounded Commands

Thomas Kollar, Stefanie Tellex, Matthew Walter, Albert Huang, Abraham Bachrach, Sachi Hemachandra, Emma Brunskill, Ashis Banerjee, Deb Roy, Seth Teller, Nicholas Roy

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Submitted to the Journal of Artificial Intelligence Research

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.10486 2017-10-02 cs.CL 78%

Symbol, Conversational, and Societal Grounding with a Toy Robot

Casey Kennington, Sarah Plane

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 2 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.00501 2017-09-05 cs.LO 78%

Computing Stable Models of Normal Logic Programs Without Grounding

Kyle Marple, Elmer Salazar, Gopal Gupta

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.01127 2016-08-04 cs.RO cs.AI cs.CV cs.LG 78%

Autonomous Grounding of Visual Field Experience through Sensorimotor Prediction

Alban Laflaquière

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG

Comments 6 pages, 4 figures, ICDL-Epirob 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1505.06289 2015-06-08 cs.CL cs.GR 78%

Text to 3D Scene Generation with Rich Lexical Grounding

Angel Chang, Will Monroe, Manolis Savva, Christopher Potts, Christopher D. Manning

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 10 pages, 7 figures, 3 tables. To appear in ACL-IJCNLP 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1402.6889 2015-02-04 cs.LO 78%

Lazy Model Expansion: Interleaving Grounding with Search

Broes De Cat, Marc Denecker, Peter Stuckey, Maurice Bruynooghe

专题命中 视觉定位与Grounding :grounding(title,abstract)

Journal ref Journal of Artificial Intelligence Research, feb 2015, volume 52, pages 235-286

详情

展开后加载摘要…

URL PDF HTML 收藏
1311.5076 2013-11-21 physics.ins-det physics.plasm-ph 78%

Design of a mechanically actuated RF grounding system for the ITER ICRH antenna

D Hancock, M Shannon, B Beaumont, P Dumortier, F Durodie, V Kyrytsya, F Louche, R McKinley, K Nicholls, the CYCLE Team

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 4 pages, 11 figures

Journal ref Proceedings of the 27th Symposium On Fusion Technology (SOFT-27); Liege, Belgium, September 24-28, 2012. Fusion Engineering and Design, Vol.88, Issues 9-10, October 2013, p.2100-2104

详情

展开后加载摘要…

URL PDF HTML 收藏
1111.1570 2011-11-08 cs.IR cs.SI 78%

Semantic Grounding Strategies for Tagbased Recommender Systems

Frederico Durao, Peter Dolog

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 13 pages, 5 figures

Journal ref International Journal of Web & Semantic Technology (IJWesT) Vol.2, No.4, 2011, 67-79

详情

展开后加载摘要…

URL PDF HTML 收藏
0906.2756 2010-11-09 cs.MA cs.LO cs.SE 78%

Norms and Commitment for iOrgs(TM) Information Systems: Direct Logic(TM) and Participatory Grounding Checking

Carl Hewitt

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments expanded article

详情

展开后加载摘要…

URL PDF HTML 收藏
astro-ph/0301095 2009-12-01 astro-ph 78%

Detection of Nine M8.0-L0.5 Binaries: The Very Low Mass Binary Population and its Implications for Brown Dwarf and VLM Star Formation

Laird M. Close, Nick Siegler, Melanie Freed, Beth Biller

专题命中 视觉定位与Grounding :VLM(title,abstract)

Comments To appear in the April 10, 2003 issue of The Astrophysical Journal 30 pages, 17 figures

Journal ref Astrophys.J. 587 (2003) 407-422

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.06179 2022-08-15 cs.CV cs.AI 77%

Exploiting Feature Diversity for Make-up Temporal Video Grounding

Xiujun Shu, Wei Wen, Taian Guo, Sunan He, Chen Wu, Ruizhi Qiao

专题命中 视觉定位与Grounding :grounding(title,comments);分类 cs.CV、cs.AI

Comments 3st Place in PIC Makeup Temporal Video Grounding (MTVG) Challenge in ACM-MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01460 2026-08-19 cs.CV 版本更新 77%

Reinforcing Consistency in Video MLLMs with Structured Rewards

通过结构化奖励强化视频MLLMs的一致性

Yihao Quan, Zeru Shi, Jinman Zhao, Ruixiang Tang

机构 * Rutgers University(罗格斯大学) University of Toronto(多伦多大学)

专题命中 视觉定位与Grounding :grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 研究通过结构化奖励提升视频MLLMs的一致性,发现传统监督不足,提出结合事实和时间单元的奖励机制,提升视频理解准确性。

Comments Accepted by COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19759 2026-08-19 cs.CV 版本更新 77%

Vision-Language Enhanced Foundation Model for Semi-Supervised Medical Image Segmentation

增强视觉-语言能力的半监督医学图像分割基础模型

Jiaqi Guo, Mingzhen Li, Hanyu Su, Keigo Healy, Lexiaozi Fan, Neda Tavakoli, Santiago López-Tapia, Daniel Kim, Aggelos K. Katsaggelos

机构 * ECE, Northwestern University(电气工程与计算机科学系,西北大学) Stats, Northwestern University(统计学系,西北大学) Radiology, Northwestern University(放射学系,西北大学)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 本文提出VESSA模型,通过增强视觉-语言能力的半监督方法提升医学图像分割精度,实验表明其在有限标注条件下表现优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13343 2026-08-14 cs.CV 新提交 77%

AmalthAI: An Open-Source Computer Vision Platform for Cultural Heritage

AmalthAI:面向文化遗产的开源计算机视觉平台

Christos Chatzisavvas, Stelios Alvanos, Efstratios Politis, Panagiotis Rigas, Thomas Pappas, Ioannis Giannoukos, Nikolaos Mitianoudis, Agata Ulanowska, Katarzyna Żebrowska, Nazarij Buławka, Christina Margariti, George Pavlidis, Chairi Kiourt, Anestis Koutsoudis, Vassilis Katsouros, George Ioannakis

机构 * Democritus University of Thrace(德谟克利特色雷斯大学) Athena Research Center(雅典娜研究中心) National and Kapodistrian University of Athens(雅典国立卡波迪斯特里亚大学) University of Warsaw(华沙大学)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 AmalthAI是面向文化遗产领域专家的开源计算机视觉平台,通过集成Kubeflow、Katib等工具,支持数据集管理、模型训练与推理,可保障敏感考古数据安全,已在黏土织物印痕数据集上验证其功能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02711 2026-08-14 cs.CV 版本更新 77%

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

Hunyuan3D-Buffalo 1.0:用于可扩展3D生成、理解与编辑的统一多模态模型

Junliang Ye, Kenkun Liu, Guocun Wang, Yang Li, Yansong Qu, Chunshi Wang, Jingwei Xu, Yunhan Yang, Zibo Zhao, Jiachen Xu, Jiaao Yu, Lifu Wang, Zhihao Liang, Xin Huang, Zhuo Chen, Chunchao Guo

机构 * Tencent Hunyuan(腾讯混元)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract_cn);分类 cs.CV

AI总结 Hunyuan3D-Buffalo 1.0是支持3D理解、文本到3D生成等功能的统一多模态模型,构建87M规模语料库训练,在相关基准上达SOTA或领先性能,验证了统一训练的有效性。

Comments Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12209 2026-08-13 cs.CV 新提交 77%

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

作为辅助监督的生成:通过解耦嵌入预测以零推理开销增强视觉理解

Zhongbin Guo, Jiahao Xie, Dongling Xiao, Qianle Wang, Ruiqi Lu, Xiaomin He, Wanxuan Sun, Cheng Yang

机构 * ByteDance(字节跳动)

专题命中 视觉定位与Grounding :multimodal large language model(abstract,abstract_cn);grounding(abstract);分类 cs.CV

AI总结 本研究提出GAS框架,将视觉生成作为辅助监督,通过解耦MoT架构的NEP实现零推理开销,提升了多模态理解尤其是感知与空间理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14098 2026-08-13 cs.CV 版本更新 77%

ForgeryVCR: Visual-Centric Reasoning via Efficient Forensic Tools in MLLMs for Image Forgery Detection and Localization

ForgeryVCR: 通过高效的取证工具在MLLMs中实现视觉中心推理用于图像伪造检测与定位

Youqi Wang, Shen Chen, Haowei Wang, Rongxuan Peng, Taiping Yao, Shunquan Tan, Changsheng Chen, Bin Li, Shouhong Ding

机构 * Shenzhen University(深圳大学) Tencent Youtu Lab(腾讯优图实验室)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 ForgeryVCR通过高效的取证工具实现视觉中心推理,提升图像伪造检测与定位的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26046 2026-08-11 cs.RO cs.CV 版本更新 77%

RoboAtlas: Contextual Active SLAM

RoboAtlas:上下文感知主动SLAM

Alexander Schperberg, Shivam K. Panda, Abraham P. Vinod, Rokaha Bhagawan, M. K. Jawed, Stefano Di Cairano

机构 * Mitsubishi Electric Research Laboratories(三菱电机研究实验室) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV

AI总结 提出RoboAtlas框架,通过上下文感知的多臂赌博机平衡几何探索与语义推理,结合3D语义地图OpenRoboVox,在真实环境中实现100%任务成功率,并在GOAT-Bench上以90.6%成功率达到SOTA。

Comments Alexander Schperberg and Shivam K. Panda made equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05699 2026-08-07 cs.CV 新提交 77%

TAU-Bench: From Anomaly Instance Tracking to Fine-Grained Video Anomaly Understanding

TAU-Bench:从异常实例跟踪到细粒度视频异常理解

Kepeng Yang, Dongxuan Liu, Rongxin Gao, Zixin Su, Rui Wu, Shuzhao Xie, Chenxin Li, Panwang Pan, Yuzhi Huang, Yue Huang, Jingyan Jiang

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV

AI总结 该研究引入TAU-Bench这一以跟踪为核心的基准,用于联合评估异常实例跟踪与细粒度视频异常理解,发现现有视觉-语言模型存在语义推理与视觉接地的差距,为构建更可靠的视频异常理解系统提供了关键评估方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27703 2026-08-05 cs.AI 版本更新 77%

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

SpatialCLI:先借助空间工具进行推理,再脱离工具进行推理

Yang Zhou, Zixuan Huang, Sunzhu Li, Zhuo Yang, Chen Zhang, Shunian Chen, Caijun Yan, Jianyao Xu, Shunyu Liu, Weijie Fu, Peiliang Li, Xiaozhi Chen, Yuxiang Cai

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.AI

AI总结 本研究提出SpatialCLI框架,通过三个阶段让视觉语言模型先借助专用空间工具推理,再内化感知能力,在MindCube上大幅提升了Qwen3-VL-8B-Instruct的性能,表现优于对比模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07916 2026-08-05 cs.CV 版本更新 77%

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation

Tarot-SAM3: 无需训练的SAM3用于任意指称表达分割

Weiming Zhang, Dingwen Xiao, Songyue Guo, Guangyu Xiang, Shiqi Wen, Minwei Zhao, Lei Chen, Lin Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Nanyang Technological University(南洋理工大学)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 Tarot-SAM3通过引入推理辅助提示和自修复机制,实现对任意指称表达的高效分割,有效解决传统方法对长或隐含表达的处理难题。

Comments We need to make a huge revision

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22768 2026-08-05 cs.CV 77%

From Pixels to Semantics: A Multi-Stage AI Framework for Structural Damage Detection in Satellite Imagery

从像素到语义:一种多阶段AI框架用于卫星图像中的结构损坏检测

Bijay Shakya, Catherine Hoier, Khandaker Mamun Ahmed

机构 * The Beacom College of Computer & Cyber Sciences, Dakota State University(达科他州立大学比康计算机与网络科学学院)

专题命中 视觉定位与Grounding :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV

AI总结 本文提出一种多阶段AI框架,结合超分辨率、目标检测和视觉-语言模型,提升灾害后建筑损坏评估的语义解读能力,通过CLIPScore和多模型策略提高检测可靠性。

Journal ref IEEE/CVF Conference on Computer Vision & Pattern Recognition Workshop (CVPRW) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11385 2026-08-04 cs.CV 版本更新 77%

DeceptionX: From Multimodal Evidence to Explainable Deception Detection

DeceptionX: 基于多模态大语言模型的可解释欺骗检测

Jiayu Zhang, Shuo Ye, Jiajian Huang, Yawen Cui, Taorui Wang, Wei Xia, Zeheng Wang, Haowen Tang, Yelin Wang, Hui Ma, Zitong Yu

机构 * Great Bay University(大湾区大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 提出DeceptionX框架,将欺骗检测从黑箱分类转变为可解释的观察-思考-总结推理过程,通过构建DeceptChain数据集和三阶段训练管道,在标准基准上超越现有方法,同时提供专家级可解释推理路径。

详情

展开后加载摘要…

URL PDF HTML 收藏