arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1561 信号源:cs.CV, cs.AI, cs.LG

1. GUI与屏幕智能体 1561 篇

2510.11014 2026-06-08 cs.RO cs.AI cs.CV 版本更新 82%

MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models

MatterDoor: 使用生成模型采样零样本空间语义先验

Subhransu S. Bhattacharjee, Hao Lu, Dylan Campbell, Rahul Shome

机构 * School of Computing, Australian National University(澳大利亚国立大学计算机学院)

专题命中 GUI与屏幕智能体 :VLM(summary_cn,abstract);分类 cs.CV、cs.AI

AI总结 针对机器人通过门缝观察时场景结构缺失的问题,提出MatterDoor方法,利用预训练生成模型(VLM引导外推、单目深度估计、语义分割)采样隐藏房间的语义3D点云先验,在Matterport3D基准上验证了零样本空间语义先验的有效性。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04806 2026-06-04 cs.CV cs.AI 82%

NoRA: Evaluating Grounded Reasonableness in Visual First-person Normative Action Reasoning

NoRA: 评估视觉第一人称规范性动作推理中的基于事实的合理性

Sichao Li, Sai Ma, Daniel Kilov, Secil Yanik Guyot, Zhuang Li, Seth Lazar

机构 * The University of Sydney(悉尼大学) Australian National University(澳大利亚国立大学) RMIT University(皇家墨尔本理工大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 GUI与屏幕智能体 :VLM(summary_cn,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI

AI总结 提出NoRA基准,通过事实-理由-动作支持图评估多模态模型生成合理动作并基于可见事实进行推理的能力,发现当前VLM在构建完整动作空间和绑定正确支持方面存在不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09196 2026-08-11 cs.RO 新提交 82%

SAIN: Structure-Aware Interactive Navigation with Active Dialogue Grounding for Mobile Robot

SAIN:面向移动机器人的、结合主动对话 grounding 的结构感知交互式导航

Yuhao Cao, Xiao Liu, Yang Xie, Lu Liu, Haoyao Chen

机构 * School of Mechanical Engineering and Automation, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)机械工程与自动化学院) Department of Mechanical Engineering, City University of Hong Kong(香港城市大学机械工程系)

专题命中 GUI与屏幕智能体 :grounding(title,title_cn)

AI总结 本文提出SAIN零样本框架,将主动对话转化为持久导航状态,在VL-LN IIGN基准上提升了导航成功率与加权成功率,无需特定任务策略训练,验证了对话转状态机制的有效性。

Comments 8 pages, 4 figures, and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18580 2026-07-22 cs.RO 新提交 82%

STeP: Signal Temporal Logic for Precise Specifications for Action Generation with Vision Language Models

STeP:用于视觉语言模型动作生成精确规范的信号时序逻辑

Kasra Torshizi, Anukriti Singh, Sidharth Mathur, Khuzema Habib, Leo Du, Pratap Tokekar

机构 * University of Maryland(马里兰大学)

专题命中 GUI与屏幕智能体 :vision language model(title);VLM(abstract,abstract_cn)

AI总结 针对视觉语言动作模型缺乏可解释性及难以遵循精确自然语言指令的问题,提出用信号时序逻辑(STL)连接高级语言理解与低级机器人执行的分层框架,经实验验证该框架能提高语言条件下机器人规划的精度、可靠性和可解释性。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21268 2026-04-24 cs.LG cs.AI cs.CV 82%

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding

两次测量,一次点击:通过强化学习共进化提议者和视觉评论者进行GUI定位

Wenkai Wang, Xiyun Li, Hongcan Guo, Wenhao Yu, Tianqing Fang, Haitao Mi, Dong Yu, Shengyu Zhang

机构 * Zhejiang University(浙江大学) Tencent AI Lab(腾讯AI实验室) The University of Hong Kong(香港大学)

专题命中 GUI与屏幕智能体 :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出通过强化学习共进化提议者和视觉评论者,以提升GUI定位的准确性和鲁棒性,通过动态平衡训练目标,增强模型在复杂界面布局中的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21142 2026-03-24 cs.RO 82%

Dynamic Control Barrier Function Regulation with Vision-Language Models for Safe, Adaptive, and Realtime Visual Navigation

动态控制屏障函数调节与视觉-语言模型用于安全、适应性和实时视觉导航

Jeffrey Chen, Rohan Chandra

机构 * Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系)

专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);VLM(abstract)

AI总结 本文提出AlphaAdj框架,利用视觉-语言模型动态调整控制屏障函数参数,以实现在动态环境中安全、高效和实时的导航。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15541 2026-03-18 cs.RO 82%

CompliantVLA-adaptor: VLM-Guided Variable Impedance Action for Safe Contact-Rich Manipulation

CompliantVLA-adaptor:基于视觉-语言模型的变量阻抗控制用于安全的高接触密度操作

Heng Zhang, Wei-Hsing Huang, Qiyi Tong, Gokhan Solak, Puze Liu, Kaidi Zhang, Sheng Liu, Jan Peters, Yu She, Arash Ajoudani

机构 * Human-Robot Interfaces and Interaction Lab, Istituto Italiano di Tecnologia, Genoa, Italy(人机交互实验室,意大利理工学院,热那亚,意大利) Ph.D. program of national interest in Robotics and Intelligent Machines (DRIM) and Università di Genova, Genoa, Italy(机器人与智能机器国家利益博士项目和热那亚大学,热那亚,意大利) Edwardson School of Industrial Engineering, Purdue University, West Lafayette, IN 47907, USA(工业工程埃德华森学校,普渡大学,西拉法克萨,印第安纳州47907,美国) Georgia Institute of Technology, Atlanta, USA(佐治亚理工学院,亚特兰大,美国) German Research Center for AI, Germany(德国人工智能研究中心,德国) TU Darmstadt, Darmstadt, Germany(图宾根大学,图宾根,德国) Karlsruhe Institute of Technology, Karlsruhe, Germany(卡尔斯鲁厄理工学院,卡尔斯鲁厄,德国)

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract)

AI总结 本文提出CompliantVLA-adaptor,通过引入基于视觉语言模型的上下文感知变量阻抗控制,提升接触密集任务的安全性和有效性。方法通过图像和自然语言解读任务上下文,调节阻抗控制器的刚度和阻尼参数,并利用实时力/扭矩反馈确保安全。实验表明在模拟和现实任务中均优于基线方法。

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11447 2026-03-13 cs.RO 82%

Enhancing Lightweight Vision Language Models through Group Competitive Learning for Socially Compliant Navigation

通过群体竞争学习增强轻量级视觉语言模型以实现符合社会规范的导航

Xinyu Zhang, Atsushi Konno, Toshihiko Yamasaki, Ling Xiao

机构 * Hokkaido University(北海道大学) The University of Tokyo(东京大学)

专题命中 GUI与屏幕智能体 :vision language model(title,abstract);VLM(abstract)

AI总结 本文提出群体竞争学习策略,通过协调全局语义与分布正则化,提升轻量级VLMs在社交导航中的推理与决策能力,实现高准确性和高效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02338 2026-03-06 cs.SE cs.RO 82%

Vision Language Model-based Testing of Industrial Autonomous Mobile Robots

基于视觉语言模型的工业自主移动机器人测试

Jiahui Wu, Chengjie Lu, Aitor Arrieta, Shaukat Ali, Thomas Peyrucain

机构 * Simula Research Laboratory and University of Oslo(Simula研究实验室和奥斯陆大学) Mondragon University(蒙dragon大学) Simula Research Laboratory(Simula研究实验室) PAL Robotics(PAL机器人技术)

专题命中 GUI与屏幕智能体 :vision language model(title,abstract);VLM(abstract)

AI总结 本文提出基于视觉语言模型的测试方法,用于生成违反功能和安全要求的机器人交互场景,以提高自主移动机器人在复杂环境中的安全性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10910 2026-02-12 cs.RO 82%

Safe mobility support system using crowd mapping and avoidance route planning using VLM

基于人群映射与避让路线规划的安全移动支持系统

Sena Saito, Kenta Tabata, Renato Miyagusuku, Koichi Ozaki

机构 * Graduate School of Regional Development and Creativity, Division of Engineering and Agriculture, Graduate School, Utsunomiya University(乌市大学研究生院区域发展与创意学院,工程与农业系)

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract)

AI总结 本文提出利用视觉-语言模型和高斯过程回归生成动态人群密度地图,以提升自主机器人在拥挤环境中的安全导航能力。

Journal ref 2025 IEEE International Conference on Real-time Computing and Robotics (RCAR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03956 2026-01-08 cs.RO 82%

CoINS: Counterfactual Interactive Navigation via Skill-Aware VLM

基于技能感知的反事实交互导航:通过技能感知视觉语言模型

Kangjie Zhou, Zhejia Wen, Zhiyong Zhuo, Zike Yan, Pengying Wu, Ieng Hou U, Shuaiyang Li, Han Gao, Kang Ding, Wenhan Cao, Wei Pan, Chang Liu

机构 * School of Advanced Manufacturing and Robotics, Peking University(北京大学先进制造与机器人学院) Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong(香港中文大学机械与自动化工程系) College of Design and Engineering, National University Of Singapore(新加坡国立大学设计与工程学院) Department of Computer Science, The University of Manchester(曼彻斯特大学计算机科学系)

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract)

AI总结 CoINS通过整合技能感知推理与稳健执行,提升机器人在复杂环境中的交互导航能力。

Comments 17 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22238 2025-12-30 cs.LG cs.AI cs.CV 82%

Masking Teacher and Reinforcing Student for Distilling Vision-Language Models

掩码教师与强化学生用于蒸馏视觉-语言模型

Byung-Kwan Lee, Yu-Chiang Frank Wang, Ryo Hachiuma

专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 Masters通过掩码教师和强化学生的方法,解决视觉-语言模型蒸馏中的大小差距问题,提升学生模型的表示学习能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20674 2025-12-25 cs.LG cs.AI cs.CV 82%

HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model

HyDRA:面向移动视觉语言模型的分层和动态秩适应

Yuanhao Xi, Xiaohuan Bing, Ramin Yahyapour

机构 * Liaoning Technical University(辽宁技术大学) University of Göttingen(哥廷根大学)

专题命中 GUI与屏幕智能体 :vision language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 HyDRA通过分层和动态秩调度优化,提升移动视觉语言模型的微调效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13974 2025-12-17 cs.RO 82%

Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline

基于移动机器人的自主施工工地安全检查:一种多层VLM-LLM流水线

Hossein Naderi, Alireza Shojaei, Philip Agee, Kereshmeh Afsari, Abiola Akanmu

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision language model(abstract)

AI总结 本文提出一种多层VLM-LLM流水线,通过机器人自主导航与人工智能结合,实现施工工地安全检查的自动化报告生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13248 2025-11-18 cs.CR 82%

DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents

Fuyao Zhang, Jiaming Zhang, Che Wang, Xiongtao Sun, Yurong Hao, Guowei Guan, Wenjie Li, Longtao Huang, Wei Yang Bryan Lim

专题命中 GUI与屏幕智能体 :MLLM(title,abstract);multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04991 2025-10-14 cs.RO 82%

Efficient Navigation in Unknown Indoor Environments with Vision-Language Models

D. Schwartz, K. Kondo, J. P. How

机构 * ETH Zürich(苏黎世联邦理工学院) Massachusetts Institute of Technology(麻省理工学院)

专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);VLM(abstract)

Comments 7 pages, 4 figures, accepted to the OWN workshop at IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09230 2025-10-13 cs.CV cs.AI cs.CL cs.LG 82%

Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras

Jindong Hong, Wencheng Zhang, Shiqin Qiao, Jianhai Chen, Jianing Qiu, Chuanyang Zheng, Qian Xu, Yun Ji, Qianyue Wen, Weiwei Sun, Hao Li, Huizhen Li, Huichao Wang, Kai Wu, Meng Li, Yijun He, Lingjie Luo, Jiankai Sun

机构 * Bytedance(字节跳动) Peking University(北京大学) Peking University People’s Hospital(北京大学人民医院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 GUI与屏幕智能体 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02723 2025-10-02 cs.RO 82%

ImpedanceGPT: VLM-driven Impedance Control of Swarm of Mini-drones for Intelligent Navigation in Dynamic Environment

Faryal Batool, Yasheerah Yaqoot, Malaika Zafar, Roohan Ahmed Khan, Muhammad Haris Khan, Aleksey Fedoseev, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory, Skolkovo Institute of Science and Technology(智能空间机器人实验室,斯克尔科沃科学与技术研究所)

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract)

Comments Accepted in IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07814 2025-08-12 cs.RO 82%

SwarmVLM: VLM-Guided Impedance Control for Autonomous Navigation of Heterogeneous Robots in Dynamic Warehousing

Malaika Zafar, Roohan Ahmed Khan, Faryal Batool, Yasheerah Yaqoot, Ziang Guo, Mikhail Litvinov, Aleksey Fedoseev, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory, Skolkovo Institute of Science and Technology(智能空间机器人实验室,斯克洛科沃科学与技术研究所)

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04482 2025-08-07 cs.AI cs.CL cs.CV cs.LG 82%

OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use

Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shenzhi Wang, Xinchen Xu, Shuofei Qiao, Zhaokai Wang, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang, Keting Yin, Zhou Zhao, Hongxia Yang, Fan Wu, Shengyu Zhang, Fei Wu

机构 * Zhejiang University(浙江大学) Fudan University(复旦大学) OPPO AI Center(OPPO AI中心) University of Chinese Academy of Sciences(中国科学院大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 GUI与屏幕智能体 :MLLM(title);grounding(abstract);分类 cs.CV、cs.AI、cs.LG

Comments ACL 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11940 2025-08-05 cs.CE 82%

MLLM-based Discovery of Intrinsic Coordinates and Governing Equations from High-Dimensional Data

Ruikun Li, Yan Lu, Shixiang Tang, Biqing Qi, Wanli Ouyang

专题命中 GUI与屏幕智能体 :MLLM(title,abstract);multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16703 2025-06-23 cs.RO 82%

VLM-Empowered Multi-Mode System for Efficient and Safe Planetary Navigation

Sinuo Cheng, Ruyi Zhou, Wenhao Feng, Huaiguang Yang, Haibo Gao, Zongquan Deng, Liang Ding

机构 * State Key Laboratory of Robotics and System, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract)

Comments accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01881 2025-06-16 cs.CV cs.AI cs.LG cs.MM cs.RO 82%

PhysNav-DG: A Novel Adaptive Framework for Robust VLM-Sensor Fusion in Navigation Applications

Trisanth Srinivasan, Santosh Patapati

机构 * Cyrion Labs(塞里昂实验室)

专题命中 GUI与屏幕智能体 :VLM(title);vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at IEEE/CVF Computer Society Conference on Computer Vision and Pattern Recognition Workshops 2025 (CVPRW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02454 2025-05-19 cs.RO 82%

VL-TGS: Trajectory Generation and Selection using Vision Language Models in Mapless Outdoor Environments

Daeun Song, Jing Liang, Xuesu Xiao, Dinesh Manocha

机构 * Department of Computer Science, George Mason University(乔治·马歇尔大学计算机科学系) Department of Computer Science, University of Maryland(马里兰大学计算机科学系)

专题命中 GUI与屏幕智能体 :vision language model(title);VLM(abstract);visual language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16073 2025-04-23 cs.CL 82%

Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation

Zhiyuan Hu, Shiyun Xiong, Yifan Zhang, See-Kiong Ng, Anh Tuan Luu, Bo An, Shuicheng Yan, Bryan Hooi

机构 * National University of Singapore(新加坡国立大学) Skywork AI Nanyang Technological University(南洋理工大学)

专题命中 GUI与屏幕智能体 :VLM(title,abstract);visual language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09000 2025-04-15 cs.RO 82%

CL-CoTNav: Closed-Loop Hierarchical Chain-of-Thought for Zero-Shot Object-Goal Navigation with Vision-Language Models

Yuxin Cai, Xiangkun He, Maonan Wang, Hongliang Guo, Wei-Yun Yau, Chen Lv

专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05225 2025-04-08 cs.RO 82%

Vision-Language Model Predictive Control for Manipulation Planning and Trajectory Generation

Jiaming Chen, Wentao Zhao, Ziyu Meng, Donghui Mao, Ran Song, Wei Pan, Wei Zhang

专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02357 2025-04-04 cs.SE 82%

ReuseDroid: A VLM-empowered Android UI Test Migrator Boosted by Active Feedback

Xiaolei Li, Jialun Cao, Yepang Liu, Shing-Chi Cheung, Hailong Wang

专题命中 GUI与屏幕智能体 :VLM(title,abstract);vision-language model(abstract)

Comments 13 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02193 2024-10-04 cs.RO 82%

Guiding Long-Horizon Task and Motion Planning with Vision Language Models

Zhutian Yang, Caelan Garrett, Dieter Fox, Tomás Lozano-Pérez, Leslie Pack Kaelbling

专题命中 GUI与屏幕智能体 :vision language model(title);vision-language model(abstract);VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00231 2024-10-02 cs.RO cs.AI cs.CV cs.LG 82%

Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models

Qi Wu, Zipeng Fu, Xuxin Cheng, Xiaolong Wang, Chelsea Finn

专题命中 GUI与屏幕智能体 :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments Project website: https://helpful-doggybot.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏