arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1565 信号源:cs.CV, cs.AI, cs.LG

1. GUI与屏幕智能体 1565 篇

2602.15400 2026-02-18 cs.RO 50%

One Agent to Guide Them All: Empowering MLLMs for Vision-and-Language Navigation via Explicit World Representation

一个引导所有代理:通过显式世界表示增强多模态大语言模型用于视觉-语言导航

Zerui Li, Hongpei Zheng, Fangguo Zhao, Aidan Chan, Jian Zhou, Sihao Lin, Shijie Li, Qi Wu

机构 * Australian Institute for Machine Learning, Adelaide University(澳大利亚机器学习研究所,阿德莱德大学) The University of Manchester(曼彻斯特大学) Zhejiang University(浙江大学) Agency for Science, Technology and Research (A*STAR)(科技研究局(A*STAR))

专题命中 GUI与屏幕智能体 :multimodal large language model(abstract)

AI总结 通过显式世界表示增强多模态大语言模型,实现视觉-语言导航的解耦框架,提升导航性能与现实应用能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14326 2026-02-17 cs.NI 50%

Diffusion-based Dynamic Contract for Federated AI Agent Construction in Mobile Metaverses

基于扩散的动态合同用于移动元宇宙中联邦AI代理的构建

Jinbo Wen, Jiawen Kang, Yang Zhang, Yue Zhong, Dusit Niyato, Jie Xu, Jianhang Tang, Chau Yuen

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

AI总结 本文提出基于扩散模型的EDMSAC算法,通过动态合同机制优化边缘-云协作下的AI代理构建,提升移动元宇宙中服务的低延迟与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11862 2026-02-13 cs.RO 50%

LAMP: Implicit Language Map for Robot Navigation

LAMP:机器人导航的隐式语言地图

Sibaek Lee, Hyeonwoo Yu, Giseop Kim, Sunwook Choi

机构 * Department of Intelligent Robotics, Sungkyunkwan University(智能机器人学系,全北国立大学) NAVER LABS Department of Robotics and Mechatronics Engineering, DGIST(机器人与机电工程系,韩国科学技术院)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

AI总结 LAMP通过隐式语言地图实现高效机器人导航,结合隐式神经场与稀疏图进行粗到细路径优化,提升大环境下的内存效率和目标精度。

Comments Accepted for publication in IEEE Robotics and Automation Letters (RA-L). Project page: https://lab-of-ai-and-robotics.github.io/LAMP/

Journal ref IEEE Robotics and Automation Letters (RA-L), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10717 2026-02-12 cs.RO 50%

Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation

说、梦、做:学习视频世界模型以驱动指令式机器人操作

Songen Gu, Yunuo Cai, Tianyu Wang, Simo Wu, Yanwei Fu

机构 * Fudan University(复旦大学)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

AI总结 本文提出一种视频条件动作框架,通过生成稳健的视频模型和对抗性蒸馏,提升机器人操作中的预测能力和空间准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09973 2026-02-11 cs.RO 50%

RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation

RoboInter: 一种面向机器人操控的综合中间表示套件

Hao Li, Ziqin Wang, Zi-han Ding, Shuai Yang, Yilun Chen, Yang Tian, Xiaolin Hu, Tai Wang, Dahua Lin, Feng Zhao, Si Liu, Jiangmiao Pang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Beihang University(北航) Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

AI总结 RoboInter通过统一的中间表示套件,提升机器人操控中视觉-语言-动作系统的泛化能力和推理能力。

Comments Published to ICLR 2026, 69 pages, 40 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07629 2026-02-11 cs.RO 50%

LCLA: Language-Conditioned Latent Alignment for Vision-Language Navigation

LCLA:语言引导的潜在对齐用于视觉-语言导航

Nitesh Subedi, Adam Haroon, Samuel Tetteh, Prajwal Koirala, Cody Fleming, Soumik Sarkar

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

AI总结 LCLA通过将视觉-语言观测对齐到专家策略的潜在空间,实现轻量级的视觉-运动学习,提升在不同环境下的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08537 2026-02-10 cs.RO 50%

UniPlan: Vision-Language Task Planning for Mobile Manipulation with Unified PDDL Formulation

UniPlan: 为移动操作统一PDDL形式的视觉-语言任务规划

Haoming Ye, Yunxiao Xiao, Cewu Lu, Panpan Cai

专题命中 GUI与屏幕智能体 :VLM(abstract)

AI总结 UniPlan通过统一PDDL形式实现视觉-语言任务规划,提升大规模室内移动操作的规划效率和效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21602 2026-02-04 cs.RO 50%

AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation

AIR-VLA:面向空中操作的视觉-语言-动作系统

Jianli Sun, Bin Tian, Qiyao Zhang, Chengxiang Li, Zihan Song, Zhiyong Cui, Yisheng Lv, Yonglin Tian

机构 * The Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Automation, Beijing Institute of Technology(北京理工大学自动化学院) School of Information and Intelligent Engineering, University of Sanya(三亚大学信息与智能工程学院) School of Mechanical and Vehicle Engineering, Hunan University(湖南大学机械与车辆工程学院) State Key Lab of Intelligent Transportation Systems, School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院)

专题命中 GUI与屏幕智能体 :VLM(abstract)

AI总结 AIR-VLA提出首个针对空中操作的视觉-语言-动作系统,通过构建仿真环境和多模态数据集,评估主流模型并揭示其在无人机移动、机械臂控制和高层规划中的能力和限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14587 2026-01-22 cs.HC cs.RO 50%

Explainable OOHRI: Communicating Robot Capabilities and Limitations as Augmented Reality Affordances

可解释的面向对象人机交互:通过增强现实 affordances 传达机器人能力和限制

Lauren W. Wang, Mohamed Kari, Parastoo Abtahi

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

AI总结 本文提出X-OOHRI,一种通过AR界面传达机器人能力和限制的可解释面向对象人机交互系统,通过视觉符号和颜色编码提升人机交互效率。

Journal ref Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction (HRI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08325 2026-01-14 cs.RO 50%

ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation

ActiveVLA: 向视觉-语言-动作模型注入主动感知以实现精确的3D机器人操控

Zhenyang Liu, Yongchong Gu, Yikai Wang, Xiangyang Xue, Yanwei Fu

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Nanyang Technological University(南洋理工大学)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

AI总结 ActiveVLA通过引入主动感知能力,提升机器人在复杂环境中的高精度3D操控性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00634 2026-01-14 cs.SE cs.HC 50%

Does GenAI Make Usability Testing Obsolete?

生成式AI会使可用性测试过时吗?

Ali Ebrahimi Pourasad, Walid Maalej

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

AI总结 本文提出UX-LLM,一种基于大视觉语言模型的可用性问题预测工具,虽无法完全替代传统测试,但可作为补充,尤其适用于资源有限的小型团队。

Comments Accepted for publication at The 47th IEEE/ACM International Conference on Software Engineering ICSE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01948 2026-01-06 cs.RO 50%

Learning Diffusion Policy from Primitive Skills for Robot Manipulation

从基本技能学习扩散策略用于机器人操作

Zhihao Gu, Ming Yang, Difan Zou, Dong Xu

机构 * Dong Xu(东旭教授)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

AI总结 本文提出SDP,一种基于技能的扩散策略,通过整合可解释的技能学习与条件动作规划,提升机器人操作中技能一致性与性能。

Comments Accepted to AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21722 2025-12-29 cs.RO 50%

MAction-SocialNav: Multi-Action Socially Compliant Navigation via Reasoning-enhanced Prompt Tuning

MAction-SocialNav: 通过推理增强的提示调优实现多动作社交合规导航

Zishuo Wang, Xinyu Zhang, Zhuonan Liu, Tomohito Kawabata, Daeun Song, Xuesu Xiao, Ling Xiao

机构 * Graduate School of Information Science and Technology, Hokkaido University(信息科学与技术研究生院,北海道大学) Graduate School of Computer Science, George Mason University(计算机科学研究生院,乔治·梅森大学)

专题命中 GUI与屏幕智能体 :vision language model(abstract)

AI总结 MAction-SocialNav通过推理增强的提示调优实现多动作社交合规导航,提升决策质量与安全性,效率高于现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20940 2025-12-25 cs.RO 50%

ETP-R1: Evolving Topological Planning with Reinforcement Fine-tuning for Vision-Language Navigation in Continuous Environments

连续环境中的视觉-语言导航(VLN-CE)需要一个具身体验的智能体在连续环境中导航至目标,遵循自然语言指令。虽然当前基于图的方法通过将环境抽象为拓扑地图并简化动作空间到路径选择,提供了一种高效、结构化的方法,但它们在利用大规模数据和先进训练范式方面落后于基于大视觉-语言模型(LVLMs)的方法。在本文中,我们通过引入ETP-R1框架,将数据扩展和强化微调(RFT)范式应用于基于图的VLN-CE模型,以弥合这一差距。

Shuhao Ye, Sitong Mao, Yuxiang Cui, Xuan Yu, Shichao Zhai, Wen Chen, Shunbo Zhou, Rong Xiong, Yue Wang

机构 * Zhejiang University(浙江大学) Huawei Technologies Co., Ltd(华为技术有限公司) Zhejiang Humanoid Robot Innovation Center(浙江人形机器人创新中心)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

AI总结 ETP-R1通过数据扩展和强化微调范式,提升基于图的连续环境视觉-语言导航性能,实现新状态最先进的表现。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19083 2025-12-25 cs.RO 50%

CoDrone: Autonomous Drone Navigation Assisted by Edge and Cloud Foundation Models

CoDrone:由边缘和云基础模型辅助的自主无人机导航

Pengyu Chen, Tao Ouyang, Ke Luo, Weijie Hong, Xu Chen

专题命中 GUI与屏幕智能体 :VLM(abstract)

AI总结 CoDrone通过整合云-边-端协作计算框架和基础模型,提升无人机自主导航性能,实现更高效和精确的环境感知与动态适应。

Comments This paper is accepted by the IEEE Internet of Things Journal (IoT-J) for publication in the Special Issue on "Augmented Edge Sensing Intelligence for Low-Altitude IoT Systems"

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20166 2025-12-24 cs.RO 50%

LoLA: Long Horizon Latent Action Learning for General Robot Manipulation

LoLA:面向通用机器人操作的长周期隐式动作学习

Xiaofan Wang, Xingyu Gao, Jianlong Fu, Zuolei Li, Dean Fortier, Galen Mullins, Andrey Kolobov, Baining Guo

机构 * Institute of Microelectronics, Chinese Academy of Sciences(中国科学院微电子研究所) University of Chinese Academy of Sciences(中国科学院大学) Microsoft Research(微软研究院)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

AI总结 LoLA通过整合长期多视角观察和机器人本体感觉,实现长周期、语言引导的机器人操作任务,显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17079 2025-12-08 cs.RO 50%

H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation

H-GAR: 一种通过目标驱动的观察-动作细化的分层交互框架用于机器人操控

Yijie Zhu, Rui Shao, Ziyang Liu, Jie He, Jizhihui Liu, Jiuru Wang, Zitong Yu

专题命中 GUI与屏幕智能体 :grounding(abstract)

AI总结 H-GAR通过目标驱动的观察-动作细化分层框架提升机器人操控的准确性和一致性。

Comments Accepted to AAAI 2026 (Oral), Project Page: https://github.com/JiuTian-VL/H-GAR

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17401 2025-11-24 cs.RO cs.HC 50%

Feasibility of Embodied Dynamics Based Bayesian Learning for Continuous Pursuit Motion Control of Assistive Mobile Robots in the Built Environment

基于具身动力学的贝叶斯学习在辅助移动机器人连续追击运动控制中的可行性

Xiaoshan Zhou, Carol C. Menassa, Vineet R. Kamat

专题命中 GUI与屏幕智能体 :grounding(abstract)

AI总结 本文提出基于具身动力学的贝叶斯学习方法,有效提升辅助移动机器人在复杂环境中的连续追击运动控制性能。

Comments 37 pages, 9 figures, and 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12436 2025-11-18 cs.RO 50%

RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation

Xiaoshuai Hao, Yingbo Tang, Lingfeng Zhang, Yanbiao Ma, Yunfeng Diao, Ziyu Jia, Wenbo Ding, Hangjun Ye, Long Chen

机构 * Xiaomi EV(小米电动车) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学 Gallagher人工智能学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07315 2025-11-11 cs.CR 50%

JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework

Yuxuan Zhou, Yang Bai, Kuofeng Gao, Tao Dai, Shu-Tao Xia

专题命中 GUI与屏幕智能体 :VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02728 2025-11-07 cs.RO 50%

Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4

Lingfeng Zhang, Erjia Xiao, Yuchen Zhang, Haoxiang Fu, Ruibin Hu, Yanbiao Ma, Wenbo Ding, Long Chen, Hangjun Ye, Xiaoshuai Hao

机构 * Tsinghua University(清华大学) Xiaomi EV(小米电动车) Georgia Institute of Technology(佐治亚理工学院) National University of Singapore(新加坡国立大学) The Chinese University of Hong Kong(香港中文大学) Renmin University of China(中国人民大学)

专题命中 GUI与屏幕智能体 :VLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24109 2025-10-29 cs.RO 50%

PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI

Wenbin Ding, Jun Chen, Mingjia Chen, Fei Xie, Qi Mao, Philip Dames

机构 * School of Electrical and Automation Engineering, Nanjing Normal University(南京师范大学电气与自动化工程学院) Department of Mechanical Engineering, Temple University(Temple大学机械工程系)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06207 2025-10-16 cs.RO 50%

EmbodiedCoder: Parameterized Embodied Mobile Manipulation via Modern Coding Model

Zefu Lin, Rongxu Cui, Chen Hanning, Xiangyu Wang, Junjia Xu, Xiaojuan Jin, Chen Wenbo, Hui Zhou, Lue Fan, Wenling Li, Zhaoxiang Zhang

机构 * University of Chinese Academy of Sciences (UCAS)(中国科学院大学) Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所) New Laboratory of Pattern Recognition (NLPR)(模式识别新实验室) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS)(多模态人工智能系统国家重点实验室) Beihang University(北航) Chinese University of Hong Kong(香港大学)

专题命中 GUI与屏幕智能体 :grounding(abstract)

Comments Demo Page: https://embodiedcoder.github.io/EmbodiedCoder/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01185 2025-10-14 cs.RO 50%

HoMeR: Learning In-the-Wild Mobile Manipulation via Hybrid Imitation and Whole-Body Control

Priya Sundaresan, Rhea Malhotra, Phillip Miao, Jingyun Yang, Jimmy Wu, Hengyuan Hu, Rika Antonova, Francis Engelmann, Dorsa Sadigh, Jeannette Bohg

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13367 2025-10-09 cs.RO 50%

Uncertainty-Informed Active Perception for Open Vocabulary Object Goal Navigation

Utkarsh Bajpai, Julius Rückin, Cyrill Stachniss, Marija Popović

机构 * Center for Robotics, University of Bonn(波恩大学机器人中心) MAVLab, TU Delft(代尔夫特理工大学MAVLab) Lamarr Institute for Machine Learning and Artificial Intelligence(机器学习与人工智能拉马尔研究所)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

Comments 7 pages, 3 figures

Journal ref Proceedings of the 2025 European Conference on Mobile Robots (ECMR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00536 2025-10-02 cs.CL 50%

GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness

Kung-Hsiang Huang, Haoyi Qiu, Yutong Dai, Caiming Xiong, Chien-Sheng Wu

机构 * Salesforce AI Research(Salesforce AI研究)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24768 2025-09-30 cs.RO 50%

IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks

Eric Hannus, Miika Malin, Tran Nguyen Le, Ville Kyrki

机构 * Intelligent Robotics Group at the Department of Electrical Engineering and Automation, School of Electrical Engineering, Aalto University(Aalto大学电气工程学院电气工程与自动化系智能机器人组) Biomimetics and Intelligent Systems Group at the Faculty of Information Technology and Electrical Engineering, University of Oulu(奥卢大学信息科技与电气工程学院仿生学与智能系统组) Section of Mechanical Technology at the Department of Engineering Technology and Didactics, Technical University of Denmark(丹麦技术大学工程技术与教学系机械技术部门)

专题命中 GUI与屏幕智能体 :vision language model(abstract)

Comments Under review for ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24387 2025-09-30 cs.RO 50%

AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation

Xin Ding, Jianyu Wei, Yifan Yang, Shiqi Jiang, Qianxi Zhang, Hao Wu, Fucheng Jia, Liang Mi, Yuxuan Yan, Weijun Wang, Yunxin Liu, Zhibo Chen, Ting Cao

机构 * University of Science and Technology of China(中国科学技术大学) Microsoft Research(微软研究院) Nanjing University(南京大学) Central South University(中南大学) Zhejiang University(浙江大学) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)

专题命中 GUI与屏幕智能体 :grounding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19391 2025-09-30 cs.RO 50%

LaVA-Man: Learning Visual Action Representations for Robot Manipulation

Chaoran Zhu, Hengyi Wang, Yik Lung Pang, Changjae Oh

机构 * Queen Mary University of London(伦敦女王学院) University College London(伦敦大学学院)

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22937 2025-09-29 cs.HC 50%

GamerAstra: Supporting 2D Non-Twitch Video Games for Blind and Low-Vision Players through a Multi-Agent Framework

Tianrun Qiu, Changxin Chen, Sizhe Cheng, Xuyang Liu, Xumeng Wang, Zhicong Lu, Yuxin Ma

专题命中 GUI与屏幕智能体 :vision-language model(abstract)

Comments 17 pages, 11 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏