Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation
通过执行反馈强化学习训练高阶调度器以实现长周期GUI自动化
Zehao Deng, Tianjie Ju, Zheng Wu, Zhuosheng Zhang, Gongshen Liu
机构
*
School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)
;
School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)
LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
LaViRA: 语言-视觉-机器人动作翻译用于连续环境中的零样本视觉语言导航
Hongyu Ding, Ziming Xu, Yudong Fang, You Wu, Zixuan Chen, Jieqi Shi, Jing Huo, Yifan Zhang, Yang Gao
机构
*
School of Computer Science, Nanjing University(南京大学计算机科学学院)
;
School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
GUI与屏幕智能体
:grounding(abstract);multimodal large language model(abstract)
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Tsinghua University(清华大学)
;
University of Adelaide(阿德莱德大学)
;
Wuhan University(武汉大学)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Beijing Jiaotong University(北京交通大学)
;
AIR Wuxi Innovation Center, Tsinghua University(清华大学无锡创新中心)
;
Lenovo(联想集团)
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract);分类 cs.CV、cs.AI