GUITrans2Act: Understanding User Operational Behaviors from Mobile GUI Interactions with Vision-Language Models
Teach-and-Repeat: 从移动屏幕演示中准确提取操作知识以赋能GUI智能体
机构 * Honor Device Co., Ltd(荣耀终端有限公司) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 GUI与屏幕智能体 :VLM(summary_cn,abstract);vision-language model(title,abstract);分类 cs.AI
AI总结 提出Teach VLM模型,通过从演示视频中提取关键帧生成操作知识,并构建数据飞轮解决训练数据稀缺问题;在基准测试中达到最优性能,并提升下游智能体的任务成功率。
Comments 20 pages, 9 figures. Yudong Zhang and Lei Hu contributed equally to this work. Zuojian Wang, and Zhilin Gao are corresponding authors