机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Dexmal
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Zhongguancun Academy(中关村学院)
;
Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Tsinghua University(清华大学)
;
Beijing Innovation Center of Humanoid Robotics Co., Ltd.(北京人形机器人创新中心有限公司)
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学多媒体信息处理国家重点实验室,计算机学院)
A Gesture-Based Visual Learning Model for Acoustophoretic Interactions using a Swarm of AcoustoBots
基于手势的视觉学习模型用于声学聚沉交互的声学机器人群
Alex Lin, Lei Gao, Narsimlu Kemsaram, Sriram Subramanian
机构
*
Department of Computer Science University College London London United Kingdom(计算机科学系伦敦大学学院伦敦英国)
;
Department of Artificial Intelligence University of Malaya Kuala Lumpur Malaysia(人工智能系马来大学吉隆坡马来西亚)
;
University College London(伦敦大学学院)
;
University of Malaya(马来大学)
CommentsThis paper has been accepted for publication in the Proceedings of the 2026 4th International Conference on Robotics, Control and Vision Engineering (RCVE 2026)
Comments5 pages, 2 figures, to be published in Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA '26), 6 pages appendix
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
CUA-Suite:大规模人工标注的视频演示用于计算机使用代理
Xiangru Jian, Shravan Nayak, Kevin Qinghong Lin, Aarash Feizi, Kaixin Li, Patrice Bechard, Spandana Gella, Sai Rajeswar
机构
*
ServiceNow
;
University of Waterloo(多伦多大学)
;
Mila
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
University of Oxford(牛津大学)
;
National University of Singapore(新加坡国立大学)
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models
ProbeFlow: 无需训练的自适应流匹配用于视觉-语言-动作模型
Zhou Fang, Jiaqi Wang, Yi Zhou, Qiongfeng Shi
机构
*
School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院,中国)
;
Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Ministry of Education, China(新一代人工智能技术及其交叉应用国家重点实验室,中华人民共和国教育部,中国)
;
School of Electronic Science & Engineering, Southeast University, China(东南大学电子科学与工程学院,中国)
OnFly: Onboard Zero-Shot Aerial Vision-Language Navigation toward Safety and Efficiency
OnFly:面向安全与效率的机载零样本空中视觉语言导航
Guiyong Zheng, Yueting Ban, Mingjie Zhang, Juepeng Zheng, Boyu Zhou
机构
*
School of Artificial Intelligence, Sun Yat-Sen University(中山大学人工智能学院)
;
Southern University of Science and Technology(南方科技大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
LaViRA: 语言-视觉-机器人动作翻译用于连续环境中的零样本视觉语言导航
Hongyu Ding, Ziming Xu, Yudong Fang, You Wu, Zixuan Chen, Jieqi Shi, Jing Huo, Yifan Zhang, Yang Gao
机构
*
School of Computer Science, Nanjing University(南京大学计算机科学学院)
;
School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
GUI与屏幕智能体
:grounding(abstract);multimodal large language model(abstract)
机构
*
Advanced Research Lab, Samsung R&D Institute China-Beijing (SRCB)(三星研发中心中国北京先进研究实验室)
;
Samsung AI Center, DS Division(三星人工智能中心,DS部门)
;
Hanyang University ERICA(翰洋大学ERICA)