Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
Video2GUI:合成大规模交互轨迹用于通用GUI代理预训练
Weimin Xiong, Shuhao Gu, Bowen Ye, Zihao Yue, Lei Li, Feifan Song, Sujian Li, Hao Tian
机构
*
National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(国家多媒体信息处理重点实验室,计算机科学学院,北京大学)
;
The University of Hong Kong(香港大学)
;
Renmin University of China(中国人民大学)
EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
EARL:一种统一的分析引导强化学习框架,用于第一人称交互推理与像素定位
Yuejiao Su, Xinshen Zhang, Zhen Ye, Lei Yao, Lap-Pui Chau, Yi Wang
机构
*
Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong SAR(香港理工大学电子与电气工程系)
;
Division of Emerging Interdisciplinary Areas (EMIA), The Hong Kong University of Science and Technology, Hong Kong SAR(香港理工大学新兴跨学科领域研究中心)
Agentic AI Ecosystems in Higher Education: A Perspective on AI Agents to Emerging Inclusive, Agentic Multi-Agent AI Framework for Learning, Teaching and Institutional Intelligence
机构
*
Fudan University, China(复旦大学)
;
University of Queensland, Australia(昆士兰大学)
;
Shanghai Academy of AI for Science, China(上海人工智能科学研究院)
;
Huashan Hospital, National Center for Neurological Disorders, Fudan University, China(华山医院,国家神经系统疾病中心,复旦大学)
;
Bioinformatics Institute (BII), Agency for Science, Technology and Research (A*STAR), Singapore(生物信息研究所(BII),科技研究局(A*STAR),新加坡)
机构
*
City University of Hong Kong(香港城市大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
CAIR, HKISI, Chinese Academy of Sciences(中国科学院计算智能研究所)
;
UESTC(电子科技大学)
;
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
机构
*
Department of Electrical and Computer Engineering, Rice University(理海大学电气与计算机工程系)
;
Center for Cardiovascular Computational and Precision Health, Department of Cardiology, DeBakey Heart and Vascular Center, Houston Methodist(休斯顿方法主义医疗中心心血管计算与精准健康中心、心内科部门、德贝基心脏和血管中心)
Characterizing the visual representation of objects from the child's view
从儿童视角表征物体的可视化分析
Jane Yang, Tarun Sepuri, Alvin Wei Ming Tan, Khai Loong Aw, Michael C. Frank, Bria Long
机构
*
Department of Psychology, University of California San Diego(加州大学圣地亚哥分校心理学系)
;
Department of Psychology, Stanford University(斯坦福大学心理学系)
;
Department of Computer Science, Stanford University(斯坦福大学计算机科学系)
Let Robots Feel Your Touch: Visuo-Tactile Cortical Alignment for Embodied Mirror Resonance
让机器人感受你的触摸:用于具身镜像共振的视觉-触觉皮层对齐
Tianfang Zhu, Ning An, Rui Wang, Jiasi Gao, Qingming Luo, Anan Li, Guyue Zhou
机构
*
Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院)
;
Key Laboratory of Biomedical Engineering of Hainan Province, School of Biomedical Engineering, Hainan University(海南省生物医学工程重点实验室,海南大学生物医学工程学院)
;
School of New Media Art and Design, Beihang University(北航艺术与设计学院)
;
MoE Key Laboratory for Biomedical Photonics, Wuhan National Laboratory for Optoelectronics, Huazhong University of Science and Technology(教育部生物医学光子学重点实验室,武汉光电研究所,华中科技大学)
机构
*
MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能大模型关键实验室,人工智能研究院,上海交通大学)
;
vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司)
A Mutual Information Lower Bound for Multimodal Regression Active Learning
多模态回归主动学习的互信息下界
Leonardo Ferreira Guilhoto, Akshat Kaushal, Paris Perdikaris
机构
*
Graduate Group on Applied Mathematics & Computational Science(应用数学与计算科学研究生组)
;
Department of Computer and Information Science(计算机与信息科学系)
;
University of Pennsylvania(宾夕法尼亚大学)
;
Department of Mechanical Engineering and Applied Mechanics(机械工程与应用力学系)