Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation
Goal-VLA: 图像生成视觉语言模型作为对象中心世界模型,赋能零样本机器人操作
Haonan Chen, Jingxiang Guo, Bangjun Wang, Tianrui Zhang, Xuchuan Huang, Boren Zheng, Yiwen Hou, Chenrui Tie, Jiajun Deng, Lin Shao
机构
*
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
;
The HKU Musketeers Foundation Institute of Data Science, The University of Hong Kong(香港大学数据科学研究院)
;
Yuanpei College, Peking University(北京大学元培学院)
;
Department of Automation, Tsinghua University(清华大学自动化系)
机构
*
Research Center for Intelligent Robotics, School of Astronautics, Northwestern Polytechnical University(西北工业大学航天学院智能机器人研究中心)
;
School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院)
;
School of Aeronautics and Astronautics, Sichuan University(四川大学空天科学与工程学院)
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Shenzhen Loop Area Institute(深圳河套学院)
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
D-Robotics(地平线机器人)
;
Southern University of Science and Technology(南方科技大学)
;
Xspark AI(星火科技)
;
The University of Hong Kong(香港大学)
;
Wuhan University(武汉大学)
;
Xiamen University Malaysia(厦门大学马来西亚分校)
;
Northeastern University(东北大学)
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
AffordGrasp:跨模态扩散用于感知意识抓取合成
Xiaofei Wu, Yi Zhang, Yumeng Liu, Yuexin Ma, Yujiao Shi, Xuming He
机构
*
ShanghaiTech University(上海科技大学)
;
Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程技术研究中心)
;
University of Science and Technology of China(中国科学技术大学)