机构
*
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院(AIR))
;
Shanghai Jiao Tong University(上海交通大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
AIR Wuxi Innovation Center, Tsinghua University(清华大学AIR无锡创新中心)
;
The University of Adelaide(阿德莱德大学)
;
Wuhan University(武汉大学)
;
Southeast University(东南大学)
;
Beijing Jiaotong University(北京交通大学)
;
Fudan University(复旦大学)
;
Li Auto(理想汽车)
;
School of Information, Renmin University of China(中国人民大学信息学院)
Commentsv5: Fixed PDF rendering compatibility issue affecting Apple PDFKit (macOS Preview/iOS PDF viewer). No changes to technical content compared to v4
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Harbin Institute of Technology(哈尔滨工业大学)
;
ShanghaiTech University(上海科技大学)
;
Shanghai Institute of Technical Physics, CAS(中国科学院上海技术物理研究所)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
MMaDA-VLA: 基于统一多模态指令与生成的大型扩散视觉-语言-动作模型
Yang Liu, Pengxiang Ding, Tengyue Jiang, Xudong Wang, Wenxuan Song, Minghui Lin, Han Zhao, Hongyin Zhang, Zifeng Zhuang, Wei Zhao, Siteng Huang, Jinkui Shi, Donglin Wang
机构
*
Westlake University(西湖大学)
;
Zhejiang University(浙江大学)
;
East China University of Science and Technology(华东理工大学)
;
Huawei Celia Team(华为Celia团队)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
OpenHelix Robotics
机构
*
Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
;
Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)
;
Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所)
;
Institute for Al Industry Research, Tsinghua University(清华大学人工智能产业研究院)
;
Department of Computer Science, Tsinghua University(清华大学计算机科学系)