Teaching Tiny VLA Models Where to Look and How to Move
XS-VLA:将粗粒度空间蒸馏与潜在流匹配相结合用于轻量级机器人控制
Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng
机构
*
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院)
;
Wuxi Dexteroushands Robotic Technology Co.(无锡灵犀机器人技术有限公司)
机构
*
Institute for AI Industry Research (AIR)(人工智能产业研究院)
;
Department of Electronic Engineering(电子工程系)
;
University of Science and Technology of China(中国科学技术大学)
机构
*
State Key Lab of Processors, Institute of Computing Technology, CAS(处理器国家重点实验室,计算技术研究所,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Nanjing University(南京大学)
;
Dexmal
机构
*
School of Computer Science, Peking University, Beijing, China(北京大学计算机科学学院)
;
School of Computer Science, South China University of Technology, Guangzhou, China(华南理工大学计算机科学学院)
;
School of Artificial Intelligence, Beijing Normal University, Beijing, China(北京师范大学人工智能学院)
;
School of EECS, Peking University, Beijing, China(北京大学电子工程与科学学院)
机构
*
Nanyang Technological University(南洋理工大学)
;
VinUniversity(文大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Tsinghua University(清华大学)
;
South China University of Technology(华南理工大学)
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models
动作QFormer:视觉-语言-动作模型中动作监督下的结构化表示塑造
Yufeng Ji, Wenhao Tang, Haoyi Niu, Koushil Sreenath, Yi Wu, Zhongyu Li
机构
*
Shanghai Qizhi Institute(上海期智研究院)
;
The Chinese University of Hong Kong(香港中文大学)
;
Hong Kong Embodied AI Lab(香港具身人工智能实验室)
;
Tsinghua University(清华大学)
;
University of California, Berkeley(加州大学伯克利分校)
LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries
LangForce: 通过潜在动作查询对视觉语言动作模型进行贝叶斯分解
Shijie Lian, Bin Yu, Xiaopeng Lin, Laurence T. Yang, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Cong Huang, Kai Chen
机构
*
Huazhong University of Science and Technology(华中科技大学)
;
Beijing Zhongguancun Academy(北京中关村学院)
;
Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Zhengzhou University(郑州大学)
;
Beihang University(北航)
;
East China Normal University(东华大学)
;
DeepCybot Co., Ltd.(DeepCybot有限公司)
专题命中
VLA模型
:VLA(summary_cn,abstract);vision language action(title);action model(title);vision-language-action(abstract)
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机科学学院多媒体信息处理技术国家重点实验室)
;
AI 2 Robotics(人工智能与机器人研究所)
;
Sun Yat-sen University(中山大学)
;
Beihang University(北京航空航天大学)