机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Beijing Zhongke Huiling Robot Technology Co.(北京中科创联机器人科技有限公司)
机构
*
Tsinghua University(清华大学)
;
Zhongguancun Academy(中关村学院)
;
Infinigence AI
;
Peking University(北京大学)
;
UC Berkeley(加州大学伯克利分校)
;
Harbin Institute of Technology(哈尔滨工程学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
CommentsAccepted to RSS 2026. This is the technical report of the RLinf Team, focusing on the algorithm side. For the system-level design, please refer to arXiv:2509.15965. The open-sourced code link: https://github.com/RLinf/RLinf
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
MMaDA-VLA: 基于统一多模态指令与生成的大型扩散视觉-语言-动作模型
Yang Liu, Pengxiang Ding, Tengyue Jiang, Xudong Wang, Wenxuan Song, Minghui Lin, Han Zhao, Hongyin Zhang, Zifeng Zhuang, Wei Zhao, Siteng Huang, Jinkui Shi, Donglin Wang
机构
*
Westlake University(西湖大学)
;
Zhejiang University(浙江大学)
;
East China University of Science and Technology(华东理工大学)
;
Huawei Celia Team(华为Celia团队)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
OpenHelix Robotics
机构
*
McGill University(麦吉尔大学)
;
Université de Montréal(蒙特利尔大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Mila – Quebec AI Institute(魁北克AI研究所)
Teaching Tiny VLA Models Where to Look and How to Move
XS-VLA:将粗粒度空间蒸馏与潜在流匹配相结合用于轻量级机器人控制
Iok Tong Lei, Ying Jie Yap, Wei Huang, Qingchen Xie, Qianzhi Li, Yujie Zhang, Xiaolong Liu, Zhidong Deng
机构
*
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院)
;
Wuxi Dexteroushands Robotic Technology Co.(无锡灵犀机器人技术有限公司)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
引导、思考、行动:面向视觉-语言-动作模型的交互式具身推理
Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang, Tianming Zhang, Xiaoke Jiang, Chuanxiu Liu, Jie Liu, Lei Zhang
机构
*
Futian Laboratory(福田实验室)
;
Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院)
;
International Digital Economy Academy (IDEA)(国际数字经济学院(IDEA))
;
School of Robotics, Hunan University(湖南大学机器人学院)
;
South China University of Technology(华南理工大学)
;
Visincept(Visincept公司)
;
National Key Laboratory of Smart Farm Technologies and Systems(智能农业技术与系统国家重点实验室)
UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models
UAOR: 面向视觉-语言-动作模型的不确定性感知观测重注入
Jiabing Yang, Yixiang Chen, Yuan Xu, Peiyan Li, Zichen Wen, Bowen Fang, Tao Yu, Xiangnan Wu, Qisen Ma, Kai Wang, Ziheng He, Yingda Li, Zhengbo Zhang, Jing Liu, Nianfeng Liu, Yan Huang, Liang Wang
机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新技术实验室)
;
Shanghai Jiao Tong University(上海交通大学)
;
FiveAges(五代)