机构
*
School of Mechanical and Electronic Engineering, Wuhan University of Technology(武汉理工大学机电工程学院)
;
Intelligent Transportation Systems Research Center, Wuhan University of Technology(武汉理工大学智能交通系统研究中心)
;
School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院)
vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models
vla.cpp:视觉-语言-动作模型的统一推理运行时
Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen, Thanh Q. Duong, Linh D. Le, Duy M. H. Nguyen, Vien A. Ngo, An T. Le
机构
*
VinRobotics
;
Center for AI Research, VinUniversity(VinUniversity 人工智能研究中心)
;
Intelligent Autonomous Systems, TU Darmstadt(达姆施塔特工业大学智能自主系统)
;
Max Planck Research School for Intelligent Systems(马克斯·普朗克智能系统研究学院)
;
University of Stuttgart(斯图加特大学)
;
German Research Center for Artificial Intelligence(德国人工智能研究中心)
SA-VLA: State-aware tokenizer for improving Vision-Language-Action Models' performance
SA-VLA: 状态感知分词器提升视觉-语言-动作模型性能
Tengyue Jiang, Chunpu Xu, Jiayue Kang, Yao Mu
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
East China University of Science and Technology(东华大学)
;
Hong Kong Polytechnic University(香港理工大学)
;
Xi’an University of Electronic Science and Technology(西安电子科技大学)
机构
*
School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院)
;
State Key Laboratory of Submarine Geoscience, School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院海底科学国家重点实验室)
G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models
G$^3$VLA:视觉-语言-动作模型的几何归纳偏置
Yue Peng, Yongzhe Zhao, Artur Habuda, Khuyen Pham, Yanheng Zhu, Tran Nguyen Le, Fares Abu-Dakka, Li Guo
机构
*
New York University Shanghai(上海纽约大学)
;
Technical University of Denmark(丹麦技术大学)
;
MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
New York University Abu Dhabi(纽约大学阿布扎比分校)
机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Beijing Zhongke Huiling Robot Technology Co.(北京中科创联机器人科技有限公司)
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization
DyGRO-VLA: 通过动态分组残差优化实现跨任务的视觉-语言-动作模型扩展
Sixu Lin, Yunpeng Qing, Litao Liu, Ming Zhou, Ruixing Jin, Xiaoyi Fan, Guiliang Liu
机构
*
School of Data Science, The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)数据科学学院)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Zhejiang University(浙江大学)
;
Rutgers University-New Brunswick(罗格斯大学新布朗斯维尔回声分校)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Jiangxing Intelligence Technology Inc.(江行智能科技有限公司)
AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models
AR-VLA:面向视觉-语言-动作模型的真自回归动作专家
Yutong Hu, Jan-Nico Zaech, Nikolay Nikolov, Yuanqi Yao, Sombit Dey, Giuliano Albanese, Renaud Detry, Luc Van Gool, Danda Paudel
机构
*
KU Leuven, Dept. Mechanical Engineering, Research unit Robotics, Automation and Mechatronics(库勒恩大学,机械工程系,机器人、自动化与机电一体化研究单位)
;
KU Leuven, Dept. Electrical Engineering, Research unit Processing Speech and Images(库勒恩大学,电气工程系,语音和图像处理研究单位)
机构
*
Sun Yat-sen University(中山大学)
;
South China University of Technology(华南理工大学)
;
Peng Cheng Laboratory(鹏城实验室)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Huawei Noah’s Ark Lab(华为诺亚实验室)