机构
*
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China(新型软件技术国家实验室,南京大学,南京,中国)
;
School of Artificial Intelligence, Nanjing University, Nanjing, China(人工智能学院,南京大学,南京,中国)
;
Mila - Quebec AI Institute(魁北克AI研究所)
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
动作草稿与验证:一种用于视觉-语言-动作模型的自验证框架
Chen Zhao, Zhuoran Wang, Haoyang Li, Shifeng Bao, Guanlin Li, Youhe Feng, Yang Li, Jie Tang, Jing Zhang
机构
*
Key Laboratory of Data Engineering(数据工程与知识工程重点实验室)
;
Tsinghua University(清华大学)
;
Beijing University of Posts(北京邮电大学)
;
School of Information, Renmin University of China(中国人民大学信息学院)
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(信息多媒体国家重点实验室,计算机科学学院,北京大学)
;
School of Software, Beihang University(软件学院,北航)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
机构
*
Institute of Trustworthy Embodied AI, Fudan University(可信具身AI研究院,复旦大学)
;
Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身AI重点实验室)
;
City University of Hong Kong(香港城市大学)
机构
*
Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences University of Chinese Academy of Sciences(人工智能安全重点实验室,计算技术研究所,中国科学院,中国科学院大学)
机构
*
Department of Civil Engineering, Stony Brook University(土木工程系,石溪大学)
;
Department of Civil and Environmental Engineering, Virginia Tech(土木与环境工程系,弗吉尼亚理工大学)
;
Department of Electrical and Computer Engineering, Virginia Tech(电气与计算机工程系,弗吉尼亚理工大学)
机构
*
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
LimX Dynamic(LimX动态)
;
Nanjing University(南京大学)
;
Nanyang Technological University(南洋理工大学)
;
Zhejiang University(浙江大学)
专题命中
VLA模型
:vision language action(title);action model(title);vision-language-action(abstract);VLA(abstract)
Habilis-$β$: A Fast-Motion and Long-Lasting On-Device Vision-Language-Action Model
Habilis-β:一种快速运动且持续运行的设备端视觉-语言-动作模型
Tommoro Robotics, :, Jesoon Kang, Taegeon Park, Jisu An, Soo Min Kimm, Jaejoon Kim, Jinu Pahk, Byungju Kim, Junseok Lee, Namheon Baek, Sungwan Ha, Hojun Baek, Eduardo Ayerve Cruz, Wontae Kim, Junghyeon Choi, Yousuk Lee, Joonmo Han, Sunghyun Cho, Sunghyun Kwon, Soyoung Lee, Jun Ki Lee, Seung-Joon Yi, Byoung-Tak Zhang, Theo Taeyeong Kim
Information-Theoretic Graph Fusion with Vision-Language-Action Model for Policy Reasoning and Dual Robotic Control
信息论图融合:基于视觉-语言-动作模型的政策推理与双臂机器人控制
Shunlei Li, Longsen Gao, Jin Wang, Chang Che, Xi Xiao, Jiuwen Cao, Yingbai Hu, Hamid Reza Karimi
机构
*
Electrical and Computer Engineering Department, University of New Mexico, Albuquerque, United States, 87106(电气与计算机工程系,新墨西哥大学,阿尔伯克基,美国,87106)
;
Dynamic Robot Systems Group, Oxford Robotics Institute, University of Oxford, United Kingdom, OX26NN(动态机器人系统组,牛津机器人研究所,牛津大学,英国,OX26NN)
;
Mechanical and Aerospace Engineering Department, The George Washington University, DC, United States, 22202(机械与航空航天工程系,乔治华盛顿大学,华盛顿特区,美国,22202)
;
Department of Computer Science, University of Alabama at Birmingham, Alabama, United States, 35294(计算机科学系,阿拉巴马大学伯明翰分校,阿拉巴马,美国,35294)
;
The School of Computation, Information and Technology, Technical University of Munich, Germany, 85748(计算、信息与技术学院,慕尼黑技术大学,德国,85748)
;
Department of Mechanical Engineering, Politecnico di Milano, Milan, Italy, 20156(机械工程系,米兰理工学院,米兰,意大利,20156)
机构
*
AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University(自动化实验室,人工智能学院,上海交通大学)
;
Anyverse Dynamics
;
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Terminal Technology Department, Alipay, Ant Group(终端技术部,蚂蚁集团)
Vision Language Action Models in Robotic Manipulation: A Systematic Review
视觉语言动作模型在机器人操作中的应用:系统综述
Muhayy Ud Din, Waseem Akram, Lyes Saad Saoud, Jan Rosell, Irfan Hussain
机构
*
Khalifa University Center for Autonomous Robotic Systems (KUCARS), Khalifa University, United Arab Emirates(卡利法大学自主机器人系统中心(KUCARS)、卡利法大学、阿拉伯联合酋长国)
;
Institute of Industrial and Control Engineering (IOC), Universitat Politecnica de Catalunya, Spain(工业与控制工程研究所(IOC)、巴塞罗那技术大学、西班牙)
专题命中
VLA模型
:vision language action(title,abstract);action model(title);VLA(abstract);分类 cs.RO、cs.CV
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
引导视觉-语言-动作模型作为反探索:一种测试时间缩放方法
Siyuan Yang, Yang Zhang, Haoran He, Ling Pan, Xiu Li, Chenjia Bai, Xuelong Li
机构
*
Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院)
;
University of Science and Technology of China(中国科学技术大学)
;
Tsinghua University(清华大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)