MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction
MVP-LAM:通过跨视角重建学习以动作为中心的潜在动作
Jung Min Lee, Dohyeok Lee, Seokhun Ju, Taehyun Cho, Jin Woo Koo, Li Zhao, Sangwoo Hong, Jungwoo Lee
机构
*
Seoul National University, Seoul, South Korea(首尔国立大学,首尔,韩国)
;
Konkuk University, Seoul, South Korea(韩国konkuk大学,首尔,韩国)
;
Microsoft Research Asia, Beijing, China(微软亚洲研究院,北京,中国)
;
HodooAI Labs, Seoul, South Korea(HodooAI实验室,首尔,韩国)
机构
*
Sun Yat-sen University(中山大学)
;
South China University of Technology(华南理工大学)
;
Peng Cheng Laboratory(鹏城实验室)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Huawei Noah’s Ark Lab(华为诺亚实验室)
Towards Backdoor-Based Ownership Verification for Vision-Language-Action Models
面向视觉-语言-动作模型的后门基于所有权验证
Ming Sun, Rui Wang, Xingrui Yu, Lihua Jing, Hangyu Du, Zhenglin Wan, Xu Pan, Ivor Tsang
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
A*STAR Institute of High Performance Computing (A*STAR IHPC)(新加坡A*STAR高性能计算研究所)
;
A*STAR Centre for Frontier AI Research (A*STAR CFAR)(新加坡A*STAR前沿人工智能研究中心)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)
;
College of Design and Engineering, Nanyang Technological University(南洋理工大学设计与工程学院)
;
Department of Computer Science, National University of Singapore(新加坡国立大学计算机科学系)
;
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing (LIESMARS), Wuhan University(武汉大学测绘遥感信息工程国家重点实验室(LIESMARS))
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
ForgeVLA:无需语言标注的联邦视觉-语言-动作学习
Yuhao Zhou, Yunpeng Zhu, Yang Zhou, Jindi Lyu, Jian Lan, Zhangyuan Wang, Dan Si, Thomas Seidl, Qing Ye, Jiancheng Lyu
机构
*
Sichuan University(四川大学)
;
Zhejiang University(浙江大学)
;
Ludwig-Maximilians-Universität München(慕尼黑路德维希-马克西米利安大学)
;
Lenovo Group Limited(联想集团有限公司)
;
Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education(教育部机器学习与工业智能工程研究中心)
机构
*
School of Computer Science, Peking University, Beijing, China(北京大学计算机科学学院)
;
School of Computer Science, China University of Geosciences (Wuhan), Wuhan, China(中国地质大学(武汉)计算机科学学院)
AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models
AugVLA-3D:基于深度的特征增强用于视觉-语言-动作模型
Zhifeng Rao, Wenlong Chen, Lei Xie, Xia Hua, Dongfu Yin, Zhen Tian, F. Richard Yu
机构
*
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳))
;
Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系)
;
School of Communication and Information Engineering, Shanghai University(上海大学通信与信息工程学院)
;
School of Information Technology, Carleton University(卡尔顿大学信息技术学院)
机构
*
Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong(香港中文大学机械与自动化工程系)
;
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学计算机科学与工程系)
;
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
Department of Mechanical and Automation Engineering, T Stone Robotics Institute, Shun Hing Institute of Advanced Engineering, Multi-Scale Medical Robotics Center, and Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(香港中文大学机械与自动化工程系、T Stone机器人研究所、Shun Hing先进工程研究所、多尺度医疗机器人中心以及医学智能与XR研究所)
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
SafeVLA: 通过约束学习实现视觉-语言-动作模型的安全对齐
Borong Zhang, Yuhao Zhang, Jiaming Ji, Yingshan Lei, Yishuai Cai, Josef Dai, Yuanpei Chen, Yaodong Yang
机构
*
Institute for Artificial Intelligence, Peking University(人工智能研究院,北京大学)
;
PKU-PsiBot Joint Lab(北京大学PsiBot联合实验室)
;
State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学)
;
Zhongguancun Academy(中关村学院)
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
统一扩散VLA:通过联合离散去噪扩散过程实现视觉-语言-动作模型
Jiayi Chen, Wenxuan Song, Pengxiang Ding, Ziyang Zhou, Han Zhao, Feilong Tang, Donglin Wang, Haoang Li
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Westlake University(西交大学)
;
Zhejiang University(浙江大学)
;
Monash University(墨尔本大学)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
PixelVLA:推进视觉-语言-动作模型中的像素级理解
Wenqi Liang, Gan Sun, Yao He, Jiahua Dong, Suyan Dai, Ivan Laptev, Salman Khan, Yang Cong
机构
*
University of Trento(特伦托大学)
;
School of Automation Science and Engineering, South China University of Technology(华南理工大学自动化科学与工程学院)
;
Mohamed bin Zayed University of Artificial Intelligence(马尔代夫人工智能大学)
;
Australian National University(澳大利亚国立大学)
Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition
制定你的策略!通过测试时的分布级组合改进扩散型或流型机器人策略
Jiahang Cao, Yize Huang, Hanzhong Guo, Rui Zhang, Mu Nan, Weijian Mai, Jiaxu Wang, Hao Cheng, Jingkai Sun, Gang Han, Wen Zhao, Qiang Zhang, Yijie Guo, Qihao Zheng, Chunfeng Song, Xiao Li, Ping Luo, Andrew F. Luo
机构
*
The University of Hong Kong(香港大学)
;
Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
;
Shanghai AI Lab(上海人工智能实验室)
;
Shanghai Jiaotong University(上海交通大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
MOTIF: Learning Action Motifs for Few-shot Cross-Embodiment Transfer
MOTIF: 为少样本跨躯体迁移学习动作动机
Heng Zhi, Wentao Tan, Lei Zhu, Fengling Li, Jingjing Li, Guoli Yang, Heng Tao Shen
机构
*
Tongji University(同济大学)
;
University of Technology Sydney(技术悉尼大学)
;
University of Electronic Science(电子科学大学)
;
Advanced Institute of Big Data(大数据高级研究所)
Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
基于成功与失败示范的分层视觉语言行动模型
Jeongeun Park, Jihwan Yoon, Byungwoo Jeon, Juhan Park, Jinwoo Shin, Namhoon Cho, Kyungjae Lee, Sangdoo Yun, Sungjoon Choi
机构
*
Department of Artificial Intelligence, Korea University(韩国大学人工智能系)
;
Kim Jaechul Graduate School of AI, KAIST(金在拙人工智能研究生院,韩国科学技术院)
;
Department of Aerospace Engineering, Seoul National University(首尔国立大学航空航天工程系)
;
Departmnet of Statistics, Korea University(韩国大学统计系)
;
NAVER AI Lab(NAVER人工智能实验室)