CAC-VLA: Context-Gated Action Conditioning for Vision-Language-Action Models
CAC-VLA:用于视觉-语言-动作模型的上下文门控动作条件调节
Yifu Xiong, Wenhao Yu, Jiaxuan Lin, Bojun Zou, Jiahao Li, Lu Zhang, Yanyong Zhang, Jianmin Ji
机构
*
University of Science and Technology of China (USTC)(中国科学技术大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机科学学院多媒体信息处理技术国家重点实验室)
;
AI 2 Robotics(人工智能与机器人研究所)
;
Sun Yat-sen University(中山大学)
;
Beihang University(北京航空航天大学)
机构
*
The Hong Kong Polytechnic University(香港理工大学)
;
Eastern Institute of Technology(东区技术研究所)
;
Great Bay University(大湾大学)
;
Northeast Forestry University(东北林业大学)
;
National University of Singapore(新加坡国立大学)
机构
*
MoE Key Lab of Artificial Intelligence, Institute of AI, Shanghai Jiao Tong University(教育部人工智能重点实验室,上海交通大学人工智能研究院)
;
Central Research Institute, Huawei(华为中央研究院)
AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning
AnchorVLA:为视觉-语言-动作规划连接离散决策与连续轨迹
Qi Liu, Yabei Li, Hongsong Wang, Heng Zhang, Lei He
机构
*
School of Vehicle and Mobility, Tsinghua University(清华大学车辆与运载学院)
;
State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University(清华大学智能绿色车辆与交通国家重点实验室)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Meituan Inc.(美团公司)
;
Dongfeng Motor Corporation(东风汽车公司)
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
Zhaohui Du, Zhe Wang, Dongzhan Zhou, Minting Pan, Hongmei Fei, Xiwen Cao, Ting Xiao, Qi Wang, Huanbo Jin, Jiaming Gu, Quan Lu, Zhe Liu
机构
*
Key Laboratory of Smart Manufacturing in Energy Chemical Process Ministry of Education, East China University of Science and Technology, Shanghai, CN(能源化工过程智能制造教育部重点实验室,东华大学,上海,中国)
;
Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, CN(东华大学计算机科学与工程系,上海,中国)
;
Department of Laboratory Medicine, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, CN(复旦大学附属瑞金医院检验医学科,上海,中国)
;
School of Information Science and Technology, Shihezi University, Shihezi, CN(石河子大学信息科学与技术学院,石河子,中国)
机构
*
Nanjing University(南京大学)
;
Australian National University(澳大利亚国立大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Nvidia(英伟达)
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
视觉语言动作模型所言即所指?论忠实性在具身推理中的作用
Matthew Foutter, Matteo Cercola, Lena Wild, Yunshan Wang, Michelle Li, Daniele Gammelli, Marco Pavone
机构
*
Stanford University(斯坦福大学)
;
Politecnico di Milano(米兰理工大学)
;
KTH Royal Institute of Technology(皇家理工学院)
;
Italian Institute of Artificial Intelligence (AI4I)(意大利人工智能研究所)
;
NVIDIA Research(英伟达研究院)
机构
*
Department of Computer Science, University College London(计算机科学系,伦敦大学学院)
;
Department of Mechanical Engineering, University College London(机械工程系,伦敦大学学院)