Yitao Jiang, Roy Xing, Luyang Zhao, Brian Plancher, Muhao Chen, Devin Balkcom
机构
*
Dept. of Computer Science, Dartmouth College(达特茅斯学院计算机科学系)
;
Dept. of Electrical and Computer Engineering, Clemson University(克莱姆森大学电气与计算机工程系)
;
Dept. of Mechanical and Aerospace Engineering, University of Houston(休斯顿大学机械与航空航天工程系)
CommentsThe results presented in this paper are preliminary. Please note that the experiments are currently ongoing, and the final data is subject to change upon the completion of the study. All ideas, results, methods, and any content herein are the sole property of the authors
机构
*
MoE Key Lab of Artificial Intelligence, Institute of AI, Shanghai Jiao Tong University(教育部人工智能重点实验室,上海交通大学人工智能研究院)
;
Central Research Institute, Huawei(华为中央研究院)
机构
*
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
Beijing Academy of Artificial Intelligence (BAAI)(北京智源人工智能研究院)
;
Beihang University(北京航空航天大学)
;
Eastern Institute of Technology, Ningbo(宁波东方理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Microsoft Research Asia (MSRA)(微软亚洲研究院)
专题命中
VLA模型
:vision-language-action(title,abstract);action model(title);VLA(abstract,abstract_cn);embodied foundation model(abstract)
Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models
基于延迟反馈的测试时扰动学习用于视觉-语言-动作模型
Zehua Zang, Xi Wang, Fuchun Sun, Xiao Xu, Lixiang Lium, Jiahuan Zhou, Jiangmeng Li
机构
*
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tsinghua University(清华大学)
;
Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)
;
National Defense University(国防大学)
A Survey on Efficient Vision-Language-Action Models
高效视觉-语言-动作模型的综述
Zhaoshu Yu, Bo Wang, Pengpeng Zeng, Haonan Zhang, Ji Zhang, Zheng Wang, Lianli Gao, Jingkuan Song, Nicu Sebe, Heng Tao Shen
机构
*
School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
;
School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机科学与人工智能学院)
;
School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)
;
Department of Information Engineering and Computer Science, University of Trento(特伦托大学信息工程与计算机科学系)
VLA-ReID: Video-Level Association for Re-Identification in Multi-Object Tracking with Highly Similar Objects
VLA-ReID:具有高度相似对象的多目标跟踪中用于重新识别的视频级关联
Yanrong Qin, Xiaoyan Cao, Yao Yao
机构
*
Glasgow College, University of Electronic Science and Technology of China(电子科技大学格拉斯哥学院)
;
Key Laboratory for Urban Habitat Environmental Science and Technology, School of Environment and Energy, Peking University Shenzhen Graduate School(北京大学深圳研究生院环境与能源学院城市人居环境科学与技术重点实验室)
;
School of Management Science and Real Estate, Chongqing University(重庆大学管理科学与房地产学院)