XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
XR-1:通过学习统一的视觉-运动表示实现多功能的视觉-语言-动作模型
Shichao Fan, Kun Wu, Zhengping Che, Xinhua Wang, Di Wu, Fei Liao, Ning Liu, Yixue Zhang, Zhen Zhao, Zhiyuan Xu, Meng Li, Qingjie Liu, Shanghang Zhang, Min Wan, Jian Tang
机构
*
Beijing Innovation Center of Humanoid Robotics, Beijing, China(北京人形机器人创新中心,北京,中国)
;
School of Mechanical Engineering and Automation, Beihang University, Beijing, China(北京航空航天大学机械工程及自动化学院,北京,中国)
;
State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,SCSE,北京航空航天大学,北京,中国)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(多媒体信息处理国家重点实验室,计算机科学学院,北京大学,北京,中国)
机构
*
School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络空间安全学院)
;
College of Computer Science, Chongqing University(重庆大学计算机科学学院)
;
School of Software and engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)
Mao-Lin Luo, Zhe-Xu Wang, Zi-Hao Zhou, Bo Ye, Jian Zhao, Min-Ling Zhang, Tong Wei
机构
*
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education(教育部计算机网络和信息集成重点实验室(东南大学))
;
Zhongguancun Academy(中关村科学城公司)
;
Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)
机构
*
Peking University(北京大学)
;
Kling Team, Kuaishou Technology(快手团队)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Sun Yat-sen University(中山大学)
机构
*
Zhejiang University(浙江大学)
;
South China University of Technology(华南理工大学)
;
Central China Normal University(华中师范大学)
;
Lenovo Group Limited(联想集团有限公司)
Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
零努力图像到音乐生成:一种可解释的基于RAG的视觉语言模型方法
Zijian Zhao, Dian Jin, Zijing Zhou
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
The University of Hong Kong(香港大学)
;
The Hong Kong University of Science(香港科学大学)
;
The Hong Kong Polytechnic University Hong Kong China(香港理工大学香港中国)
;
The University of Hong Kong Hong Kong China(香港大学香港中国)
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
长期体育视频中的时间组成推理方法
Siyu Cao, Lu Zhang, Ruizhe Zeng, Zhi-yong Liu
机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院)
;
AnyverseDynamics