Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models
面向大语言模型时代的空中机器人视觉-语言导航
Xingyu Xia, Lekai Zhou, Yujie Tang, Xiaozhou Zhu, Hai Zhu, Wen Yao
机构
*
Defense Innovation Institute, Chinese Academy of Military Sciences(军事科学院国防创新研究院)
;
Intelligent Game and Decision Laboratory(智能博弈与决策实验室)
;
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院)
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
ST4VLA:基于空间引导的视觉-语言-动作模型训练
Jinhui Ye, Fangjing Wang, Ning Gao, Junqiu Yu, Yangkun Zhu, Bin Wang, Jinyu Zhang, Weiyang Jin, Yanwei Fu, Feng Zheng, Yilun Chen, Jiangmiao Pang
机构
*
Shanghai AI Laboratory(上海人工智能实验室)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Southern University of Science and Technology(南方科技大学)
;
Fudan University(复旦大学)
机构
*
Beijing Institute of Technology(北京理工大学)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)
;
DataCanvas
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Shenzhen MSU-BIT University(深圳MSU-BIT大学)
SCLARO: A Dataset for Grounded Scenario-Level Scene Understanding and ScenarioCLIP for Benchmarking
SCLARO:面向场景级场景理解的数据集及用于基准测试的ScenarioCLIP
Advik Sinha, Saurabh Atreya, Aashutosh A, Sk Aziz Ali, Abhijit Das
机构
*
Machine Intelligence Group, Department of CS&IS, Birla Institute of Technology and Science, Pilani – Hyderabad Campus(人工智能组,计算机科学与信息系,比拉理工学院和科学学院,比拉-海得拉巴校区)
专题命中
GUI与屏幕智能体
:visual language model(title);分类 cs.CV
Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots
多模态大语言模型驱动机器人中的躯体自我识别
Iñaki Dellibarda Varela, Pablo Romero-Sorozabal, Diego Torricelli, Gabriel Delgado-Oleas, Jose Ignacio Serrano, Maria Dolores del Castillo Sobrino, Eduardo Rocon, Manuel Cebrian
机构
*
Center for Automation and Robotics(自动化与机器人中心)
;
University of Azuay(阿祖亚大学)
专题命中
GUI与屏幕智能体
:multimodal large language model(title);分类 cs.AI
ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
ActionFlow:面向边缘设备的视觉语言模型流水线动作加速
Yuntao Dai, Hang Gu, Teng Wang, Qianyu Cheng, Yifei Zheng, Zhiyong Qiu, Lei Gong, Wenqi Lou, Xuehai Zhou
机构
*
School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学)
;
Suzhou Institute for Advanced Research, University of Science and Technology of China(苏州市先进研究院,中国科学技术大学)
;
IEIT SYSTEMS Co., Ltd.(IEIT SYSTEMS公司)
专题命中
GUI与屏幕智能体
:vision language model(title);分类 cs.AI
机构
*
Department of Mechanical Engineering, Stanford University(斯坦福大学机械工程系)
;
Department of Aeronautics and Astronautics, Stanford University(斯坦福大学航空航天工程系)