Can Vision-Language Models Understand Construction Workers? An Exploratory Study
视觉-语言模型能理解建筑工人吗?一项探索性研究
Hieu Bui, Nathaniel E. Chodosh, Arash Tavakoli
机构
*
Department of Electrical and Computer Engineering(电气与计算机工程系)
;
Villanova University(维拉诺瓦大学)
;
Department of Computing Sciences(计算科学系)
;
Department of Civil and Environmental Engineering(土木与环境工程系)
Navigation with VLM framework: Towards Going to Any Language
Zecheng Yin, Chonghao Cheng, and Yao Guo, Zhen Li
机构
*
Future Network of Intelligence Institute(Shenzhen)(智能未来网络研究所(深圳))
;
University of Technology Sydney(悉尼技术大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Chinese University of Hongkong(Shenzhen)(香港中文大学(深圳))
专题命中
GUI与屏幕智能体
:VLM(title,abstract);vision language model(abstract);分类 cs.CV、cs.AI
机构
*
Columbia University(哥伦比亚大学)
;
Toyota Research Institute(丰田研究院)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Tsinghua University(清华大学)
CommentsAccepted to Robotics: Science and Systems (RSS) 2025. The first three authors contributed equally. Project Page: https://robopil.github.io/code-diffuser/
机构
*
Renmin University of China(中国人民大学)
;
Tsinghua University(清华大学)
;
Xiamen University(厦门大学)
;
BUPT(北邮)
;
ModelBest
;
Chinese Academy of Sciences(中国科学院)
;
National University of Singapore(新加坡国立大学)
;
Shanghai Qi Zhi Institute(上海启智研究院)
专题命中
GUI与屏幕智能体
:vision language model(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
机构
*
College of Computer Science, Chongqing University(重庆大学计算机学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)
;
School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)
专题命中
GUI与屏幕智能体
:vision-language model(title);vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI
Enhancing Screen Time Identification in Children with a Multi-View Vision Language Model and Screen Time Tracker
Xinlong Hou, Sen Shen, Xueshen Li, Xinran Gao, Ziyi Huang, Steven J. Holiday, Matthew R. Cribbet, Susan W. White, Edward Sazonov, Yu Gan
机构
*
Department of Biomedical Engineering, Stevens Institute of Technology(生物医学工程系,史蒂文斯理工学院)
;
Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学)
;
Department of Electrical Engineering, Columbia University in the City of New York(电气工程系,哥伦比亚大学(纽约市))
;
Nokia Bell Labs(诺基亚贝尔实验室)
;
Department of Communication & Information Science, University of Alabama(传播与信息科学系,阿拉巴马大学)
;
Department of Psychology, University of Alabama(心理学系,阿拉巴马大学)
;
Department of Electrical and Computer Engineering, University of Alabama(电气与计算机工程系,阿拉巴马大学)
专题命中
GUI与屏幕智能体
:vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.AI