CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
CUA-Suite:大规模人工标注的视频演示用于计算机使用代理
Xiangru Jian, Shravan Nayak, Kevin Qinghong Lin, Aarash Feizi, Kaixin Li, Patrice Bechard, Spandana Gella, Sai Rajeswar
机构
*
ServiceNow
;
University of Waterloo(多伦多大学)
;
Mila
;
Université de Montréal(蒙特利尔大学)
;
McGill University(麦吉尔大学)
;
University of Oxford(牛津大学)
;
National University of Singapore(新加坡国立大学)
机构
*
Department of Computer Science and Technology, Tsinghua University, China(计算机科学与技术系,清华大学,中国)
;
College of Computer Science, Zhejiang University, China(浙江大学计算机科学学院,中国)
;
College of Computer Science, Beijing University of Posts and Telecommunications, China(北京邮电大学计算机科学学院,中国)
机构
*
Human-Robot Interfaces and Interaction Lab, Istituto Italiano di Tecnologia(人机交互实验室,意大利理工学院)
;
Ph.D. program of national interest in Robotics and Intelligent Machines (DRIM) and Università di Genova(机器人与智能机器国家级博士项目(DRIM)和热那亚大学)
;
Edwardson School of Industrial Engineering, Purdue University(工业工程学院,普渡大学)
机构
*
MoE Key Lab of BIPC, University of Science and Technology of China(北京信息科技大学MoE关键实验室,中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
;
Nanyang Technological University(南洋理工大学)
机构
*
Tongji University(同济大学)
;
Computer Network Information Center, CAS(计算机网络信息中心,中国科学院)
;
HIAS, University of Chinese Academy of Sciences(高等研究所,中国科学院大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
Learning To Guide Human Decision Makers With Vision-Language Models
通过视觉-语言模型引导人类决策者
Debodeep Banerjee, Stefano Teso, Burcu Sayin, Andrea Passerini
机构
*
DI, University of Pisa(比萨大学DI)
;
DISI, University of Trento(特伦托大学DISI)
;
CIMEC, University of Trento(特伦托大学CIMEC)
;
Rajib Pal Bankura Sammilani Medical College(拉吉布·帕尔班克尔医学学院)
机构
*
MoE Key Lab of BIPC, University of Science and Technology of China(脑信息处理实验室,中国科学技术大学)
;
Opus AI Research(Opus AI研究院)
;
Southeast University(东南大学)
;
Nanyang Technological University(南洋理工大学)
机构
*
Texas A&M University(德克萨斯A&M大学)
;
University of Minnesota(明尼苏达大学)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
Abaka AI(Abaka人工智能)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
机构
*
Shandong University(山东大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Donghua University(东华大学)
;
Shanghai Normal University(上海师范大学)
;
East China Normal University(华东师范大学)
;
Hefei University of Technology(合肥工业大学)
Vision-Language Models vs Human: Perceptual Image Quality Assessment
视觉-语言模型与人类:感知图像质量评估
Imran Mehmood, Imad Ali Shah, Ming Ronnier Luo, Brian Deegan
机构
*
School of Engineering, University of Galway(Galway大学工程学院)
;
State Key Laboratory of Extreme Photonics and Instrumentation, Zhejiang University(浙江大学极端光信息获取国家重点实验室)
专题命中
VLM训练与架构
:vision-language model(title);vision language model(abstract);VLM(abstract);分类 cs.CV
Will It Zero-Shot?: Predicting Zero-Shot Classification Performance For Arbitrary Queries
零样本预测是否可行?:预测任意查询的零样本分类性能
Kevin Robbins, Xiaotong Liu, Yu Wu, Le Sun, Grady McPeak, Abby Stylianou, Robert Pless
机构
*
Computer Science George Washington University Washington, DC, USA(计算机科学 华盛顿大学 华盛顿特区, 美国)
;
Computer Science Saint Louis University Saint Louis, USA(计算机科学 圣路易斯大学 圣路易斯, 美国)
机构
*
Sun Yat-sen University(中山大学)
;
Guangdong Key Laboratory of Big Data Analysis(广东大数据分析与处理重点实验室)
;
X-Era AI Lab(X-Era人工智能实验室)
;
Guangdong University of Technology(广东工业大学)
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
从面板到像素:从生物医学科学文献中进行缩放视觉-语言预训练
Kun Yuan, Min Woo Sun, Zhen Chen, Alejandro Lozano, Xiangteng He, Shi Li, Nassir Navab, Xiaoxiao Sun, Nicolas Padoy, Serena Yeung-Levy
机构
*
University of Strasbourg, CNRS, INSERM, ICube, UMR7357, Strasbourg, France(斯特拉斯堡大学、法国国家科学研究中心、法国国家医学研究院、ICube、UMR7357、斯特拉斯堡,法国)
;
Technical University of Munich(慕尼黑技术大学)
;
Stanford University(斯坦福大学)
;
DSAI, The Hong Kong Polytechnic University(香港理工大学DSAI实验室)
;
University of British Columbia(不列颠哥伦比亚大学)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
IHU Strasbourg, Strasbourg(斯特拉斯堡IHU医院,斯特拉斯堡)
;
Vector Institute for AI(人工智能向量研究所)
;
Yale University(耶鲁大学)
Cross-Modal Prototype Alignment and Mixing for Training-Free Few-Shot Classification
跨模态原型对齐与混合用于无训练少样本分类
Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bartłomiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer
机构
*
Department of Computer Science, Universitat Autònoma de Barcelona, Spain(巴塞罗那自治大学计算机科学系)
;
Computer Vision Center, Barcelona, Spain(巴塞罗那计算机视觉中心)
;
Media Integration and Communication Center, University of Florence, Italy(佛罗伦萨大学媒体整合与传播中心)
;
Bernoulli Institute, University of Groningen, the Netherlands(格罗宁根大学伯努利学院)
;
IDEAS Research Institute, Poland(波兰IDEAS研究所)
;
ESAT-PSI, KU Leuven, Belgium(比利时KU莱顿大学ESAT-PSI)