ArchGPT: Understanding the World's Architectures with Large Multimodal Models
Yuze Wang, Luo Yang, Junyi Wang, Yue Qi
机构
*
State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室)
;
School of Computer Science and Engineering(计算机科学与工程学院)
;
Beihang University(北京航空航天大学)
;
School of Computer Science and Technology(计算机科学与技术学院)
;
Shandong University(山东大学)
Long Video Understanding with Learnable Retrieval in Video-Language Models
Jiaqi Xu, Cuiling Lan, Wenxuan Xie, Xuejin Chen, Yan Lu
机构
*
School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学)
;
Microsoft Research Asia(微软亚洲研究院)
专题命中
视觉问答
:VLM(abstract);分类 cs.CV
CommentsAccepted by IEEE Transactions on Multimedia (TMM)
VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
Ziyang Zhang, Yang Yu, Xulei Yang, Si Yong Yeo
机构
*
MedVisAI Lab Department of ECE Northwestern University(MedVisAI实验室 电子工程系 西北大学)
;
Institute for Infocomm Research (I 2 R) A*STAR, Singapore(信息与通信研究所(I 2 R)A*STAR,新加坡)
;
MedVisAI Lab Lee Kong Chian School of Medicine, Nanyang Technological University(MedVisAI实验室 李科田医学院,南洋理工大学)
机构
*
University of Sydney(悉尼大学)
;
University of Wollongong(沃林戈大学)
;
University of Adelaide(阿德莱德大学)
;
First Clinical Medical College, Guangzhou University of Chinese Medicine(广州中医药大学第一临床学院)
A Systematic Review of Key Retrieval-Augmented Generation (RAG) Systems: Progress, Gaps, and Future Directions
Agada Joseph Oche, Ademola Glory Folashade, Tirthankar Ghosal, Arpan Biswas
机构
*
Bredesen Center for Interdisciplinary Research, University of Tennessee, Knoxville, USA(跨学科研究中心,田纳西大学,肯塔基州 Knoxville)
;
National Center for Computational Sciences, Oak Ridge National Laboratory, Oak Ridge, USA(计算科学国家中心,橡树岭国家实验室,橡树岭)
;
University of Tennessee-Oak Ridge Innovation Institute, University of Tennessee, Knoxville, USA(田纳西大学橡树岭创新研究所,田纳西大学,肯塔基州 Knoxville)
Long-Form Answers to Visual Questions from Blind and Low Vision People
Mina Huh, Fangyuan Xu, Yi-Hao Peng, Chongyan Chen, Hansika Murugu, Danna Gurari, Eunsol Choi, Amy Pavel
机构
*
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Hong Kong University of Science and Technology(香港科学与技术大学)
;
University of Colorado Boulder(科罗拉多大学博尔德分校)
专题命中
视觉问答
:vision language model(abstract);分类 cs.CV