StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
StreamSense: 基于选择性视觉-语言模型路由的流式社交任务检测
Han Wang, Deyi Ji, Lanyun Zhu, Jiebo Luo, Roy Ka-Wei Lee
机构
*
Singapore University of Technology and Design(新加坡科技设计大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Nanyang Technological University(南洋理工大学)
;
University of Rochester(罗切斯特大学)
Leveraging Data to Say No: Memory Augmented Plug-and-Play Selective Prediction
利用数据说不:基于记忆的插拔式选择预测
Aditya Sarkar, Yi Li, Jiacheng Cheng, Shlok Mishra, Nuno Vasconcelos
机构
*
University of Maryland, College Park(马里兰大学)
;
University of California, San Diego(加州大学圣地亚哥分校)
;
Qualcomm AI(高通人工智能)
;
Yale University(耶鲁大学)
;
Meta AI
机构
*
Department of Computer Science and Engineering, Lehigh University(计算机科学与工程系,莱恩大学)
;
Department of Computer Science(计算机科学系)
;
Engineering, Lehigh University(工程系,莱恩大学)
;
Qualcomm AI Research(高通人工智能研究)
NLPrompt: Noise-Label Prompt Learning for Vision-Language Models
NLPrompt: 噪声标签提示学习用于视觉-语言模型
Bikang Pan, Qun Li, Xiaoying Tang, Wei Huang, Zhen Fang, Feng Liu, Jingya Wang, Jingyi Yu, Ye Shi
机构
*
ShanghaiTech University(上海科技大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
RIKEN Center for Advanced Intelligence Project(日本RIKEN高级智能项目中心)
;
University of Technology Sydney(悉尼科技大学)
;
University of Melbourne(墨尔本大学)
机构
*
Department of Computer Science and Engineering, The Chinese University of Hong Kong(计算机科学与工程系,香港中文大学)
;
Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(医学智能与XR研究院,香港中文大学)
;
Department of Computer Science, Khalifa University(计算机科学系,哈利法大学)
Knowledge-enhanced Pretraining for Vision-language Pathology Foundation Model on Cancer Diagnosis
基于知识增强的视觉语言病理基础模型用于癌症诊断
Xiao Zhou, Luoyi Sun, Dexuan He, Wenbin Guan, Ge Wang, Ruifen Wang, Lifeng Wang, Xiaojun Yuan, Xin Sun, Ya Zhang, Kun Sun, Yanfeng Wang, Weidi Xie
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Shanghai Jiao Tong University(上海交通大学)
;
Xinhua Hospital Affiliated to Shanghai Jiao Tong University School of Medicine(上海交通大学医学院附属新华医院)
;
School of Artificial Intelligence(人工智能学院)
;
Department of Pathology(病理学部)
;
Department of Oral Pathology(口腔病理学部)
;
Department of Pediatric Hematology/Oncology(儿科血液肿瘤科)
;
Clinical Research and Innovation Unit(临床研究与创新单元)
;
Department of Pediatric Cardiology(儿童心脏病科)
Minh Duc Chu, Kshitij Pawar, Zihao He, Roxanna Sharifi, Ross Sonnenblick, Magdalayna Curry, Laura D'Adamo, Lindsay Young, Stuart B Murray, Kristina Lerman
机构
*
USC Information Sciences Institute(USC信息科学研究所)
;
Keck School of Medicine, USC(USC凯克医学院)
;
Department of Clinical Psychology, Drexel University(德雷塞尔大学临床心理学系)
;
Department of Psychiatry and Biobehavioral Sciences, UCLA(UCLA精神病学与生物行为科学系)
机构
*
School of Science, Southern University of Science and Technology(南方科技大学科学学院)
;
China University of Geosciences(中国地质大学)
;
University of Hong Kong(香港大学)
;
Hubei Key Laboratory of Planetary Geology, China University of Geosciences(湖北省行星地质重点实验室,中国地质大学)
;
The Chinese University of Hong Kong(香港中文大学)
Can Synthetic Images Serve as Effective and Efficient Class Prototypes?
合成图像能否作为有效的高效类别原型?
Dianxing Shi, Dingjie Fu, Yuqiao Liu, Jun Wang
机构
*
Beijing Research Institute of Uranium Geology(铀矿北京研究机构)
;
Huazhong University of Science and Technology(华中科技大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Great Bay University(大湾大学)
Proxy Robustness in Vision Language Models is Effortlessly Transferable
视觉语言模型中的代理鲁棒性可以轻易转移
Xiaowei Fu, Fuxiang Huang, Lei Zhang
机构
*
Chongqing Key Laboratory of Bio-perception(重庆生物感知实验室)
;
Multimodal Intelligent Information Processing, Chongqing University, Chongqing 401331, China(多模态智能信息处理,重庆大学,重庆401331,中国)
;
School of Microelectronics(微电子学院)
;
Communication Engineering, Chongqing University, Chongqing 401331, China(通信工程,重庆大学,重庆401331,中国)
;
School of Data Science, Lingnan University, Hong Kong(数据科学学院,岭南大学,香港)
REF-VLM: Triplet-Based Referring Paradigm for Unified Visual Decoding
基于三元组的引用范式用于统一视觉解码
Yan Tai, Luhao Zhu, Yunan Ding, Yiying Dong, Guangtao Zhai, Xiaohong Liu, Guodong Guo
机构
*
School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China(上海交通大学计算机科学学院)
;
Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo, China(宁波数字孪生研究院)
;
School of Information Science and Electronic Engineering, Shanghai Jiao Tong University, Shanghai, 200240, China(上海交通大学信息科学与电子工程学院)
SVII-3D: Advancing Roadside Infrastructure Inventory with Decimeter-level 3D Localization and Comprehension from Sparse Street Imagery
SVII-3D:利用厘米级3D定位与稀疏街道影像的综合理解,推进道路基础设施库存建设
Chong Liu, Luxuan Fu, Yang Jia, Zhen Dong, Bisheng Yang
机构
*
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing (LIESMARS), Wuhan University, Wuhan 430079, China(信息工程测绘遥感国家重点实验室(LIESMARS),武汉大学)
;
Research Institute Ltd, Chengdu 610000, China(四川省公路规划设计研究有限公司)
Studying Illustrations in Manuscripts: An Efficient Deep-Learning Approach
研究手稿中的插图:一种高效的深度学习方法
Yoav Evron, Michal Bar-Asher Siegal, Michael Fire
机构
*
Faculty of Computer and Information Science, Ben-Gurion University of the Negev, Be’er Sheva, Israel(计算机与信息科学学院,内盖夫本· Gurion大学,贝尔谢巴,以色列)
;
The Goldstein-Goren Department of Jewish Thought, Ben-Gurion University of the Negev, Be’er Sheva, Israel(犹太思想Goldstein-Goren部门,内盖夫本· Gurion大学,贝尔谢巴,以色列)
机构
*
Institute of Information Science, Beijing Jiaotong University(北京交通大学信息科学学院)
;
National University of Singapore(新加坡国立大学)
;
Meituan(美团)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Picsart AI Research (PAIR)(Picsart AI研究(PAIR))