机构
*
Australian National University(澳大利亚国立大学)
;
The University Of Queensland(昆士兰大学)
;
Peking University(北京大学)
;
GE research(通用电气研究院)
;
CSIRO(澳大利亚联邦科学与工业研究组织)
Jona Ruthardt, Manu Gaur, Deva Ramanan, Makarand Tapaswi, Yuki M. Asano
机构
*
University of Technology Nuremberg(纽伦堡工业大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
International Institute of Information Technology, Hyderabad(海得拉巴国际信息技术学院)
Through the Looking Glass: A Dual Perspective on Weakly-Supervised Few-Shot Segmentation
透过镜中奇境:弱监督小样本分割的双重视角
Jiaqi Ma, Guo-Sen Xie, Fang Zhao, Zechao Li
机构
*
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
;
School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院)
UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving
UniDrive: 面向自动驾驶可解释风险理解的统一视觉-语言与定位框架
Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
机构
*
organization= Department of Earth Science \& Engineering, Imperial College London , city= London , postcode= SW7 2AZ , country= United Kingdom
;
organization= SpaceTimeLab, Department of Civil, Environmental
;
Geomatic Engineering, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom
;
organization= Department of Computing, The Hong Kong Polytechnic University , city= Hong Kong , country= China
;
organization= Trinity College, University of Oxford , city= Oxford , postcode= OX1 3BH , country= United Kingdom
;
organization= Department of Geography, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom
;
organization= Centre for Global Infrastructure Resilience, The Bartlett School of Sustainable Construction, University College London , city= London , postcode= WC1E 7HB , country= United Kingdom
Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing
基于百年跨度扫描文档档案微调的页面图像分类器,用于进一步的内容特定处理
Kateryna Lutsai, Dana Křivánková, Pavel Straňák, David Novák
机构
*
Institute of Formal and Applied Linguistics, Charles University MFF(查尔斯大学数学与物理学院形式与应用语言学研究所)
;
Institute of Archaeology, Czech Academy of Sciences(捷克科学院考古研究所)
机构
*
College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院)
;
Department of Bioengineering and Imperial-X, Imperial College London(帝国理工学院伦敦校区生物工程系)
;
Department of Pathology, Xiangtan Maternal and Child Health Hospital(湘潭 maternal and child health hospital pathology department)
;
Department of Pathology, The First People’s Hospital of Xiangtan City(湘潭市第一人民医院病理科)
Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology
面向放射学的空间定位2D视觉-语言模型的可扩展训练
Yusuf Salcan, Simon Ging, Robin Tibor Schirrmeister, Philipp Arnold, Elmar Kotter, Behzad Bozorgtabar, Thomas Brox
机构
*
Computer Vision Group, University of Freiburg, Germany(德国弗莱堡大学计算机视觉组)
;
Department of Radiology, Medical Center -- University of Freiburg, Germany(德国弗莱堡大学医学中心放射科)
;
CRIION-AI Lab, Freiburg, Germany(德国弗莱堡CRIION-AI实验室)
MirrorCheck: Efficient Adversarial Defense for Vision-Language Models
MirrorCheck: 视觉-语言模型的高效对抗防御
Samar Fares, Klea Ziu, Toluwani Aremu, Nikita Durasov, Martin Takáč, Pascal Fua, Ivan Laptev, Karthik Nandakumar
机构
*
Mohamed Bin Zayed University of Artificial Intelligence(莫扎伊德大学人工智能大学)
;
NVIDIA
;
École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)
;
Michigan State University(密歇根州立大学)
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
GeoWorld-VLM:从世界模型中获取几何结构用于视觉-语言模型
Renjie Gu, Kaichen Zhou, Yan Luo, Mengyu Wang
机构
*
Harvard AI and Robotics Lab(哈佛人工智能与机器人实验室)
;
Kempner Institute for the Study of Natural and Artificial Intelligence(凯普纳自然与人工智能研究 institute)
;
Harvard University(哈佛大学)