Neuroscience-Inspired Analyses of Visual Interestingness in Multimodal Transformers
受神经科学启发的多模态Transformer中视觉趣味性分析
Mathis Immertreu, Fitim Abdullahu, Thomas Kinfe, Helmut Grabner, Patrick Krauss, Achim Schilling
机构
*
Cognitive Computational Neuroscience Group, Patter Recognition Lab, University Erlangen-Nürnberg(认知计算神经科学组、模式识别实验室、埃尔兰根-纽伦堡大学)
;
IDS Institut für Data Science, ZHAW School of Engineering, Winterthur, Switzerland(IDS数据科学研究所、ZHAW工程学院、温特图尔,瑞士)
;
Mannheim Center for Neuromodulation and Neuroprosthetics, University Hospital Mannheim, Heidelberg University(曼海姆神经调制与神经假体中心、曼海姆大学医院、海德堡大学)
;
BGU Ludwigshafen, Germany(BGU路易斯港,德国)
;
Physics and Cognition Group, MCNN, University Hospital Mannheim, Heidelberg University(物理与认知组、MCNN、曼海姆大学医院、海德堡大学)
;
NeuroAI and BCI Group, MCNN, University Hospital Mannheim, Heidelberg University(神经AI与BCI组、MCNN、曼海姆大学医院、海德堡大学)
CapCLIP: A Vision-Language Representation Alignment Approach for Wireless Capsule Endoscopy Analysis
CapCLIP:一种用于无线胶囊内镜分析的视觉-语言表示对齐方法
Haroon Wahab, Irfan Mehmood, Hassan Ugail
机构
*
School of Computer Science, AI and Electronics Faculty of Engineering and Digital Technologies(计算机科学与电子工程学院,工程与数字技术学院)
;
School of Management Faculty of Mgmt, Law & Social Sciences(管理学院,管理、法律与社会科学学院)
;
Centre for Visual Computing and Intelligent Systems(视觉计算与智能系统中心)
机构
*
Qwen Large Model Application Team, Alibaba(阿里巴巴大模型应用团队)
;
Alibaba Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所阿里巴巴分所)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
LVLMs中的词汇劫持:通过排除惰性标记来揭示关键注意力头以缓解幻觉
Yangneng Chen, Junlin Li, Weijun Yao, Xilai Ma, Guodong Du, Wenya Wang, Jing Li
机构
*
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Huawei Technologies Co., Ltd.(华为技术有限公司)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Nanyang Technological University(南洋理工大学)
机构
*
Beijing University of Posts and Telecommunications, China(北京邮电大学)
;
National Institute of Informatics, Japan(日本国立信息机构)
;
Peking University, China(北京大学)
;
The University of Tokyo, Japan(东京大学)
机构
*
IRMV Lab, Shanghai Jiao Tong University(上海交通大学IMV实验室)
;
Meta Reality Labs(Meta现实实验室)
;
Department of Electronic Engineering, Shanghai Jiao Tong University(上海交通大学电子工程系)
;
China University of Mining and Technology(中国矿业大学)
;
College of Intelligence Science and Technology, National University of Defense Technology(国防科技大学智能科学与技术学院)
机构
*
School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications(智能工程与自动化学院,北京邮电大学)
;
State Grid Corporation of China(国家电网公司)
;
School of Artificial Intelligence, Beijing University of Posts and Telecommunications(人工智能学院,北京邮电大学)
Qing Zhong, Guodong Ding, Lingqiao Liu, Zaiwen Feng, Lin Yuanbo Wu, Angela Yao
机构
*
College of Informatics, Huazhong Agricultural University, Wuhan, China(华中农业大学信息学院)
;
School of Computing, National University of Singapore, Singapore(新加坡国立大学计算机学院)
;
School of Computer Science, Adelaide University, Australia(阿德莱德大学计算机科学学院)
;
School of Engineering, University of Warwick, Coventry, UK(沃里克大学工程学院)
;
Zhejiang Yuexiu University, Shaoxing, China(浙江越秀大学)
Towards Conversational Medical AI with Eyes, Ears and a Voice
面向有眼睛、耳朵和声音的对话式医疗AI
Meet Shah, Jason Gusdorf, Anil Palepu, Chunjong Park, Jack W. O'Sullivan, Vishnu Ravi, Tim Strother, Pavel Dubov, Aliya Rysbek, Toshiyuki Fukuzawa, Yana Lunts, Jan Freyberg, Michael B. Chang, Aniruddh Raghu, David Stutz, Devora Berlowitz, Eliseo Papa, Taylan Cemgil, JD Velasquez, Jack Chen, Arthur Chen, Doug Fritz, Charlie Taylor, Katya Tregubova, Jing Rong Lim, Richard Green, Sara Mahdavi, Mahvish Nagda, Jihyeon Lee, Craig Schiff, Liviu Panait, Sukhdeep Singh, Valentin Liévin, David G. T. Barrett, Hannah Gladman, Anna Cupani, Francesca Pietra, Uchechi Okereke, Katherine Tong, Clemens Meyer, Erwan Rolland, Mili Sanwalka, Michael D. Howell, Shixiang Shane Gu, Bibo Xu, Euan A. Ashley, S. M. Ali Eslami, Gregory Wayne, Pushmeet Kohli, Vivek Natarajan, Adam Rodman, Alan Karthikesalingam, Ryutaro Tanno
机构
*
Google DeepMind(谷歌DeepMind)
;
Google Research(谷歌研究)
;
Beth Israel Deaconess Medical Center, Harvard Medical School(贝塞斯达医院, 哈佛医学院)
;
Stanford University(斯坦福大学)
Geospatial-Temporal Sensemaking of Remote Sensing Activity Detections with Multimodal Large Language Model
基于多模态大语言模型的遥感活动检测的时空感知
David F. Ramirez, Tim Overman, Kristen Jaskie, Andreas Spanias
机构
*
SenSIP Center, School of ECEE, Arizona State University(SenSIP中心,电子与计算机工程学院,亚利桑那州立大学)
;
Prime Solutions Group Inc(Prime Solutions Group公司)
;
Intelligence Advanced Research Projects Activity(智能高级研究计划局)