Mechanistic Interpretability of Antibody Language Models Using SAEs
使用 SAE 对抗体语言模型的机制可解释性研究
Rebonto Haque, Oliver M. Turnbull, Anisha Parsan, Nithin Parsan, John J. Yang, Anna L. Beukenhorst, Charlotte M. Deane
机构
*
Department of Statistics, University of Oxford, UK(英国牛津大学统计系)
;
Reticular, San Francisco, USA(美国旧金山Reticular公司)
;
EECS, MIT, Cambridge MA, USA(美国麻省理工学院电子工程与计算机科学系)
;
Leyden Laboratories BV, Leiden, The Netherlands(荷兰莱顿实验室)
Commentsv3: 15 pages; corrected author list and affiliations in the main text; minor text changes; updated steering results following minor code changes; conclusions and findings remain unchanged; included link to data and code in the Data Availability section
机构
*
Shanghai Academy of Artificial Intelligence for Science, Shanghai, China.(上海人工智能科学研究院)
;
School of Biomedical Engineering, Shanghai Jiao Tong University, Shanghai, China.(上海交通大学生物医学工程学院)
;
Incubation Institute, Fudan University, Shanghai, China.(复旦大学孵化院)
Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility
这仅仅是幻想吗?语言模型表示反映了人类对事件可能性的判断
Michael A. Lepori, Jennifer Hu, Ishita Dasgupta, Roma Patel, Thomas Serre, Ellie Pavlick
机构
*
Department of Computer Science(计算机科学系)
;
Department of Cognitive Science(认知科学系)
;
Google(谷歌)
;
DeepMind(深度Mind)
;
Brown University(布朗大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
Department of Cognitive(认知系)
;
Department of Computer Science & Psychological Sciences(计算机科学与心理学科学系)
MIMIC: A Generative Multimodal Foundation Model for Biomolecules
MIMIC:一种生成式多模态基础模型用于生物分子
Siavash Golkar, Jake Kovalic, Irina Espejo Morales, Samuel Sledzieski, Minhuan Li, Ksenia Sokolova, Geraud Krawezik, Alberto Bietti, Claudia Skok Gibbs, Roman Klypa, Shengwei Xiong, Francois Lanusse, Liam Parker, Kyunghyun Cho, Miles Cranmer, Tom Hehir, Michael McCabe, Lucas Meyer, Rudy Morel, Payel Mukhopadhyay, Mariel Pettee, Helen Qu, Jeff Shen, David Fouhey, Hadi Sotoudeh, Vikram Mulligan, Pilar Cossio, Sonya M. Hanson, Alisha N. Jones, Olga G. Troyanskaya, Shirley Ho
机构
*
Polymathic AI Center for Data Science, New York University(多学科人工智能数据科学中心,纽约大学)
;
Polymathic AI Department of Applied Physics, Yale University(多学科人工智能应用物理系,耶鲁大学)
;
Center for Computational Mathematics, Flatiron Institute(计算数学中心,Flatiron研究所)
;
Center for Computational Biology, Flatiron Institute(计算生物学中心,Flatiron研究所)
;
Princeton Precision Health, Princeton University(普林斯顿精准健康,普林斯顿大学)
;
Department of Chemistry, New York University(化学系,纽约大学)
;
Department of Computer Science, Princeton University(计算机科学系,普林斯顿大学)
;
Lewis-Sigler Institute for Integrative Genomics, Princeton University(刘易斯-西格尔整合基因组学研究所,普林斯顿大学)
;
Center for Computational Astrophysics, Flatiron Institute(计算天文学中心,Flatiron研究所)
;
Department of Astrophysical Sciences, Princeton University(天体物理科学系,普林斯顿大学)
;
Department of Physics, New York University(物理学系,纽约大学)
scpFormer: A Foundation Model for Unified Representation and Integration of the Single-Cell Proteomics
scpFormer:单细胞蛋白质组学的统一表示与整合基础模型
Qifeng Zhou, Lei Yu, Yuzhi Guo, Yuwei Miao, Hehuan Ma, Wenliang Zhong, Lin Xu, Junzhou Huang
机构
*
Department of Computer Science and Engineering, The University of Texas at Arlington(德克萨斯理工大学计算机科学与工程系)
;
Quantitative Biomedical Research Center, Department of Health Data Science and Biostatistics, Peter O’Donnell Jr. School of Public Health, University of Texas Southwestern Medical Center(德克萨斯西南医学中心量化生物医学研究中心、健康数据科学与生物统计学系、彼得·奥·donnell Jr. 公共卫生学院)
On the Predictive Power of Representation Dispersion in Language Models
语言模型中表示分散度的预测能力研究
Yanhong Li, Ming Li, Karen Livescu, Jiawei Zhou
机构
*
University of Chicago(芝加哥大学)
;
University of Maryland(马里兰大学)
;
Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)
;
Stony Brook University(史泰福布鲁克大学)
;
Allen Institute for AI(人工智能研究院)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
视觉-语言模型中提示诱导幻觉的机制
William Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov, Ritambhara Singh, Carsten Eickhoff, Kyle Mahowald
机构
*
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
Brown University(布朗大学)
;
Technion(技术学院)
;
University of Tübingen(图宾根大学)
;
Harvard University(哈佛大学)
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
SCRIPT: 一种用于韩语预训练语言模型的子字符组合表示注入模块
SungHo Kim, Juhyeong Park, Eda Atalay, SangKeun Lee
机构
*
Department of Artificial Intelligence, Korea University, Seoul, South Korea(人工智能系,韩国大学,首尔,韩国)
;
Department of Computer Science and Engineering, Korea University, Seoul, South Korea(计算机科学与工程系,韩国大学,首尔,韩国)
Dzianis Piatrashyn, Nikita Kotelevskii, Kirill Grishchenkov, Nikita Glazkov, Ivan Nasonov, Ilya Makarov, Timothy Baldwin, Preslav Nakov, Roman Vashurin, Maxim Panov
机构
*
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Ocapital
;
National University of Science and Technology (NUST) MISIS(国立科技大学(NUST)MISIS)
;
AXXX
;
Ivannikov Institute for System Programming of the Russian Academy of Sciences(俄罗斯科学院伊万尼科夫系统编程研究所)
;
Trusted AI Center, RAS(俄罗斯科学院可信人工智能中心)