Conditional Representation Learning for Customized Tasks
基于条件的表示学习用于定制任务
Honglin Liu, Chao Sun, Peng Hu, Yunfan Li, Xi Peng
机构
*
College of Computer Science, Sichuan University(四川大学计算机学院)
;
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空航天信息研究所)
;
National Key Laboratory of Fundamental Algorithms and Models for Engineering Numerical Simulation, Sichuan University(四川大学工程数值模拟基础算法与模型国家重点实验室)
专题命中
知识编辑与模型理解
:LLM(abstract);large language model(abstract);language model(abstract)
Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
Laura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer, Anna Hedström, Marina M. -C. Höhne, Oliver Eberle
机构
*
Technische Universität Berlin(柏林技术大学)
;
BIFOLD
;
UMI Lab(UMI实验室)
;
Fraunhofer Heinrich-Hertz-Institute(弗劳恩霍夫海因里希-赫兹研究所)
;
ETH AI Center(苏黎世联邦理工学院人工智能中心)
;
Universität Potsdam(波茨坦大学)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
专题命中
知识编辑与模型理解
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
Jerry Huang, Prasanna Parthasarathi, Mehdi Rezagholizadeh, Boxing Chen, Sarath Chandar
机构
*
Mila & Université de Montréal(Mila与蒙特利尔大学)
;
Noah’s Ark Lab(Noah’s Ark实验室)
;
Advanced Micro Devices
;
Chandar Research Lab(Chandar研究实验室)
;
Polytechnique Montréal(蒙特利尔理工学院)
;
CIFAR AI Chair(CIFAR人工智能主席)
专题命中
知识编辑与模型理解
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
CommentsAccepted to Findings of The 63rd Annual Meeting of the Association for Computational Linguistics (ACL) 2025. Official proceedings version available at https://aclanthology.org/2025.findings-acl.60/
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
Jianhui Chen, Xiaozhi Wang, Zijun Yao, Yushi Bai, Lei Hou, Juanzi Li
机构
*
Department of Computer Science and Technology, BNRist(计算机科学与技术系,BNRist)
;
Shenzhen International Graduate School(深圳国际研究生院)
;
KIRC, Institute for Artificial Intelligence, Tsinghua University(人工智能研究院,清华大学)
专题命中
知识编辑与模型理解
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
T. Ed Li, Junyu Ren
机构
*
Yale University(耶鲁大学)
;
University of Chicago(芝加哥大学)
专题命中
知识编辑与模型理解
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
CommentsFirst submitted on February 10th, 2025 to ICLR 2025 Workshop (XAI4Science: From Understanding Model Behavior to Discovering New Scientific Knowledge). The paper was accepted but the workshop does not generate proceedings. Now uploading to arXiv to make the paper publicly available
Learning to Reason for Hallucination Span Detection
Hsuan Su, Ting-Yao Hu, Hema Swetha Koppula, Kundan Krishna, Hadi Pouransari, Cheng-Yu Hsieh, Cem Koc, Joseph Yitan Cheng, Oncel Tuzel, Raviteja Vemulapalli
机构
*
National Taiwan University(国立台湾大学)
;
Apple(苹果公司)
专题命中
知识编辑与模型理解
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
Xu Wang, Yan Hu, Benyou Wang, Difan Zou
机构
*
School of Computing and Data Science, The University of Hong Kong(计算与数据科学学院,香港大学)
;
School of Data Science, The Chinese University of Hong Kong, Shenzhen(数据科学学院,香港中文大学(深圳))
专题命中
知识编辑与模型理解
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
Hallucinated Span Detection with Multi-View Attention Features
Yuya Ogasa, Yuki Arase
机构
*
Grad. Sch. of Information Science and Tech.(信息科学与技术研究生院)
;
The University of Osaka(大阪大学)
;
School of Computing(计算学部)
;
Institute of Science(科学研究所)
;
LY Corporation(LY公司)
专题命中
知识编辑与模型理解
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG