机构
*
MBZUAI(穆罕默德·本·扎耶德人工智能大学)
;
Texas A&M University(德州农工大学)
;
National University of Singapore(新加坡国立大学)
;
UCLA(加州大学洛杉矶分校)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
Mass General Hospital(麻省总医院)
;
Harvard Medical School(哈佛医学院)
专题命中
知识编辑与模型理解
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
CommentsAccepted at the IJCAI-ECAI Joint Workshop on Planning for Complex Real-World Applications and Bridging the Gap Between AI Planning and (Reinforcement) Learning
What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics
中间层知道什么:从熵动力学检测越狱
Sofiia Nikolenko, Michele Papucci, Mina Rezaei, Shireen Kudukkil Manchingal
机构
*
LMU Munich(慕尼黑大学)
;
relAI – Konrad Zuse School of Excellence in Reliable AI(relAI – 康拉德·楚泽可靠人工智能卓越学校)
;
University of Pisa(比萨大学)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
School of Engineering, Computing and Mathematics, Oxford Brookes University(牛津布鲁克斯大学工程、计算与数学学院)
专题命中
知识编辑与模型理解
:LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
CommentsAccepted at the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD) 2026. A short version accepted at EIML@ICML 2026
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
通过内部归因图对大语言模型越狱进行机制性可解释性研究
Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang, Longwei Wang
机构
*
Department of Computer Science, University of South Dakota(南达科他大学计算机科学系)
;
School of Information and Artificial Intelligence, Yangzhou University(扬州大学信息与人工智能学院)
专题命中
知识编辑与模型理解
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
InA-Probe: Instruction-Aware Active Probing for Time Series Forecasting with LLMs
InA-Probe:面向LLM时间序列预测的指令感知主动探测
Peiliang Gong, Emadeldeen Eldele, Chenyu Liu, Ziyu Jia, Yi Ding, Xinliang Zhou, Lianchao Gu, Qi Zhu, Yang Liu, Daoqiang Zhang, Xiaoli Li
机构
*
Nanyang Technological University(南洋理工大学)
;
Khalifa University(哈利法大学)
;
Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
;
Singapore University of Technology and Design(新加坡科技设计大学)
专题命中
知识编辑与模型理解
:LLM(title_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI
Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia
基于减法映射的扰动型区域可解释性方法(PRISM):语言模型与卒中后失语症中的命名错误分离现象
Xiang Guan, Roger D. Newman-Norlund, Yong Yang, Saeed Ahmadi, Regan Willis, Nadra Salman, Kalil Warren, Srihari Nelakuditi, Chris Rorden, Leonardo Bonilha, Julius Fridriksson
机构
*
University of South Carolina(南卡罗来纳大学)
;
ALLT.AI, LLC(ALLT.AI有限责任公司)
;
USC School of Medicine(南卡罗来纳大学医学院)
专题命中
知识编辑与模型理解
:language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.LG