Comments7 pages, with one page for appendix. Accepted for publication at the 2025 21th International Conference on Wireless and Mobile Computing, Networking and Communications (WiMob)
Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability
大语言模型中的道德敏感性:通过行为剖析和机制可解释性对上下文偏见进行分层评估
Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur
机构
*
University of Maryland, College Park(马里兰大学学院公园分校)
;
Purdue University(普渡大学)
;
Meta
;
Apple(苹果公司)
;
Wright State University(怀特州立大学)
;
University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)
专题命中
知识编辑与模型理解
:SLM(abstract,abstract_cn);LLM(abstract_cn);large language model(abstract);language model(abstract)
ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents
ASA:无需骨干训练的工具调用智能体表示工程
Youjin Wang, Run Zhou, Yingjie Ma, Rong Fu, Jiani Liang, Shuaishuai Cao, Min Huang, Tao Fang, Liangming Pan
机构
*
Renmin University of China(中国人民大学)
;
University of Macau(澳门大学)
;
Central South University(中南大学)
;
Jiangxi Normal University(江西师范大学)
;
Macau Millennium College(澳门 millennium 学院)
;
Peking University(北京大学)
CommentsThe manuscript consists of 24 pages formatted in the ACL style. Youjin Wang, Run Zhou, and Yingjie Ma contributed equally to this work. Tao Fang and Liangming Pan are the co-corresponding authors
机构
*
Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究所)
;
Shanghai Innovation Institute(上海创新研究院)
;
University of Maryland, College Park(马里兰大学帕克分校)