CommentsProceedings of the 6th Workshop on Trustworthy NLP (TrustNLP 2026), ACL 2026, San Diego, California, USA. Available at https://openreview.net/forum?id=WJCalficPT
机构
*
Center of Excellence for Generative AI, KAUST(KAUST生成式人工智能卓越中心)
;
Jilin University(吉林大学)
;
Zhejiang University(浙江大学)
;
The Swiss AI Lab, IDSIA-USI/SUPSI(瑞士人工智能实验室 IDSIA-USI/SUPSI)
;
NNAISENSE
Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
评估稀疏自编码器与概念标注的可解释性
Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Begüm Demir
机构
*
The Berlin Institute for the Foundations of Learning and Data (BIFOLD)(柏林学习与数据基础研究所)
;
Technische Universität Berlin(柏林工业大学)
;
INRAE(法国国家农业、食品与环境研究院)
;
Inria, EVERGREEN(法国国家信息与自动化研究所,EVERGREEN)
;
UMR TETIS, Univ Montpellier(UMR TETIS,蒙彼利埃大学)
Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs
业务即规则:面向LLM的业务规则流建模基准与框架
Chen Yang, Ruping Xu, Ruizhe Li, Bin Cao, Jing Fan
机构
*
Zhejiang University of Technology(浙江工业大学)
;
Zhejiang Key Laboratory of Visual Information Intelligent Processing(浙江省视觉信息智能处理重点实验室)
;
University of Aberdeen(阿伯丁大学)
;
University of Birmingham(伯明翰大学)
Societal Alignment Frameworks Can Improve LLM Alignment
社会对齐框架可以改进大语言模型对齐
Karolina Stańczak, Nicholas Meade, Mehar Bhatia, Hattie Zhou, Konstantin Böttinger, Jeremy Barnes, Jason Stanley, Jessica Montgomery, Richard Zemel, Nicolas Papernot, Nicolas Chapados, Denis Therien, Timothy P. Lillicrap, Ana Marasović, Sylvie Delacroix, Gillian K. Hadfield, Siva Reddy
机构
*
ETH Zurich(苏黎世联邦理工学院)
;
Mila, McGill University(麦吉尔大学米尔人工智能实验室)
;
University of Cambridge(剑桥大学)
;
Columbia University(哥伦比亚大学)
;
University of Toronto, Google DeepMind(多伦多大学与DeepMind)
;
McGill University, ServiceNow(麦吉尔大学与ServiceNow)
;
Google DeepMind(谷歌DeepMind)
;
University of Utah(犹他大学)
;
King's College London(伦敦国王学院)
;
Johns Hopkins University(约翰霍普金斯大学)
;
Mila, McGill University, ServiceNow(麦吉尔大学米尔人工智能实验室与ServiceNow)
Operationalizing Fairness: Post-Hoc Threshold Optimization Under Hard Resource Limits
将公平性操作化:在硬性资源限制下的事后阈值优化
Moirangthem Tiken Singh, Amit Kalita, Sapam Jitu Singh
机构
*
Department of Computer Science and Engineering, DUIET, Dibrugarh University, Assam, India(迪布鲁大学计算机科学与工程系,DUIET,阿萨姆,印度)
;
Department of Computer Science and Engineering, MIT, Manipur University, 795003, India(曼尼普尔大学计算机科学与工程系,MIT,曼尼普尔大学,印度)