机构
*
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Department of Computer Sciences, University of Wisconsin–Madison(威斯康星大学麦迪逊分校计算机科学系)
;
Department of Statistics, University of Wisconsin–Madison(威斯康星大学麦迪逊分校统计学系)
Comments17 pages, 5 figures, 9 tables. v2 corrects scorer and taxonomy defects, adds no-model baselines showing label leakage, re-runs the Lithuanian cells on de-leaked text, and withdraws the claim that few-shot helps on judgment-form classification everywhere; all tables and figures regenerated. Dataset: https://huggingface.co/datasets/overthelex/multi-legal-bench
Role Steering of Language Models for Social Simulations
用于社会模拟的语言模型角色引导
Isaac Song, Mohammed Rehan Parwani, Glenn Matlin, Emile Anand, Akhil Theerthala, Arjun Chatterjee, Anthony Wen-Ming Zang, Maria Kostylew, Yonadav G. Shavit, Sebastien Krier, Mark Riedl
机构
*
Georgia Institute of Technology(佐治亚理工学院)
;
ML Alignment & Theory Scholars (MATS)(ML对齐与理论学者组织(MATS))
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Oxford(牛津大学)
;
Google DeepMind(谷歌DeepMind)
;
OpenAI(开放人工智能公司)
机构
*
The University of Hong Kong(香港大学)
;
Nanjing University(南京大学)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
;
Fudan University(复旦大学)
Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
我们真的需要参数超过10亿的多模态情感语言模型吗?
Kaiwen Zheng, Junchen Fu, Wenhao Deng, Hu Han, Joemon M. Jose, Xuri Ge
机构
*
University of Glasgow(格拉斯哥大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
School of Artificial Intelligence, Shandong University(山东大学人工智能学院)
机构
*
School of Computing, Informatics Institute of Technology(计算学院,信息学院技术研究所)
;
School of Mathematics and Computer Science, Swansea University(数学与计算机科学学院,萨默塞特大学)
;
R&D, Zame AI(Zame AI研发部)
Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis
适配器合并重新激活潜在推理痕迹:一种机制分析
Junyi Zou
机构
*
Zjydiary Group(Zjydiary小组)
专题命中
安全评测
:alignment(abstract);分类 cs.CL、cs.AI
AI总结
研究揭示适配器合并后推理痕迹的重新激活机制,通过几何方法减少泄漏并提升准确性。
CommentsWithdrawn by the authors after identifying an implementation error in adapter coefficient scaling that materially affects the main empirical results and invalidates the current conclusions. A corrected reanalysis is in progress