Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI
机制设计是不够的:面向合作AI的亲社会智能体
Xuanqiang Angelo Huang, Charlie Tharas, Samuele Marro, Van Q. Truong, Bernhard Schölkopf, Emanuele La Malfa, Zhijing Jin
机构
*
ETH Zürich(苏黎世联邦理工学院)
;
University of Oxford(牛津大学)
;
Institute for Decentralized AI(去中心化人工智能研究所)
;
Jinesis Lab, University of Toronto & Vector Institute(多伦多大学Jinesis实验室及向量研究所)
;
EuroSafeAI
;
University of Pennsylvania(宾夕法尼亚大学)
;
Max Planck Institute for Intelligent Systems, Tübingen, Germany(德国图宾根最大计划智能系统研究所)
;
ELLIS Institute Tübingen(图宾根ELLIS研究所)
From monoliths to modules: Decomposing transducers for efficient world modelling
从整体到模块:分解转换器以实现高效的world建模
Alexander Boyd, Franz Nowak, David Hyland, Manuel Baltieri, Fernando E. Rosas
机构
*
Department of Informatics, University of Sussex(Sussex大学信息学院)
;
Beyond Institute for Theoretical Science (BITS)(理论科学研究所)
;
ETH Zürich(苏黎世联邦理工学院)
;
Principles of Intelligent Behaviour in Biological and Social Systems (PIBBSS)(生物和社会系统智能行为原理研究所)
;
Department of Computer Science, University of Oxford(牛津大学计算机科学系)
;
Araya Inc.(Araya公司)
;
Sussex AI and Sussex Centre for Consciousness Science, University of Sussex(Sussex大学人工智能与意识科学中心)
;
Centre for Complexity Science and Center for Psychedelic Research, Department of Brain Sciences, Imperial College London(复杂科学中心和迷幻研究中心,伦敦帝国理工学院脑科学系)
;
Center for Eudaimonia and Human Flourishing, University of Oxford(幸福与人类繁荣中心,牛津大学)
Dung V. Nguyen, Hieu M. Vu, Nhi Y. Pham, Lei Zhang, Tan M. Nguyen
机构
*
Department of Mathematics(数学系)
;
Center for AI Research(人工智能研究中心)
;
National University of Singapore(新加坡国立大学)
;
VinUniversity(文大学)
;
Torilab(Torilab实验室)
Escaping Plato's Cave: JAM for Aligning Independently Trained Vision and Language Models
走出洞穴:JAM用于对齐独立训练的视觉和语言模型
Lauren Hyoseo Yoon, Yisong Yue, Been Kim
机构
*
Computation and Neural Systems(计算与神经系统)
;
California Institute of Technology(加利福尼亚理工学院)
;
Computation and Mathematical Sciences(计算与数学科学)
;
Google DeepMind(谷歌DeepMind)
From Fuzzy to Formal: Scaling Hospital Quality Improvement with AI
从模糊到正式:利用AI扩展医院质量改进
Patrick Vossler, Jean Feng, Venkat Sivaraman, Robert Gallo, Hemal Kanzaria, Dana Freiser, Christopher Ross, Amy Ou, James Marks, Susan Ehrlich, Christopher Peabody, Lucas Zier
机构
*
University of California, San Francisco(加州大学旧金山分校)
;
Zuckerberg San Francisco General Hospital(扎克伯格旧金山总医院)
Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks
揭示并对齐异常注意力头以防御NLP后门攻击
Haotian Jin, Yang Li, Haihui Fan, Lin Shen, Xiangfang Li, Bo Li
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
State Key Laboratory of Cyberspace Security Defense(网络空间安全防御国家重点实验室)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
Comments12 pages, v2: added correction to Polanyi on why tacit knowledge is tacit (structural vs quantitative), unified three independent intellectual threads (Smolensky, Dreyfus, dynamical systems theory)
Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient Influence
Antibody: 通过削弱有害梯度影响来加强对抗有害微调的防御
Quoc Minh Nguyen, Trung Le, Jing Wu, Anh Tuan Bui, Mehrtash Harandi
机构
*
Department of Electrical and Computer Systems Engineering, Monash University, Australia(莫纳什大学电气与计算机系统工程系)
;
Department of Data Science and AI, Monash University, Australia(莫纳什大学数据科学与人工智能系)