Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences
机构 * Stanford University(斯坦福大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Stanford University(斯坦福大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
机构 * HKUST(香港科技大学) ; Microsoft Research Asia(微软亚洲研究院) ; Duke University(杜克大学) ; Northwestern University(西北大学) ; Johns Hopkins University(约翰霍普金斯大学) ; Microsoft(微软)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Tencent(腾讯)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
机构 * College of Computing(计算学院) ; Georgia Institute of Technology(佐治亚理工学院) ; University of Illinois(伊利诺伊大学) ; Neuro Industry Research(神经产业研究) ; Neuro Industry, Inc.(神经产业公司)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG
机构 * IIIT Dharwad(德瓦德理工学院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) ; Western University(西部大学)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG
机构 * The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室) ; Nanyang Technological University(南洋理工大学) ; University of Maryland(马里兰大学) ; IBM Research(IBM研究院) ; CISPA Helmholtz Center for Information Security(CISPA 欧洲信息安全部分) ; University of Oxford(牛津大学)
专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Perspective paper for a broader scientific audience. The first two authors contributed equally to this paper. 13 pages
机构 * Nokia Bell Labs(诺基亚贝尔实验室)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
Comments This paper has been accepted in AIES 2025
机构 * School of Computing, ADAPT Centre, Dublin City University, Dublin, Ireland(计算学院、ADAPT中心、都柏林城市大学、都柏林、爱尔兰)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
机构 * Chalmers University of Technology, Sweden Carl von Ossietzky Universität Oldenburg, Germany Universit\'e Grenoble Alpes, France University of Warwick, United Kingdom CSX-AI, France SRI International, United States
专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
Comments 32 pages
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Accepted to CogSci 2025. Code can be found at https://github.com/dylanwaldner/BeGoodOrSurvive
机构 * AI Standards Lab(AI标准实验室) ; Technical University of Munich(慕尼黑技术大学) ; Institute of Data Science(数据科学研究所) ; Digital Technologies(数字技术) ; Trajectory Labs(轨迹实验室)
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments 10 pages, 1 table. Accepted to the ICML 2025 Technical AI Governance Workshop
机构 * University Tübingen(图宾根大学) ; CZS Institute for Artificial Intelligence and Law(法律与人工智能研究所)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
机构 * Tongji University(同济大学) ; Johns Hopkins University(约翰霍普金斯大学) ; University of California, Los Angeles(加州大学洛杉矶分校) ; Microsoft Research Asia(微软亚洲研究院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
Comments ICML 2025
机构 * The Beacom College of Computer and Cyber Sciences(贝科姆计算机与网络科学学院) ; Dakota State University(达科他州立大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
Comments 6 pages, 0 figures, 8 tables
Journal ref 2025 IEEE 13th International Symposium on Digital Forensics and Security (ISDFS)
机构 * Institute of Information Science, Academia Sinica(学术院信息研究所)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
Journal ref Ethics and Information Technology 27(2): 20 (2025)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Cooperative AI Foundation, Technical Report #1
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
Comments 30 pages
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY
Comments Preprint to be published in Proceedings of PACLIC38
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
Journal ref Transformers and large language models in healthcare: A review, Artificial Intelligence in Medicine, Volume 154, 2024, 102900,
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
Comments 19 pages, 6 figures, submitted to Conference ICADCML2025
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments MICCAI 2024 Workshop on Fairness of AI in Medical Imaging
专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments 28 pages, 4 figures, 2 tables
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Accepted by ACM Multimedia 2024. The dataset and code can be found at https://github.com/achernarwang/LiVO