IMMACULATE: A Practical LLM Auditing Framework via Verifiable Computation
IMMACULATE: 通过可验证计算实现的实用LLM审计框架
Yanpei Guo, Wenjie Qu, Linyu Wu, Shengfang Zhai, Lionel Z. Wang, Ming Xu, Yue Liu, Binhang Yuan, Dawn Song, Jiaheng Zhang
机构
*
National University of Singapore(新加坡国立大学)
;
Nanyang Technological University(南洋理工大学)
;
Independent Researcher(独立研究者)
;
University of California, Berkeley(加州大学伯克利分校)
TxRay: Agentic Postmortem of Live Blockchain Attacks
TxRay:活区块链攻击的代理事后分析
Ziyue Wang, Jiangshan Yu, Kaihua Qin, Dawn Song, Arthur Gervais, Liyi Zhou
机构
*
Decentralized Intelligence AG(去中心化智能AG)
;
The University of Sydney(悉尼大学)
;
University of Warwick(沃里克大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
University College London(伦敦大学学院)
How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
恶意AI群如何威胁民主:代理AI与大语言模型的融合标志着信息战争的新前沿
Daniel Thilo Schroeder, Meeyoung Cha, Andrea Baronchelli, Nick Bostrom, Nicholas A. Christakis, David Garcia, Amit Goldenberg, Yara Kyrychenko, Kevin Leyton-Brown, Nina Lutz, Gary Marcus, Filippo Menczer, Gordon Pennycook, David G. Rand, Maria Ressa, Frank Schweitzer, Dawn Song, Christopher Summerfield, Audrey Tang, Jay J. Van Bavel, Sander van der Linden, Jonas R. Kunst
机构
*
Department of Sustainable Communication Technologies, SINTEF Digital(可持续通信技术系,SINTEF数字)
;
Max Planck Institute for Security and Privacy(安全与隐私研究所)
;
Department of Mathematics, City St George’s University of London(数学系,圣乔治大学)
;
Macrostrategy Research Initiative(战略研究计划)
;
Human Nature Lab, Yale University(人性实验室,耶鲁大学)
;
Department of Politics and Public Administration, University of Konstanz(政治与公共管理系,康斯坦茨大学)
;
Harvard Business School, Harvard University(哈佛商学院,哈佛大学)
;
Department of Psychology, University of Cambridge(心理学系,剑桥大学)
;
Department of Computer Science, University of British Columbia(计算机科学系,不列颠哥伦比亚大学)
;
Department of Human Centered Design & Engineering, University of Washington(以人为本设计与工程系,华盛顿大学)
;
Department of Psychology, New York University(心理学系,纽约大学)
;
Observatory on Social Media and Luddy School of Informatics, Computing, and Engineering, Indiana University(社交媒体观察所和信息、计算与工程学院,印第安纳大学)
Comments5 Pages, This is the author's version of the work. It is posted here by permission of the AAAS for personal use, not for redistribution. The definitive version was published in Science on January 22, 2026, DOI: 10.1126/science.adz1697
Canyu Chen, Baixiang Huang, Zekun Li, Zhaorun Chen, Shiyang Lai, Xiongxiao Xu, Jia-Chen Gu, Jindong Gu, Huaxiu Yao, Chaowei Xiao, Xifeng Yan, William Yang Wang, Philip Torr, Dawn Song, Kai Shu
机构
*
Northwestern University(西北大学)
AI总结
本文提出编辑攻击作为LLM安全威胁,揭示了通过编辑注入虚假信息和偏见的风险及高隐蔽性。
CommentsAccepted to Proceedings of AAAI 2026. The first two authors contributed equally. 7 pages for main paper, 31 pages including appendix. The code, results, dataset for this paper and more resources are on the project website: https://llm-editing.github.io
How and Why LLMs Generalize: A Fine-Grained Analysis of LLM Reasoning from Cognitive Behaviors to Low-Level Patterns
LLMs如何泛化:从认知行为到低级模式的细粒度分析
Haoyue Bai, Yiyou Sun, Wenjie Hu, Shi Qiu, Maggie Ziyu Huan, Peiyang Song, Robert Nowak, Dawn Song
机构
*
University of Wisconsin, Madison(威斯康星大学麦迪逊分校)
;
University of California, Berkeley(加州大学伯克利分校)
;
University of Pennsylvania(宾夕法尼亚大学)
;
California Institute of Technology(加州理工学院)
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
VulnLLM-R:基于代理架构的专用推理LLM用于漏洞检测
Yuzhou Nie, Hongwei Li, Chengquan Guo, Ruizhe Jiang, Zhun Wang, Bo Li, Dawn Song, Wenbo Guo
机构
*
Department of Computer Science, University of California, Santa Barbara, CA, USA(加州大学圣芭芭拉分校计算机科学系)
;
Department of Computer Science, University of Chicago, Chicago, IL, USA(芝加哥大学计算机科学系)
;
Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, CA, USA(加州大学伯克利分校电子工程与计算机科学系)
;
Department of Computer Science, University of Illinois Urbana-Champaign, Champaign, IL, USA(伊利诺伊大学厄巴纳-香槟分校计算机科学系)
Dan Hendrycks, Dawn Song, Christian Szegedy, Honglak Lee, Yarin Gal, Erik Brynjolfsson, Sharon Li, Andy Zou, Lionel Levine, Bo Han, Jie Fu, Ziwei Liu, Jinwoo Shin, Kimin Lee, Mantas Mazeika, Long Phan, George Ingebretsen, Adam Khoja, Cihang Xie, Olawale Salaudeen, Matthias Hein, Kevin Zhao, Alexander Pan, David Duvenaud, Bo Li, Steve Omohundro, Gabriel Alfour, Max Tegmark, Kevin McGrew, Gary Marcus, Jaan Tallinn, Eric Schmidt, Yoshua Bengio
机构
*
Center for AI Safety(AI安全中心)
;
University of California, Berkeley(加州大学伯克利分校)
;
Virtue AI
;
Morph Labs(Morph实验室)
;
University of Michigan(密歇根大学)
;
LG AI Research(LG人工智能研究)
;
University of Oxford(牛津大学)
;
Stanford University(斯坦福大学)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Gray Swan AI
;
Carnegie Mellon University(卡内基梅隆大学)
;
Cornell University(康奈尔大学)
;
Hong Kong Baptist University(香港 Baptist大学)
;
HKUST(香港科技大学)
;
Nanyang Technological University(南洋理工大学)
;
KAIST(韩国科学技术院)
;
University of California, Santa Cruz(加州大学圣克鲁兹分校)
;
Massachusetts Institute of Technology(麻省理工学院)
;
University of Tübingen(图宾根大学)
;
University of Washington(华盛顿大学)
;
University of Toronto(多伦多大学)
;
Vector Institute(向量研究所)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Beneficial AI Research(有益AI研究)
;
Conjecture
;
Institute for Applied Psychometrics(应用心理测量研究所)
;
New York University(纽约大学)
;
CSER
;
Université de Montréal(蒙特利尔大学)
;
LawZero
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
BountyBench: AI代理攻击者和防御者对现实世界网络安全系统的影响
Andy K. Zhang, Joey Ji, Celeste Menders, Riya Dulepet, Thomas Qin, Ron Y. Wang, Junrong Wu, Kyleen Liao, Jiliang Li, Jinghan Hu, Sara Hong, Nardos Demilew, Shivatmica Murgai, Jason Tran, Nishka Kacheria, Ethan Ho, Denis Liu, Lauren McLane, Olivia Bruvik, Dai-Rong Han, Seungwoo Kim, Akhil Vyas, Cuiyuanxiu Chen, Ryan Li, Weiran Xu, Jonathan Z. Ye, Prerit Choudhary, Siddharth M. Bhatia, Vikram Sivashankar, Yuxuan Bao, Dawn Song, Dan Boneh, Daniel E. Ho, Percy Liang
VMDT: Decoding the Trustworthiness of Video Foundation Models
Yujin Potter, Zhun Wang, Nicholas Crispino, Kyle Montgomery, Alexander Xiong, Ethan Y. Chang, Francesco Pinto, Yuqi Chen, Rahul Gupta, Morteza Ziyadi, Christos Christodoulopoulos, Bo Li, Chenguang Wang, Dawn Song
机构
*
University of California, Berkeley(加州大学伯克利分校)
;
University of California, Santa Cruz(加州大学圣克ruz分校)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Chicago(芝加哥大学)
;
Amazon(亚马逊)
;
Information Commissioner’s Office(信息专员办公室)
Predicting Task Performance with Context-aware Scaling Laws
Kyle Montgomery, David Park, Jianhong Tu, Michael Bendersky, Beliz Gunel, Dawn Song, Chenguang Wang
机构
*
UC Santa Cruz(加州大学圣克ruz分校)
;
Washington University in St. Louis(华盛顿大学圣路易斯分校)
;
Databricks(Databricks公司)
;
Google DeepMind(谷歌DeepMind)
;
UC Berkeley(加州大学伯克利分校)