Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
大语言模型的越狱与漏洞缓解
Benji Peng, Hanxuan Chen, Keyu Chen, Qian Niu, Ziqian Bi, Ming Liu, Pohsun Feng, Tianyang Wang, Lawrence K. Q. Yan, Yizhu Wen, Yichao Zhang, Caitlyn Heqi Yin, Xinyuan Song, Riyang Bao, Jiacheng Shi
机构
*
Hunan University Changsha, PRC
;
Georgia Institute of Technology Atlanta, USA
;
Kyoto University Kyoto, Japan
;
Purdue University West Lafayette, USA
;
National Taiwan Normal University Taipei, ROC
;
University of Liverpool Suzhou, PRC
;
Hong Kong University of Science
;
University of Hawaii Honolulu, USA
;
The University of Texas at Dallas Dallas, USA
;
University of Wisconsin-Madison Madison, USA
;
Emory University Atlanta, USA
;
College of William \& Mary Williamsburg, USA
机构
*
John A. Paulson School of Engineering And Applied Sciences, Harvard University(哈佛大学约翰·A·保罗森工程与应用科学学院)
;
Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology(麻省理工学院脑科学与认知科学系)
;
Speech and Hearing Bioscience and Technology, Harvard Medical School(哈佛医学院语音与听力生物科学与技术系)
;
Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究学院)
;
Center for Brain Science, Harvard University(哈佛大学脑科学中心)
Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents
Agora: 面向生产级共识协议中自主漏洞检测的LLM智能体
Xiang Liu, Sa Song, Zhaowei Zhang, Huiying Lan, Jason Zeng, Ming Wu, Michael Heinrich, Yong Sun, Ceyao Zhang
机构
*
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
;
School of Information and Telecommunication Engineering, Beijing University of Posts and Telecommunications(北京邮电大学信息与电信工程学院)
;
Peking University(北京大学)
;
G Labs(0G实验室)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
测量基于LLM的简历筛选中真实世界的提示注入攻击
Mohan Zhang, Yuqi Jia, Zhen Tan, Steven Jiang, Neil Zhenqiang Gong, Tianlong Chen, Dawn Song
机构
*
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
Duke University(杜克大学)
;
Arizona State University(亚利桑那州立大学)
;
hireEZ
;
University of California, Berkeley(加州大学伯克利分校)
Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software
物理学就是一切?物理学家监督人工智能开发科学软件的案例研究
Nhat-Minh Nguyen
机构
*
Kavli IPMU (WPI), UTIAS, The University of Tokyo(Kavli研究所(WPI)、UTIAS、东京大学)
;
Center for Data-Driven Discovery(数据驱动发现中心)
;
Institute For Interdisciplinary Research in Science(科学跨学科研究中心)
Comments10 pages, 2 figures, 2 tables, 1 physicist and a few AI agents. Accepted by ICML 2026 AI for Science Workshop. Code and development log are available at this repo: https://github.com/MinhMPA/clax-pt
Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
校准还不够:评估语言变化下的置信度估计
Yuxi Xia, Dennis Ulmer, Terra Blevins, Yihong Liu, Hinrich Schütze, Benjamin Roth
机构
*
Faculty of Computer Science, UniVie Doctoral School Computer Science(计算机科学系,维也纳大学计算机科学博士学院)
;
Faculty of Philological and Cultural Studies, University of Vienna, Austria(文学与文化研究系,维也纳大学,奥地利)
;
ILLC, University of Amsterdam, Netherlands(阿姆斯特丹大学ILLC,荷兰)
;
Khoury College of Computer Sciences, Northeastern University, USA(东北大学计算机科学学院,美国)
;
LMU Munich, Munich Center for Machine Learning (MCML), Germany(慕尼黑大学,慕尼黑机器学习中心(MCML),德国)
Journal refProceedings of The fourth international workshop on the role of resources in the age of large language models RESOURCEFUL-2026 at LREC 2026, Palma de Mallorca, Spain, 2026
Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization
对齐但脆弱:通过零阶优化增强LLM安全鲁棒性
Zhihao Liu, Yifan Wu, Jian Lou, Di Wang, Yuxi Zhou, Yuke Hu
机构
*
The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室)
;
Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高科技园区(滨江)区块链与数据安全研究院)
;
Sun Yat-sen University(中山大学)
;
KAUST(卡塔尔大学)