Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
将AI代理与网络安全专家在真实世界渗透测试中进行比较
Justin W. Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Jun-shen Ho, Anna Wu, Arnold Tianyi Yang, Neil Perry, Andy Zou, Matt Fredrikson, J. Zico Kolter, Percy Liang, Dan Boneh, Daniel E. Ho
Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models
思想图景:可视化大语言模型的推理过程
Zhanke Zhou, Zhaocheng Zhu, Xuan Li, Mikhail Galkin, Xiao Feng, Sanmi Koyejo, Jian Tang, Bo Han
机构
*
TMLR Group, Hong Kong Baptist University(香港 Baptist 大学 TMLR 团体)
;
Stanford University(斯坦福大学)
;
Mila - Québec AI Institute(魁北克 AI 院)
;
Université de Montréal(蒙特利尔大学)
;
HEC Montréal(蒙特利尔 HEC 学院)
;
Intel AI Lab(英特尔 AI 实验室)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
我难以相信它不稳健:在嵌入漂移下安全分类器的灾难性崩溃
Subramanyam Sahoo, Vinija Jain, Divya Chaudhary, Aman Chadha
机构
*
Independent(独立研究者)
;
Meta AI
;
AWS Generative AI Innovation Center, Amazon Web Services(AWS生成式AI创新中心,亚马逊网络服务)
;
Northeastern University, Seattle, WA, USA(东北大学,西雅图,华盛顿州,美国)
;
Stanford University(斯坦福大学)
CommentsEqual Contribution: Xiaochuang Yuan and Hui Xu contributed equally to this work. All correspondence should be directed to yxc20098@gmail.com. Submitted to Agents in the Wild Workshop, ICLR2026
Mental Models of Autonomy and Sentience Shape Reactions to AI
自主性与意识的内心模型影响对AI的反应
Janet V. T. Pauketat, Daniel B. Shank, Aikaterina Manoli, Jacy Reese Anthis
机构
*
Sentience Institute(意识研究所)
;
Missouri University of Science and Technology(密苏里科技大学)
;
Max Planck Institute for Human Cognitive and Brain Sciences(人类认知与脑科学Max Planck研究所)
;
Stanford University(斯坦福大学)
;
University of Chicago(芝加哥大学)
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
ThinkMorph:多模态交错链式推理中的涌现特性
Jiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li, Michael Qizhe Shieh, Yejin Choi, Ranjay Krishna, Yu Cheng
机构
*
National University of Singapore(新加坡国立大学)
;
Zhejiang University(浙江大学)
;
University of Washington(华盛顿大学)
;
Stanford University(斯坦福大学)
;
absolute AI
;
The Chinese University of Hong Kong(香港中文大学)
Digital Companionship: Overlapping Uses of AI Companions and AI Assistants
数字陪伴:AI陪伴与AI助手的重叠使用
Aikaterina Manoli, Janet V. T. Pauketat, Ali Ladak, Hayoun Noh, Angel Hsing-Chi Hwang, Jacy Reese Anthis
机构
*
Max Planck Institute for Human Cognitive and Brain Sciences(人类认知与脑科学研究所)
;
Sentience Institute(意识研究所)
;
University of Edinburgh(爱丁堡大学)
;
University of Oxford(牛津大学)
;
University of Southern California(南加州大学)
;
Stanford University(斯坦福大学)
机构
*
School of Electrical and Computer Engineering(电气与计算机工程学院)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Department of Electrical Engineering(电气工程系)
;
Stanford University(斯坦福大学)
FMIP: Joint Continuous-Integer Flow For Mixed-Integer Linear Programming
FMIP: 混合整数线性规划的联合连续-整数流
Hongpei Li, Hui Yuan, Han Zhang, Jianghao Lin, Dongdong Ge, Mengdi Wang, Yinyu Ye
机构
*
Shanghai University of Finance and Economics(上海财经大学)
;
Princeton University(普林斯顿大学)
;
National University of Singapore(国立新加坡大学)
;
Antai College of Economics and Management(经济管理学院)
;
Shanghai Institute for Mathematics and Interdisciplinary Sciences(上海数学与交叉科学研究院)
;
Stanford University(斯坦福大学)
CommentsAccepted at the International Conference on Learning Representations (ICLR), 2025. A generative framework for MILP that jointly models integer and continuous variables, achieving 41% primal gap reduction with broad solver compatibility
Suppressing Prior-Comparison Hallucinations in Radiology Report Generation via Semantically Decoupled Latent Steering
通过语义解耦潜在引导抑制放射科报告生成中的先验比较幻觉
Ao Li, Rui Liu, Mingjie Li, Sheng Liu, Lei Wang, Xiaodan Liang, Lina Yao, Xiaojun Chang, Lei Xing
机构
*
University of New South Wales(新南威尔士大学)
;
Australian Artificial Intelligence Institute, University of Technology Sydney(澳大利亚人工智能研究所,技术悉尼大学)
;
Stanford University(斯坦福大学)
;
School of Computing and Information Technology of University of Wollongong Australia(沃林根澳大利亚大学计算与信息科技学院)
;
Sun Yat-sen University(中山大学)
;
University of Science and Technology of China(中国科学技术大学)
CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers
CMT-Benchmark:由专家研究人员构建的凝聚态理论基准
Haining Pan, James V. Roggeveen, Erez Berg, Juan Carrasquilla, Debanjan Chowdhury, Surya Ganguli, Federico Ghimenti, Juraj Hasik, Henry Hunt, Hong-Chen Jiang, Mason Kamb, Ying-Jer Kao, Ehsan Khatami, Michael J. Lawler, Di Luo, Titus Neupert, Xiaoliang Qi, Michael P. Brenner, Eun-Ah Kim
机构
*
Rutgers University(罗格斯大学)
;
Harvard University(哈佛大学)
;
Weizmann Institute of Science(魏茨曼科学研究所)
;
ETH Zürich(苏黎世联邦理工学院)
;
Cornell University(康奈尔大学)
;
Stanford University(斯坦福大学)
;
University of Zürich(苏黎世大学)
;
Stanford Institute for Materials and Energy Sciences(斯坦福材料与能源科学研究所)
;
SLAC National Accelerator Laboratory(斯坦福直线加速器实验室)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
National Taiwan University(台湾大学)
;
San José State University(圣何塞州立大学)