Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It
为什么多步工具使用强化学习会崩溃以及监督信号如何修复它
Yupu Hao, Zhuoran Jin, Huanxuan Liao, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统认知与决策智能重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
机构
*
The Hong Kong Polytechnic University(香港理工大学)
;
Nanyang Technological University(南洋理工大学)
;
Tsinghua University(清华大学)
;
Royal Melbourne Institute of Technology(皇家墨尔本理工大学)
;
Beijing University of Aeronautics and Astronautics(北京航空航天大学)
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents
智能体技能可能有害:LLM智能体中技能诱导故障的实证研究
Gen Dong, Yanjie Gao, Liqun Li, Tianyin Xu, Yu Hua, Fan Yang
机构
*
Huazhong University of Science and Technology(华中科技大学)
;
Microsoft Research(微软研究院)
;
Microsoft(微软)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
E-Bench:在现实世界产品场景中对多步工具使用智能体进行基准测试
Weihuang Zheng, Tianyuan Zou, Eileen Ye, Alphet Liu, Youyong Kong, Ya-Qin Zhang, Duran Zheng, Maxm Pan
机构
*
Hunyuan Team, Tencent(腾讯混元团队)
;
Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp
UNIBROWSE:用于多模态浏览比较的数据到智能体框架
Xiyu Wei, Qingwei Zong, Zhuocheng Yu, Sujian Li
机构
*
Key Laboratory of Computational Linguistics, MOE, Peking University(教育部计算语言学重点实验室,北京大学)
;
School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)
;
School of Computer Science, Peking University(北京大学计算机科学学院)
Bridging the Agent-World Gap: Text World Models for LLM-based Agents
弥合智能体-世界鸿沟:面向基于LLM的智能体的文本世界模型
Yixia Li, Hongru Wang, Peng Lai, Zhiwen Ruan, He Zhu, Youxin Zhu, Ganlong Zhao, Minda Hu, Yun Chen, Sibei Yang, Peng Li, Jeff Z. Pan, Jia Pan, Guanhua Chen, Yang Liu, Guanbin Li
机构
*
Southern University of Science and Technology(南方科技大学)
;
University of Edinburgh(爱丁堡大学)
;
Peking University(北京大学)
;
Sun Yat-sen University(中山大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai University of Finance and Economics(上海财经大学)
;
Tsinghua University(清华大学)
;
The University of Hong Kong(香港大学)
Comments18 pages, 3 figures. Published at the Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD) and the Workshop on Failure Modes of Agentic AI (FAGEN) at ICML 2026