MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
MCPEvol-Bench:跨MCP服务器动态演化对大语言模型智能体性能进行基准测试
Huanxi Liu, Kun Hu, Jiaqi Liao, Qiang Wang, Pengfei Qian, YuanZhao Zhai, Dawei Feng, Bo Ding, Huaimin Wang
机构
*
College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机科学与技术学院)
;
State Key Laboratory of Complex & Critical Software Environment(复杂关键软件环境国家重点实验室)
;
National Key Laboratory of Parallel and Distributed Computing(并行与分布式计算国家重点实验室)
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
SEED:用于智能体强化学习的自进化在线策略蒸馏
Jinyang Wu, Shuo Yang, Zhengxi Lu, Fan Zhang, Yuhao Shen, Lang Feng, Haoran Luo, Zheng Lian, Shuai Zhang, Zhengqi Wen, Jianhua Tao
机构
*
Tsinghua University(清华大学)
;
Zhejiang University(浙江大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Nanyang Technological University(南洋理工大学)
;
Tongji University(同济大学)
机构
*
School of Software Technology, Zhejiang University(浙江大学软件学院)
;
State Key Laboratory of CAD&CG, Zhejiang University(浙江大学CAD&CG国家重点实验室)
;
Stomatology Hospital, Zhejiang University School of Medicine(浙江大学医学院附属口腔医院)
Comments43 pages, 9 figures. Oral presentation at the Robotics: Science and Systems 2026 Workshop on Foundation Models for Robot Planning (FM4RoboPlan)
Catch, Throw, Repeat: Planning for Human-Robot Partner Juggling
接住、抛出、重复:人机搭档杂耍的规划
Jonathan Rainer Lippert, Kai Ploeger, Abir Chowdhury, Hermann Müller, Jan Peters, Alap Kshirsagar
机构
*
Technical University of Darmstadt(达姆施塔特工业大学)
;
Justus-Liebig University Gießen(吉森尤斯 - 李比希大学)
;
German Research Center for AI (DFKI)(德国人工智能研究中心(DFKI))
;
Robotics Institute Germany(德国机器人研究所)
;
Indian Institute of Technology Delhi - Abu Dhabi(印度理工学院德里分校 - 阿布扎比)
Yining Xing, Zhiyuan Liu, Zehong Ke, Wenhao Yu, Jianqiang Wang
机构
*
School of Vehicle and Mobility, Tsinghua University(清华大学车辆与运载学院)
;
State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University(清华大学智能网联汽车与交通国家重点实验室)
CommentsPresented at the 2nd Causal Neuro-symbolic Artificial Intelligence (Causal NeSy): Toward Agentic LLMs with Neuro-Symbolic and Graph Based Reasoning Workshop @ ESWC2026
Observational planning for the 2026 August 5 Falcon 9 Upper Stage lunar impact
2026年8月5日猎鹰9号上面级月球撞击的观测规划
Benjamin Fernando, Jennifer Heldmann, Bill Grey, John Ortiz, Bryan Euser, Darryl Z. Seligman, Eunhyeuk Kim, Anthony Colaprete, Elisa Maria Alessi, Detlef Koschny, Anthony Cook, Joel Green, Patrick King, Stacy Teng, Dawn Graninger, Arnold Goldberg, William Cooke, Mike F. Skrutskie, Kevin Schlaufman, Nicholas Schmerr, Carly M. Donahue, Carl A. Schmidt, Nancy J. Chanover
机构
*
Department of Architecture, National University of Singapore, Singapore 117566, Singapore
;
Cambridge Centre for Advanced Research
;
Sustainable Design Group, Department of Architecture, University of Cambridge, Cambridge, United Kingdom
;
Department of City
;
Regional Planning, University of Pennsylvania, Philadelphia, PA 19104, USA
;
Urban Analytics Subject Group, Urban Studies \& Social Policy Division, University of Glasgow
;
Laboratory for Earth Surface Processes, Ministry of Education, College of Urban
;
Environmental Sciences, Peking University, Beijing 100871, China
CXRAgent: Director-Orchestrated Multi-Stage Reasoning for Chest X-Ray Interpretation
CXRAgent:用于胸部X光解读的由主任编排的多阶段推理
Jinhui Lou, Yan Yang, Zhou Yu, Zhenqi Fu, Weidong Han, Qingming Huang, Jun Yu
机构
*
School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机科学学院)
;
Department of Automation, Tsinghua University(清华大学自动化系)
;
Department of Colorectal Medical Oncology, Zhejiang Cancer Hospital(浙江省肿瘤医院结直肠医学肿瘤科)
;
School of Computer and Control Engineering, University of Chinese Academy of Sciences(中国科学院大学计算机与控制工程学院)
;
School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院)
Comments13 pages. Extended version of the paper accepted at IEEE SMC-IT/SCC 2026 (Space Mission Challenges for Information Technology / Space Computing Conference); adds a unified GPU+CPU DRA quota and a scheduler-level accelerator fallback re-validated on Kueue v0.18.3
CommentsExtended version of a paper accepted at the 2026 IEEE Conference on Decision and Control (CDC). Contains complete proofs, background material, and full experimental details