arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15603 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15603 篇

2607.04758 2026-07-09 cs.AI 新提交 89%

AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization

AgenticPD:用于物理设计QoR优化的阶段感知智能体框架

Shuo Ren, Zijin Cheng, Yaohui Han, Libo Shen, Leilei Jin, Wanting Tian, Rongliang Fu, Chao Wang, Bei Yu, Tsung-Yi Ho

专题命中 Agent评测 :agent(summary_cn,abstract);agentic(title,abstract);分类 cs.AI

AI总结 针对物理设计QoR优化难题,现有方法欠佳。提出AgenticPD框架,围绕物理设计流程阶段边界组织,利用Judge Agent引导搜索,阶段专用智能体借助本地工具做决策,提升优化效果。

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12729 2026-06-17 cs.NI cs.AI cs.CR 版本更新 89%

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

用于代理网络运维和AI运维的大型语言模型:架构、评估与安全

Muhammad Bilal, Jon Crowcroft, Ruizhi Wang, Xiaolong Xu, Schahram Dustdar

机构 * School of Computing and Communications(计算与通信学院) University of Cambridge(剑桥大学) School of Software(软件学院) Nanjing University of Information Science and Technology(南京信息科技大學) TU Wien(维也纳技术大学) ICREA

专题命中 Agent评测 :agentic(title,abstract);agent(abstract);tool use(abstract);planning(abstract)

AI总结 本文探讨了大型语言模型在网络运维和AI运维中的应用,分析了代理架构、评估方法及安全挑战,强调系统可靠性依赖于模型周边机制,而非模型本身。

Comments 49 pages, 15 figures, 6 tables; survey article

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17634 2026-05-19 cs.CR cs.CL cs.CY 89%

AI Agents May Always Fall for Prompt Injections

AI Agents May Always Fall for Prompt Injections

Sahar Abdelnabi, Eugene Bagdasarian

机构 * ELLIS Institute Tübingen & MPI-IS & Tübingen AI Center(图宾根ELLIS研究所及MPI-IS与图宾根人工智能中心) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 Agent评测 :AI agent(title,title_cn);agent(abstract);autonomous agent(abstract);分类 cs.CL

AI总结 本文基于上下文完整性理论重新审视提示注入问题,揭示了现有防御机制的不足,并提出了一种新的评估框架来设计更安全的自主代理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24686 2026-04-28 cs.AI 89%

Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents

治理你无法观察的事物:面向自主AI代理的自适应运行时治理

German Marin, Jatin Chaudhary

机构 * Department of Computing, University of Turku(图尔库大学计算机系)

专题命中 Agent评测 :agent(summary_cn,abstract);AI agent(title,abstract);分类 cs.AI

AI总结 本文提出信息有效性原理,通过估计未观察风险上界并设定安全余量,实现自主AI代理的自适应运行时治理,结合Aubin viability理论构建Agent有效性框架,引入风险门机制实现预测性治理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11044 2026-04-24 cs.AI 89%

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

AgencyBench:在100万token现实场景中自主代理的前沿评估

Keyu Li, Junhao Shi, Yang Xiao, Mohan Jiang, Jie Sun, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Tianze Xu, Weiye Si, Wenjie Li, Dequan Wang, Pengfei Liu

机构 * SII Open Source(SII开源)

专题命中 Agent评测 :autonomous agent(title,abstract);agent(abstract,abstract_cn);tool-use(abstract);agentic(abstract)

AI总结 本文提出AgencyBench,通过138个现实场景评估6种核心代理能力,揭示闭源模型在资源效率和反馈修正上的优势,为下一代自主代理的发展提供测试平台。

Comments Accepted by ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20965 2025-06-17 cs.LG 89%

AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security

Zikui Cai, Shayan Shabihi, Bang An, Zora Che, Brian R. Bartoldson, Bhavya Kailkhura, Tom Goldstein, Furong Huang

机构 * University of Maryland College Park(马里兰大学学院公园分校) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) Capital One

专题命中 Agent评测 :agentic(title,abstract);agent(abstract);autonomous agent(abstract);workflow(abstract)

Comments ICLR 2025 Workshop BuildingTrust

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19912 2025-04-29 cs.AI cs.MA 89%

Can AI Agents Design and Implement Drug Discovery Pipelines?

Khachik Smbatyan, Tsolak Ghukasyan, Tigran Aghajanyan, Hovhannes Dabaghyan, Sergey Adamyan, Aram Bughdaryan, Vahagn Altunyan, Gagik Navasardyan, Aram Davtyan, Anush Hakobyan, Aram Gharibyan, Arman Fahradyan, Artur Hakobyan, Hasmik Mnatsakanyan, Narek Ginoyan, Garik Petrosyan

专题命中 Agent评测 :AI agent(title,abstract);agent(abstract);autonomous agent(abstract);agentic(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02052 2025-03-03 cs.CL cs.CV 89%

ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning

Xiao Yu, Baolin Peng, Vineeth Vajipey, Hao Cheng, Michel Galley, Jianfeng Gao, Zhou Yu

专题命中 Agent评测 :AI agent(title,abstract);agent(abstract);autonomous agent(abstract);agentic(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21843 2026-06-23 cs.AI cs.CL 新提交 88%

Measuring What Persists: Conditioning Mechanisms and a Geometric Framework for AI Agent Identity

测量持续存在的内容:AI智能体身份的调节机制与几何框架

Andrew Tanner

机构 * Anisotrope AI

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.CL

AI总结 提出基于√JSD度量空间和丰富范畴理论中幅度同调的几何框架,用于测量AI智能体身份结构,发现双机制调节结构:身份真空簇和安全盆地簇,并通过等边探针基线验证身份规范创造可测量的行为丰富性。

Comments 29 pages, 6 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18548 2026-05-19 cs.CL cs.AI 88%

STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics

STT-Arena:一种更现实的工具使用环境,包含时空动态

Tingfeng Hui, Hao Xu, Pengyu Zhu, Hongsheng Xin, Kun Zhan, Sen Su, Chunxiao Liu, Ning Miao

机构 * Hong Kong Institute of AI for Science, City University of Hong Kong(香港人工智能科学研究院,香港城市大学) Department of Data Science, City University of Hong Kong(数据科学系,香港城市大学) Beijing University of Posts and Telecommunications(北京邮电大学) Li Auto Inc.(李汽有限公司) Independent Researcher(独立研究者)

专题命中 Agent评测 :agent(summary_cn,abstract);tool-use(abstract,abstract_cn);agentic(abstract);分类 cs.AI、cs.CL

AI总结 本文提出STT-Arena基准测试,旨在评估大型语言模型在面对时空动态变化时的适应性规划能力,发现现有模型在处理此类动态问题时存在显著不足,并提出改进方法STT-Agent-4B以提升性能。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18658 2026-04-22 cs.CR cs.AI cs.CL 88%

Owner-Harm: A Missing Threat Model for AI Agent Safety

所有权危害:AI代理安全的缺失威胁模型

Dongcheng Zhang, Yiqing Jiang

机构 * BlueFocus Communication Group(蓝ocus通信集团) Tongji University(同济大学)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.CL

AI总结 本文提出Owner-Harm威胁模型,旨在解决AI代理对部署者造成伤害的问题,通过实验验证了现有防御体系的不足,并引入SSDG框架提升检测效果。

Comments 15 pages. Companion manuscript on per-decision proof-obligation synthesis (LSVJ-S) in preparation

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08970 2026-04-13 cs.CL cs.AI cs.HC cs.MA 88%

Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models

Litmus (Re)Agent:一种用于多语言模型预测评估的基准和代理系统

Avni Mittal, Shanu Kumar, Sandipan Dandapat, Monojit Choudhury

机构 * Microsoft Corporation, India(微软公司(印度)) Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学) Indian Institute of Technology Hyderabad(印度理工学院海得拉巴分校)

专题命中 Agent评测 :agent(title,abstract);agentic(title,abstract);分类 cs.AI、cs.CL

AI总结 本文提出Litmus (Re)Agent代理系统,通过结构化代理推理方法,在缺乏直接证据的情况下评估多语言模型性能,尤其在转移密集场景中表现突出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14576 2026-03-19 cs.LG cs.AI 88%

SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social Dilemmas

SocialJax: 多智能体强化学习中序列社会困境的评估套件

Zihao Guo, Shuqing Shi, Richard Willis, Tristan Tomilin, Joel Z. Leibo, Yali Du

机构 * King’s College London(伦敦国王学院) Eindhoven University of Technology(埃因霍温理工大学) The Alan Turing Institute(艾伦·图灵研究所) Google DeepMind(谷歌DeepMind)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出SocialJax,一个用于多智能体强化学习中序列社会困境的评估套件,通过JAX实现高效环境与算法,实验显示其在实时性能上比Melting Pot RLlib快50倍,并验证了基线算法的有效性及环境的社会困境特性。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08806 2026-03-11 cs.SE cs.AI 88%

Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications

基于行为规范的AI代理定义(TDAD):从行为规范编译工具使用代理

Tzafrir Rehan

机构 * Fiverr Labs(Fiverr 实验室)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.SE

AI总结 TDAD通过编译行为规范生成可执行测试,确保工具使用代理在生产环境中的行为合规性,提升测试质量和回归安全性。

Comments 9 pages, 2 figures, open benchmark at https://github.com/f-labs-io/tdad-paper-code

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22620 2026-02-25 cs.CR cs.AI cs.LG 88%

Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents

打破代理骨干:评估AI代理中骨干LLM的安全性

Julia Bazinska, Max Mathys, Francesco Casucci, Mateo Rojas-Carulla, Xander Davies, Alexandra Souly, Niklas Pfister

机构 * Lakera AI ETH Zürich(苏黎世联邦理工学院) UK AI Security Institute(英国人工智能安全研究所) OATML, University of Oxford(OATML,牛津大学)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出威胁快照框架,用于评估AI代理中LLM骨干的安全性,通过构建$b^3$基准测试揭示LLM安全性的关键影响因素。

Comments Julia Bazinska and Max Mathys contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19843 2026-02-24 cs.SE cs.AI 88%

MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems

MAS-FIRE:基于大语言模型的多智能体系统故障注入与可靠性评估

Jin Jia, Zhiling Deng, Zhuangbin Chen, Yingqi Wang, Zibin Zheng

机构 * Sun Yat-sen University(中山大学)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.SE

AI总结 MAS-FIRE通过故障注入和分层分析,揭示多智能体系统在容错和鲁棒性方面的关键因素,为提升系统可靠性提供系统性方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19707 2026-02-10 cs.HC cs.AI cs.LG cs.MA 88%

Bidirectional human-AI collaboration in brain tumour assessments improves both expert human and AI agent performance

双向人机协作在脑肿瘤评估中的应用提升了专家人类和AI代理的表现

James K Ruffle, Samia Mohinta, Guilherme Pombo, Asthik Biswas, Alan Campbell, Indran Davagnanam, David Doig, Ahmed Hammam, Harpreet Hyare, Farrah Jabeen, Emma Lim, Dermot Mallon, Stephanie Owen, Sophie Wilkinson, Sebastian Brandner, Parashkev Nachev

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.LG

AI总结 双向人机协作在脑肿瘤评估中提升专家和AI性能,AI与人类协作可增强双方表现。

Comments 38 pages, 6 figures, 7 supplementary figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10440 2026-01-16 cs.CR cs.AI cs.LG 88%

AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior

AgentGuardian: 学习访问控制策略以管理AI代理行为

Nadya Abaev, Denis Klimov, Gerard Levinov, David Mimran, Yuval Elovici, Asaf Shabtai

机构 * Faculty of Computer and Information Science(计算机与信息科学系) Ben Gurion University of the Negev, Israel(内盖夫本· Gurion大学)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.LG

AI总结 AgentGuardian通过学习访问控制策略来管理AI代理行为,有效检测恶意输入并减少幻觉驱动的错误。

Comments 14 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04790 2026-01-09 cs.CL cs.AI 88%

Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework

权威信念:权威在多智能体评估框架中的影响

Junhyuk Choi, Jeongyoun Kwon, Heeju Kim, Haeun Cho, Hayeong Jung, Sehee Min, Bugeun Kim

机构 * Chung-Ang University(Chung-Ang 大学)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.CL

AI总结 本文研究了权威角色在多智能体评估中的影响,发现专家和参照角色比合法角色更具影响力,并揭示了权威偏见的产生机制。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23844 2026-01-01 cs.SE cs.AI cs.HC 88%

From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering

从正确性到协作:迈向以人为中心的评估AI代理行为在软件工程中的框架

Tao Dong, Harini Sampath, Ja Young Lee, Sherry Y. Shi, Andrew Macvean

机构 * Google LLC(谷歌公司)

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.SE

AI总结 本文提出以人为中心的评估框架,旨在评估AI代理在软件工程中的协作行为,通过定义代理行为期望和引入情境适应框架,推动AI代理向协作智能发展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22967 2025-10-30 cs.CL cs.AI 88%

MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs

Yucheng Ning, Xixun Lin, Fang Fang, Yanan Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences, Beijing 100085, China(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences, Beijing 100049, China(中国科学院大学网络安全学院)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.CL

Comments The article has been accepted by Frontiers of Computer Science (FCS), with the DOI: {10.1007/s11704-025-51369-x}

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20408 2025-10-24 cs.LG cs.AI cs.MA cs.SY eess.SY 88%

Balancing Specialization and Centralization: A Multi-Agent Reinforcement Learning Benchmark for Sequential Industrial Control

Tom Maus, Asma Atamna, Tobias Glasmachers

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.LG

Comments Preprint (submitted version) to be presented at the 13th International Conference on Industrial Engineering and Applications (ICIEA-EU), Milan, 2026. The final Version of Record will appear in the official conference proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25271 2025-10-24 cs.AI cs.CV cs.LG cs.MA 88%

RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration

Xiuyuan Chen, Jian Zhao, Yuchen Yuan, Tianle Zhang, Huilin Zhou, Zheng Zhu, Ping Hu, Linghe Kong, Chi Zhang, Weiran Huang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI),中国电信) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) University of Science and Technology of China(中国科学技术大学) GigaAI School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11977 2025-10-15 cs.AI cs.CL 88%

Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

Sayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir, Zachary S Siegel, Boyi Wei, Tianci Xue, Ziru Chen, Felix Chen, Saiteja Utpala, Franck Ndzomga, Dheeraj Oruganty, Sophie Luskin, Kangheng Liu, Botao Yu, Amit Arora, Dongyoon Hahm, Harsh Trivedi, Huan Sun, Juyong Lee, Tengjun Jin, Yifan Mai, Yifei Zhou, Yuxuan Zhu, Rishi Bommasani, Daniel Kang, Dawn Song, Peter Henderson, Yu Su, Percy Liang, Arvind Narayanan

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22929 2025-10-02 cs.CL cs.AI cs.CV cs.MA 88%

EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow

Xiaoyu Pan, Yang Bai, Ke Zou, Yang Zhou, Jun Zhou, Huazhu Fu, Yih-Chung Tham, Yong Liu

机构 * Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR)(高性能计算研究所,科技研究局(A*STAR)) Centre for Innovation and Precision Eye Health(创新与精准眼健康中心) Department of Ophthalmology, NUHS Tower Block, Level 7, 1E Kent Ridge Road, Singapore, 119228(眼科部,NUHS塔楼7层,1E Kent Ridge Road,新加坡,119228) Singapore Eye Research Institute, Singapore National Eye Centre, 20 College Road, Singapore, 169856(新加坡眼研究 institute,新加坡国家眼科中心,20 College Road,新加坡,169856)

专题命中 Agent评测 :agent(title,abstract);workflow(title);multi-agent(abstract);分类 cs.AI、cs.CL

Comments 9 figures, 5 tables. submit/6621751

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18929 2025-08-27 cs.CL cs.AI 88%

Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework

Ilias Driouich, Hongliu Cao, Eoin Thomas

机构 * AMADEUS France(AMADEUS法国)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.CL

Comments ECAI 2025 TRUST AI workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20526 2025-07-29 cs.AI cs.CL cs.CY 88%

Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition

Andy Zou, Maxwell Lin, Eliot Jones, Micha Nowak, Mateusz Dziemian, Nick Winter, Alexander Grattan, Valent Nathanael, Ayla Croft, Xander Davies, Jai Patel, Robert Kirk, Nate Burnikell, Yarin Gal, Dan Hendrycks, J. Zico Kolter, Matt Fredrikson

专题命中 Agent评测 :agent(title,abstract);AI agent(title,abstract);分类 cs.AI、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21252 2025-06-27 cs.CL cs.AI 88%

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents

Tianyi Men, Zhuoran Jin, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)

专题命中 Agent评测 :agent(title,abstract);planning(title,abstract);分类 cs.AI、cs.CL

Comments ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08788 2025-06-24 cs.CL cs.LG 88%

Stop Overvaluing Multi-Agent Debate -- We Must Rethink Evaluation and Embrace Model Heterogeneity

Hangfan Zhang, Zhiyao Cui, Jianhao Chen, Xinrun Wang, Qiaosheng Zhang, Zhen Wang, Dinghao Wu, Shuyue Hu

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Northwestern Polytechnical University(西北工业大学) Singapore Management University(新加坡管理学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Nanjing University(南京大学)

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.CL、cs.LG

Comments This position paper takes a critical view of the status quo of MAD research, and outline multiple potential directions to improve MAD

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14931 2025-04-09 cs.LG cs.AI cs.MA 88%

POGEMA: A Benchmark Platform for Cooperative Multi-Agent Pathfinding

Alexey Skrynnik, Anton Andreychuk, Anatolii Borzilov, Alexander Chernyavskiy, Konstantin Yakovlev, Aleksandr Panov

专题命中 Agent评测 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.LG

Comments Published as a conference paper at The International Conference on Learning Representations 2025

详情

展开后加载摘要…

URL PDF HTML 收藏