arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

2026-04-29 至 2026-04-29 共收录 126 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 规划决策 36 篇

2604.25085 2026-04-29 cs.GT cs.AI cs.CY 57%

Optimally Auditing Adversarial Agents

最优审计对抗性智能体

Sanmay Das, Fang-Yi Yu, Yuang Zhang

机构 * Virginia Tech(弗吉尼亚理工大学) George Mason University(乔治·梅森大学)

专题命中 规划决策 :agent(abstract);分类 cs.AI

AI总结 本文提出一个主代理博弈模型,研究如何设计最优审计策略以应对智能体的欺诈行为,提出适应性和非适应性设置下的高效算法,并扩展至有限审计预算场景。

Comments Published in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2026, pages 16787-16794

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 2026, pages 16787-16794

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24913 2026-04-29 cs.LG q-bio.PE 57%

Generative diffusion models for spatiotemporal influenza forecasting

生成扩散模型用于时空流感预测

Joseph Lemaitre, Justin Lessler

机构 * Department of Epidemiology, Gillings School of Global Public Health, University of North Carolina at Chapel Hill(流行病学系,全球公共卫生学院,北卡罗来纳大学查佩尔希尔分校) Department of Epidemiology, Johns Hopkins Bloomberg School of Public Health, Baltimore, MD 21205, USA(流行病学系,约翰·霍普金斯伯恩斯坦公共卫生学院,巴尔的摩,MD 21205, USA) Carolina Population Center, University of North Carolina at Chapel Hill(卡罗来纳人口中心,北卡罗来纳大学查佩尔希尔分校)

专题命中 规划决策 :planning(abstract);分类 cs.LG

AI总结 本文提出Influpaint模型,利用扩散模型对流感流行病进行时空预测,通过将流感季节编码为时空图像,学习疾病动态的丰富分布,并在回顾评估中实现了与领先集成方法相当的预测精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24790 2026-04-29 cs.CR cs.AI 57%

Semantic Denial of Service in LLM-controlled robots

语义拒绝服务在LLM控制的机器人中

Jonathan Steinberg, Oren Gal

机构 * Swarms & AI Lab (SAIL) University of Haifa(群集与人工智能实验室(SAIL)海法大学)

专题命中 规划决策 :agent(abstract);分类 cs.AI

AI总结 本文研究了LLM控制机器人中语义拒绝服务攻击,发现通过注入安全可信的短语可触发安全推理中断,提出防御策略需平衡攻击抑制与真实危险响应,强调系统架构而非提示级别的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25690 2026-04-29 cs.CE 50%

Data Driven Calibration of Analytical Concrete Creep Models Considering Preloading Effects Using Gaussian Processes

基于数据驱动的分析性混凝土蠕变模型校准:考虑预加载效应的高斯过程方法

Leonie Heller, Christopher Taube, Gledson Rodrigo Tondo, Guido Morgenthal

专题命中 规划决策 :planning(abstract)

AI总结 本文利用高斯过程回归校准混凝土蠕变模型,考虑预加载强度、时机和混凝土龄期的影响,提升模型精度并量化不确定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08323 2026-04-29 cs.CV 50%

Detecting Dental Landmarks from Intraoral 3D Scans: the 3DTeethLand challenge

从口内3D扫描中检测牙齿标志:3DTeethLand挑战

Achraf Ben-Hamadou, Nour Neifar, Ahmed Rekik, Oussama Smaoui, Firas Bouzguenda, Sergi Pujades, Niels van Nistelrooij, Shankeeth Vinayahalingam, Kaibo Shi, Hairong Jin, Youyi Zheng, Tibor Kubík, Oldřich Kodym, Petr Šilling, Kateřina Trávníčková, Tomáš Mojžiš, Jan Matula, Jeffry Hartanto, Xiaoying Zhu, Kim-Ngan Nguyen, Tudor Dascalu, Huikai Wu, and Weijie Liu, Shaojie Zhuang, Guangshun Wei, Yuanfeng Zhou

机构 * Department of Oral Maxillofacial Surgery, Radboud University Medical Center, Geert Grooteplein Zuid 10, 6525 GA Nijmegen, Netherlands organization= State Key Lab of CAD\&CG, Zhejiang University, Hangzhou, 310058 , country= China organization= Department of Computer Graphics Multimedia, Brno University of Technology , city= Brno , country= Czech Republic text= National Dental Centre Singapore organization= Guangxi Colleges Universities Key Laboratory of Intelligent Software ,country= Wuzhou University text= National University of Singapore organization= Department of Computer Science , country = University of Copenhagen organization= School of Software, Shandong University , country= China Centre de Recherche en Num\' e rique de Sfax, Laboratory of Signals, Systems, Artificial Intelligence Inria, Univ. Grenoble Alpes, CNRS, Grenoble INP, LJK, France

专题命中 规划决策 :planning(abstract)

AI总结 本文提出3DTeethLand挑战,通过公开数据集评估牙齿标志检测算法,推动临床应用。挑战引入340个口内3D扫描数据,49支队伍参与,最终6支进入决赛,展示高精度与召回率的平衡方法。

Comments MICCAI 2024, 3DTeethLand, Challenge report, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07446 2026-04-29 cs.CR 50%

Electric Democracy: Proof of Work to secure Elections

电子民主:用工作证明保障选举

Vitaly Zuevsky

专题命中 规划决策 :agent(abstract)

AI总结 本文提出利用工作证明算法与其它安全机制构建零信任选举系统,以解决电子投票的安全问题,提升选举透明度与参与度。

Comments Novel application of cryptographic primitives for electronic voting. Discussion is invited

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25447 2026-04-29 physics.ao-ph 50%

Evaluating local climate in global storm-resolving models with the Köppen-Geiger classification

利用柯本-盖格尔分类评估全球风暴解析模型的局部气候

Chiel C. van Heerwaarden, Menno A. Veerman, Imme Benedict, Lukas Brunner, Edgar Dolores-Tesillos, Emanuel Dutra, Erich Fischer, Junhong Lee, Olivia Martius, Xabier Pedruzo-Bagazgoitia, Ulrike Proske, Sarah N. Warnau, Jonathan D. Willie, Cathy Hohenegger

专题命中 规划决策 :planning(abstract)

AI总结 研究评估了两种全球风暴解析模型在柯本-盖格尔气候分类中的表现,发现模型在区域偏倚方面仍有不足,但整体上在许多地区表现良好,提出该分类作为标准诊断工具。

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25393 2026-04-29 math.OC 50%

An uncertainty model for positive-valued parameters with application to robust optimization

正值参数的不确定性模型及其在稳健优化中的应用

Tatsuya Tanaka, Huimin Li, Shota Yamanaka, Ellen H. Fukuda, Nobuo Yamashita

专题命中 规划决策 :planning(abstract)

AI总结 本文提出一种正值参数的不确定性集模型,以解决传统方法可能包含非正值的问题,通过凸函数度量参数变化并提供可计算的对偶形式,应用于光伏-电池运营规划和支持向量机问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25323 2026-04-29 cs.RO 50%

ANCHOR: A Physically Grounded Closed-Loop Framework for Robust Home-Service Mobile Manipulation

ANCHOR:一种基于物理的闭环框架,用于鲁棒的家庭服务移动操作

Jinhao Jiang, Shengyu Fang, Sibo Zuo, Yujie Tang, Yirui Li

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 规划决策 :planning(abstract)

AI总结 ANCHOR通过物理 grounding 和结构化故障处理,提升家庭服务机器人在动态环境中的任务成功率和扰动恢复能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25016 2026-04-29 eess.SY cs.SY 50%

A Novel Two-Step Approach for Reactive Power Demand Calculation Using Integrated Voltage Stability Analysis

一种基于整合电压稳定性分析的反应功率需求计算新两步方法

Hassan Abouelgheit, Hendrik Lens

专题命中 规划决策 :planning(abstract)

AI总结 本文提出一种新两步方法,通过整合电压稳定性分析计算反应功率需求,结合迭代准动态仿真、Q-V分析和动态仿真,全面评估电压稳定性,最终通过案例验证方法有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24907 2026-04-29 cs.LO cs.RO 50%

Logic of Fuzzy Paths

模糊路径的逻辑

Kush Grover, Pratham Gupta, Jan Křetínský

机构 * Indian Institute of Science(印度科学研究院) Masaryk University(马萨里克大学)

专题命中 规划决策 :planning(abstract)

AI总结 本文提出一种新的时间逻辑,用于运动规划中的规范。该逻辑基于信号时间逻辑,但以路径为基本元素,简化了公式并提升了对行为偏好的表达能力,适用于人类指定的规范和演示学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09981 2026-04-29 cs.CV cs.RO 50%

ReSim: Reliable World Simulation for Autonomous Driving

ReSim:面向自动驾驶的可靠世界模拟

Jiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen, Yuqian Shao, Xiaosong Jia, Hongyang Li, Andreas Geiger, Xiangyu Yue, Li Chen

机构 * The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) OpenDriveLab at Shanghai AI Lab(上海人工智能实验室OpenDrive实验室) NVIDIA Research(NVIDIA研究) Xiaomi EV(小米电动车) Shanghai Jiao Tong University(上海交通大学) University of Tübingen, Tübingen AI Center(图宾根大学,图宾根人工智能中心) HKUST(香港科技大学)

专题命中 规划决策 :planning(abstract)

AI总结 本文提出ReSim模型,通过整合真实数据与模拟数据提升自动驾驶场景模拟的可靠性与可控性,提升规划和策略选择性能。

Comments NeurIPS 2025 Spotlight. Project page: https://opendrivelab.com/ReSim

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多智能体 23 篇

2604.09618 2026-04-29 cs.DC cs.AI cs.CR 90%

HearthNet: Edge Multi-Agent Orchestration for Smart Homes

HearthNet:边缘多智能体协调用于智能家居

Zhonghao Zhan, Krinos Li, Yefan Zhang, Hamed Haddadi

机构 * Imperial College London(伦敦帝国学院) Independent Researcher(独立研究员)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);planning(abstract);分类 cs.AI

AI总结 HearthNet通过边缘多智能体系统解决智能家居中自然语言控制、设备故障和持续协调的问题,采用MQTT、Git共享状态和授权租赁实现设备管理。

Comments (CAIS 2026) Proceedings of the ACM Conference on AI and Agentic Systems, Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25737 2026-04-29 cs.SE cs.AI 88%

SAFEdit: Does Multi-Agent Decomposition Resolve the Reliability Challenges of Instructed Code Editing?

SAFEdit:多智能体分解能否解决受指导代码编辑的可靠性挑战?

Noam Tarshish, Nofar Selouk, Daniel Hodisan, Bar Ezra Gafniel, Yuval Elovici, Asaf Shabtai, Eliya Nachmani

机构 * Ben-Gurion University of the Negev(巴伊兰大学)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.SE

AI总结 SAFEdit通过多智能体框架分解编辑过程,提升代码编辑的可靠性与准确性,其迭代改进机制显著提高了任务成功率。

Comments Accepted to the EQUISA (Evaluation of Qualitative Aspects of Intelligent Software Assistants) workshop at EASE (Evaluation and Assessment in Software Engineering) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25088 2026-04-29 cs.AI cs.CL 88%

Cooperate to Compete: Strategic Coordination in Multi-Agent Conquest

合作以竞争:多智能体征服中的战略协调

Abigail O'Neill, Alan Zhu, Mihran Miroyan, Narges Norouzi, Joseph E. Gonzalez

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.CL

AI总结 本文提出C2C多智能体环境,研究在混合动机设定下智能体如何通过合作达成长期竞争目标,发现人类与AI在谈判行为上有显著差异,并通过优化提升AI胜率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25268 2026-04-29 cs.CL cs.AI 88%

CRAFT: Grounded Multi-Agent Coordination Under Partial Information

CRAFT:在部分信息下的 grounded 多智能体协调

Abhijnan Nath, Hannah VanderHoeven, Nikhil Krishnaswamy

机构 * Department of Computer Science, Colorado State University(计算机科学系,科罗拉多州立大学)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI、cs.CL

AI总结 CRAFT 是评估大语言模型在严格部分信息下实用沟通能力的多智能体基准,通过自然语言构建共享3D结构。研究发现更强推理能力不必然带来更好协调,小模型常表现更优,表明多智能体协调仍是当前语言模型的挑战。

Comments Added revisions, corrected typos and additional analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25264 2026-04-29 cs.CR cs.SE 88%

MARD: A Multi-Agent Framework for Robust Android Malware Detection

MARD:一种用于鲁棒Android恶意软件检测的多智能体框架

Xueying Zeng, Youquan Xian, Sihao Liu, Xudong Mou, Yanze Li, Lei Cui, Bo Li

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.SE

AI总结 MARD通过结合LLM的语义理解和传统静态分析,提出多智能体框架,有效降低APK深度分析成本,实现高可解释性检测,F1得分达93.46%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24881 2026-04-29 cs.AI 88%

Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate

潜在代理:一种事后训练过程用于内部化多代理辩论

John Seon Keun Yi, Aaron Mueller, Dokyun Lee

机构 * Boston University(波士顿大学)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 本文提出一种两阶段微调流程,将多代理辩论转化为单个LLM,通过动态奖励调度和长度截断实现内部化,减少93%的token使用,同时通过激活引导发现代理特定子空间,并展示通过内部化辩论抑制恶意代理的应用。

Comments ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24808 2026-04-29 cs.MA cs.AI cs.CY cs.DC 88%

ITAS: A Multi-Agent Architecture for LLM-Based Intelligent Tutoring

ITAS:基于大语言模型的智能辅导系统多智能体架构

Iizalaarab Elhaimeur, Nikos Chrisochoides

机构 * Center for Real-Time Computing(实时计算中心) Computer Science Department(计算机科学系) Old Dominion University(旧 Dominion 大学) Physics Departments(物理系)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 本文提出ITAS多智能体架构,用于解决大语言模型在真实课程中运行的挑战,通过三层架构实现教学、操作和反馈功能,展示了系统在实际应用中的表现。

Comments Companion papers: arXiv:Q-ID (Quantum deployment), arXiv:L-ID (Latency analysis)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24477 2026-04-29 cs.CR cs.AI cs.MA 88%

GAMMAF: A Common Framework for Graph-Based Anomaly Monitoring Benchmarking in LLM Multi-Agent Systems

GAMMAF:用于LLM多智能体系统基于图的异常监控基准测试的通用框架

Pablo Mateo-Torrejón, Alfonso Sánchez-Macián

机构 * University Carlos III of Madrid(卡洛斯三世大学)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract);分类 cs.AI

AI总结 本文提出GAMMAF框架,用于评估LLM多智能体系统中的异常检测方法,通过生成合成数据集和动态隔离攻击节点,验证了其在拓扑扩展性和执行效率上的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24971 2026-04-29 cs.LG cs.CL cs.DC 88%

PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference

PolyKV:一种用于多智能体LLM推理的共享非对称压缩KV缓存池

Ishan Patel, Ishan Joshi

机构 * Independent Researcher(独立研究者)

专题命中 多智能体 :agent(title,abstract);multi-agent(title,comments);分类 cs.CL、cs.LG

AI总结 PolyKV通过非对称压缩技术实现多个并发推理代理共享单一KV缓存池,提升内存效率并减少推理延迟。

Comments 10 pages, 6 tables. Code: https://github.com/ishan1410/PolyKV Keywords: KV cache compression, multi-agent LLM inference, asymmetric quantization, FWHT, TurboQuant, shared memory

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19016 2026-04-29 cs.HC 88%

CHORUS: Effort-Aware Multi-Agent Human-AI Collaboration for Professional Translation

CHORUS:面向专业翻译的有意识多智能体人机协作

George X. Wang, Jiaqian Hu, Guande Wu Jing Qian

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract)

AI总结 CHORUS通过多智能体协作系统支持翻译流程与个人风格,减少译者认知负担,提升翻译质量,降低完成时间33.8%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15292 2026-04-29 cs.MA 88%

Adversarial Attack on Black-Box Multi-Agent by Adaptive Perturbation

对黑盒多智能体系统的对抗性攻击:自适应扰动

Jianming Chen, Yawen Wang, Junjie Wang, Xiaofei Xie, Yuanzhe Hu, Qing Wang, Fanjiang Xu

专题命中 多智能体 :agent(title,abstract);multi-agent(title,abstract)

AI总结 本文提出AdapAM框架,通过自适应选择策略和基于代理的扰动诱导恶意行为,提升多智能体系统攻击的隐蔽性和有效性,实验表明其在不同扰动率下表现最佳。

Journal ref AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25040 2026-04-29 cs.AI cs.CL 84%

Leverage Laws: A Per-Task Framework for Human-Agent Collaboration

利用率法则:一种面向任务的人机协作框架

Stan Loosmore

专题命中 多智能体 :agent(title,abstract);planning(abstract);分类 cs.AI、cs.CL

AI总结 本文提出一种面向任务的利用率比率,用于衡量人机协作中人类工作被代理所替代的比例,结合信息需求的流向和时间成本,探讨了利用率的渐进行为及任务新颖性对规划的影响。

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24831 2026-04-29 cs.SE cs.LG 84%

FGDM: Reasoning Aware Multi-Agentic Framework for Software Bug Detection using Chain of Thought and Tree of Thought Prompting

FGDM:基于推理的多智能体框架用于软件缺陷检测的链式思维和树式思维提示

Srita Padmanabhuni, Bhargavi Karuturi, Jerusha Karen Indupalli, Santhan Reddy Chilla, Vivek Yelleti

机构 * SRM University, Andhra Pradesh(安得拉邦SRM大学)

专题命中 多智能体 :agentic(title);agent(abstract);multi-agent(abstract);分类 cs.LG、cs.SE

AI总结 本文提出FGDM框架,利用链式思维和树式思维提示,通过流程图识别代码错误并生成修复代码,有效提升大型代码库中缺陷检测的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25724 2026-04-29 cs.AI 82%

Scalable Inference Architectures for Compound AI Systems: A Production Deployment Study

可扩展的推理架构用于复合AI系统:一种生产部署研究

Srikanta Prasad S, Utkarsh Arora

机构 * Agentforce AI Platform, Salesforce India Pvt Ltd(Agentforce AI平台,Salesforce印度私人有限公司)

专题命中 多智能体 :agentic(abstract,comments);agent(abstract);AI agent(abstract);multi-agent(abstract)

AI总结 本文研究了一种模块化、平台无关的推理架构,用于支持复合AI应用,如Agentforce和ApexGuru,展示了其在降低延迟、提升吞吐量和节省成本方面的成效。

Comments Accepted to the ACM Conference on AI and Agentic Systems (ACM CAIS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25727 2026-04-29 cs.AI 81%

Toward Scalable Terminal Task Synthesis via Skill Graphs

通过技能图实现可扩展的终端任务合成

Zhiyuan Fan, Tinghao Yu, Yuanjun Cai, Jiangtao Guan, Yun Yang, Dingxin Hu, Jiang Zhou, Xing Wu, Zhuo Han, Feng Zhang, Lilin Wang

机构 * Hunyuan Team, Tencent(文生图团队,腾讯)

专题命中 多智能体 :agent(abstract);workflow(abstract);agentic(abstract);multi-agent(abstract)

AI总结 本文提出SkillSynth框架,通过场景介导的技能图生成多样化终端任务,提升训练效率和任务多样性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24807 2026-04-29 cs.CY cs.AI cs.MA 77%

From Prototype to Classroom: An Intelligent Tutoring System for Quantum Education

从原型到课堂:面向量子教育的智能辅导系统

Iizalaarab Elhaimeur, Nikos Chrisochoides

机构 * Center for Real-Time Computing(实时计算中心) Computer Science Department(计算机科学系) Old Dominion University(旧 Dominion 大学) Physics Departments(物理系)

专题命中 多智能体 :agent(abstract);planning(abstract);multi-agent(abstract);分类 cs.AI

AI总结 本文提出ITAS系统,通过多智能体架构和课程设计解决量子教育中的可靠性问题,验证系统在真实课堂中的可行性及教学分析能力。

Comments 10 pages, 6 figures, 1 table. Submitted to IEEE QCE 2026. Companion papers (in preparation): ITAS architecture and latency analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25832 2026-04-29 cs.AI 70%

TrialCalibre: A Fully Automated Causal Engine for RCT Benchmarking and Observational Trial Calibration

TrialCalibre:一种完全自动化的因果引擎用于RCT基准测试和观察性试验校准

Amir Habibdoust, Xing Song

专题命中 多智能体 :agent(abstract);workflow(abstract);分类 cs.AI

AI总结 本文提出TrialCalibre,一种多智能体系统,用于自动化和扩展BenchExCal流程,通过专门的智能体协调整个过程,实现适应性、可审计和透明的因果效应估计。

Comments 5 pages , 2 figures

Journal ref Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025. Copyright 2025 by the author(s)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25076 2026-04-29 cs.LG 70%

Zero Shot Coordination for Sparse Reward Tasks with Diverse Reward Shapings

稀疏奖励任务中具有多样化奖励塑造的零 shot 协调

Keenan Powell, Peihong Yu, Pratap Tokekar

机构 * Department of Computer Science(计算机科学系) University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 多智能体 :agent(abstract);multi-agent(abstract);分类 cs.LG

AI总结 本文提出了一种基于随机奖励塑造的多智能体强化学习方法,通过4种算法选择策略训练多个方法,提升在Overcooked环境中的稀疏奖励表现,改进率达62.2%-119.2%。

详情

展开后加载摘要…

URL PDF HTML 收藏