arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15729 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15729 篇

2103.08057 2021-03-16 cs.LG cs.AI cs.IR 73%

RecSim NG: Toward Principled Uncertainty Modeling for Recommender Ecosystems

Martin Mladenov, Chih-Wei Hsu, Vihan Jain, Eugene Ie, Christopher Colby, Nicolas Mayoraz, Hubert Pham, Dustin Tran, Ivan Vendrov, Craig Boutilier

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.05079 2021-03-10 cs.LG cs.AI 73%

Domain-Robust Visual Imitation Learning with Mutual Information Constraints

Edoardo Cetin, Oya Celiktutan

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG

Comments Presented at ICLR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.05672 2021-01-22 cs.LG cs.AI cs.MA 73%

Imitating Interactive Intelligence

Josh Abramson, Arun Ahuja, Iain Barr, Arthur Brussee, Federico Carnevale, Mary Cassin, Rachita Chhaparia, Stephen Clark, Bogdan Damoc, Andrew Dudzik, Petko Georgiev, Aurelia Guy, Tim Harley, Felix Hill, Alden Hung, Zachary Kenton, Jessica Landon, Timothy Lillicrap, Kory Mathewson, Soňa Mokrá, Alistair Muldal, Adam Santoro, Nikolay Savinov, Vikrant Varma, Greg Wayne, Duncan Williams, Nathaniel Wong, Chen Yan, Rui Zhu

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.07844 2021-01-21 cs.LG cs.AI cs.SY eess.SY 73%

Scalable Optimization for Wind Farm Control using Coordination Graphs

Timothy Verstraeten, Pieter-Jan Daems, Eugenio Bargiacchi, Diederik M. Roijers, Pieter J. K. Libin, Jan Helsen

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.05507 2021-01-15 cs.LG cs.AI cs.HC cs.MA 73%

Evaluating the Robustness of Collaborative Agents

Paul Knott, Micah Carroll, Sam Devlin, Kamil Ciosek, Katja Hofmann, A. D. Dragan, Rohin Shah

专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.09729 2020-12-16 cs.AI cs.LG 73%

Mastering Complex Control in MOBA Games with Deep Reinforcement Learning

Deheng Ye, Zhao Liu, Mingfei Sun, Bei Shi, Peilin Zhao, Hao Wu, Hongsheng Yu, Shaojie Yang, Xipeng Wu, Qingwei Guo, Qiaobo Chen, Yinyuting Yin, Hao Zhang, Tengfei Shi, Liang Wang, Qiang Fu, Wei Yang, Lanxiao Huang

专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.LG

Comments AAAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.10134 2020-12-15 cs.CV cs.AI cs.LG 73%

m2caiSeg: Semantic Segmentation of Laparoscopic Images using Convolutional Neural Networks

Salman Maqbool, Aqsa Riaz, Hasan Sajid, Osman Hasan

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG

Comments 16 pages, 5 figures, Code available at: https://github.com/salmanmaq/segmentationNetworks, Dataset available at: https://www.kaggle.com/salmanmaq/m2caiseg

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.08299 2020-12-10 cs.AI cs.LG 73%

GLIB: Efficient Exploration for Relational Model-Based Reinforcement Learning via Goal-Literal Babbling

Rohan Chitnis, Tom Silver, Joshua Tenenbaum, Leslie Pack Kaelbling, Tomas Lozano-Perez

专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG

Comments AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.02403 2020-11-05 cs.AI cs.LG cs.RO 73%

IDE-Net: Interactive Driving Event and Pattern Extraction from Human Data

Xiaosong Jia, Liting Sun, Masayoshi Tomizuka, Wei Zhan

专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.11952 2020-10-02 cond-mat.mes-hall cs.AI cs.LG cs.RO 73%

Autonomous robotic nanofabrication with reinforcement learning

Philipp Leinen, Malte Esders, Kristof T. Schütt, Christian Wagner, Klaus-Robert Müller, F. Stefan Tautz

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG

Comments 3 figures

Journal ref Sci. Adv. 6, eabb6987 (2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.09453 2020-09-29 cs.LG cs.AI cs.GT cs.MA 73%

OpenSpiel: A Framework for Reinforcement Learning in Games

Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Pérolat, Sriram Srinivasan, Finbarr Timbers, Karl Tuyls, Shayegan Omidshafiei, Daniel Hennes, Dustin Morrill, Paul Muller, Timo Ewalds, Ryan Faulkner, János Kramár, Bart De Vylder, Brennan Saeta, James Bradbury, David Ding, Sebastian Borgeaud, Matthew Lai, Julian Schrittwieser, Thomas Anthony, Edward Hughes, Ivo Danihelka, Jonah Ryan-Davis

专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.13221 2020-09-01 cs.LG cs.AI cs.HC cs.RO stat.ML 73%

Human-in-the-Loop Methods for Data-Driven and Reinforcement Learning Systems

Vinicius G. Goecks

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG

Comments PhD thesis, Aerospace Engineering, Texas A&M (2020). For more information, see https://vggoecks.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.04105 2020-08-11 cs.NI cs.AI cs.LG cs.MA 73%

Distributed Deep Reinforcement Learning for Functional Split Control in Energy Harvesting Virtualized Small Cells

Dagnachew Azene Temesgene, Marco Miozzo, Deniz Gündüz, Paolo Dini

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

Comments Submitted to IEEE transaction on sustainable computing. arXiv admin note: text overlap with arXiv:1906.05735

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.06306 2020-02-18 cs.LG cs.AI cs.MA stat.ML 73%

Jelly Bean World: A Testbed for Never-Ending Learning

Emmanouil Antonios Platanios, Abulhair Saparov, Tom Mitchell

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

Comments Published as a conference paper at ICLR 2020

Journal ref International Conference on Learning Representations 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.01562 2019-11-06 cs.LG cs.AI cs.RO 73%

DeepRacer: Educational Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning

Bharathan Balaji, Sunil Mallya, Sahika Genc, Saurabh Gupta, Leo Dirac, Vineet Khare, Gourav Roy, Tao Sun, Yunzhe Tao, Brian Townsend, Eddie Calleja, Sunil Muralidhara, Dhanasekar Karuppasamy

专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.01401 2019-09-18 cs.LG cs.AI cs.RO stat.ML 73%

Unsupervised Emergence of Egocentric Spatial Structure from Sensorimotor Prediction

Alban Laflaquière, Michael Garcia Ortiz

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG

Comments 27 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.11788 2019-07-30 cs.LG cs.AI stat.ML 73%

On Hard Exploration for Reinforcement Learning: a Case Study in Pommerman

Chao Gao, Bilal Kartal, Pablo Hernandez-Leal, Matthew E. Taylor

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

Comments AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE) 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.06508 2019-07-16 cs.AI cs.LG stat.ML 73%

General Board Game Playing for Education and Research in Generic AI Game Learning

Wolfgang Konen

专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.LG

Comments 8 pages, for: Conference on Games (CoG), London, 2019. Index Terms: game learning, general game playing, AI, temporal difference learning, board games, n-tuple systems

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.08809 2019-06-24 cs.LG cs.AI stat.ML 73%

A Deep Reinforcement Learning Approach for Global Routing

Haiguang Liao, Wentai Zhang, Xuliang Dong, Barnabas Poczos, Kenji Shimada, Levent Burak Kara

专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG

Comments Preprint submitted to ASME JMD

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.03967 2019-06-11 cs.LG cs.AI cs.NE cs.RO stat.ML 73%

Autonomous Goal Exploration using Learned Goal Spaces for Visuomotor Skill Acquisition in Robots

Adrien Laversanne-Finot, Alexandre Péré, Pierre-Yves Oudeyer

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.03765 2019-02-12 cs.LG cs.AI cs.CV cs.RO stat.ML 73%

Latent Space Reinforcement Learning for Steering Angle Prediction

Qadeer Khan, Torsten Schön, Patrick Wenzel

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.05098 2018-11-07 cs.LG cs.AI cs.NE 73%

DiCE: The Infinitely Differentiable Monte-Carlo Estimator

Jakob Foerster, Gregory Farquhar, Maruan Al-Shedivat, Tim Rocktäschel, Eric P. Xing, Shimon Whiteson

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.05172 2017-11-30 cs.LG cs.AI cs.HC cs.RO stat.ML 73%

Emotion in Reinforcement Learning Agents and Robots: A Survey

Thomas M. Moerland, Joost Broekens, Catholijn M. Jonker

专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.LG

Comments To be published in Machine Learning Journal

Journal ref Machine Learning 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1304.2024 2014-03-18 cs.LG cs.AI cs.MA stat.ML 73%

A General Framework for Interacting Bayes-Optimally with Self-Interested Agents using Arbitrary Parametric Model and Model Prior

Trong Nghia Hoang, Kian Hsiang Low

专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG

Comments 23rd International Joint Conference on Artificial Intelligence (IJCAI 2013), Extended version with proofs, 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24010 2026-07-28 cs.LG 新提交 72%

When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost

主动检索生成式人工智能何时应进行检索?效用、校准和成本的预算感知评估

Pin Qian, Su Wang, Chong Peng, Junxian You, Lifei Liu, Haoran Yu, Yihang Chen, Xiaochong Jiang

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Glasgow(格拉斯哥大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 Agent评测 :agentic(abstract,abstract_cn);分类 cs.LG

AI总结 研究主动检索生成式人工智能何时检索,通过将其重述为效用估计进行预算感知评估,并分离出相关三个问题,利用多种方法实现,在多数据集和模型中验证,强调评估应报告多方面指标。

Comments Accepted at the ACM SIGKDD KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI; 7 pages, 1 figure, and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23975 2026-07-28 cs.AI q-bio.QM 新提交 72%

Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks

柏拉图生物:通过时间再发现和结构基准进行验证优先的生物新奇性筛选

Stefan G. Creadore

机构 * Praxa Labs(普拉克斯实验室)

专题命中 Agent评测 :agent(abstract,comments);workflow(abstract);分类 cs.AI

AI总结 研究开发柏拉图生物,扩展开放架构并结合多种功能,修复评估缺陷。通过Python套件等验证其有效性,评估两个用例,如历史再发现任务及蛋白质结构比较,提供可重复软件契约和筛选基准,为生物研究提供支持。

Comments 16 pages, 6 figures, 3 tables. Companion code and data: https://github.com/Eldergenix/Plato-Scientific-Research-Autonomous-Agent. This fork-specific validation study cites, but does not duplicate, arXiv:2510.26887

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.08986 2026-07-21 cs.AI cs.LO math-ph math.AP math.MP 版本更新 72%

A Formalization of the Mean-Field Derivation of the Vlasov Equation

弗拉索夫方程平均场推导的形式化:作为策略游戏的人工智能辅助精益形式化

Joseph K. Miller

机构 * Stanford University(斯坦福大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 Agent评测 :agent(abstract,comments);AI agent(abstract);分类 cs.AI

AI总结 该研究以数学家指导AI在Lean 4中形式化研究成果为案例,将其构建为形式化游戏。通过此方式对非线性弗拉索夫方程适定性完整形式化,展示了开发过程及成果,还介绍了最优传输机制的分离情况及开发时间等,为形式化研究提供了新方法。

Comments 26 pages, 4 figures. Lean 4 development, blueprint site, and agent logs: https://github.com/Hydrodynamical/Vlasov_Meanfield_Formalization

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11619 2026-07-16 cs.AI 版本更新 72%

When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents

当智能体与自身意见相左:测量基于LLM的智能体的行为一致性

Aman Mehta

机构 * Aman Mehta

专题命中 Agent评测 :agentic(abstract,comments);agent(abstract);分类 cs.AI

AI总结 研究发现基于LLM的智能体在相同任务上运行结果不一致,且这种不一致与任务成功率密切相关,通过监控行为一致性可提升智能体可靠性。

Comments Accepted at the ICML 2026 Workshop on Statistical Frameworks for Uncertainty in Agentic Systems. 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11672 2026-06-11 cs.CR cs.AI 新提交 72%

Can Open-Source LLM Agents Replace Static Application Security Testing Tools? An Empirical Assessment

开源LLM代理能否取代静态应用安全测试工具?一项实证评估

Derek Yohn, Luke Flancher, Mirajul Islam, Khaled Slhoub

机构 * College of Engineering and Science, Florida Institute of Technology(工程学院与科学学院,佛罗里达理工学院)

专题命中 Agent评测 :agentic(abstract,comments);agent(abstract);分类 cs.AI

AI总结 评估基于开源LLM的代理在静态应用安全测试中的性能,与SAST工具Bandit对比,发现当前不适合实际应用。

Comments Keywords: Agentic AI, Cybersecurity, Large Language Models, Static Application Security Testing, Model performance evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04229 2025-10-20 cs.HC cs.AI 72%

When AI Gets Persuaded, Humans Follow: Inducing the Conformity Effect in Persuasive Dialogue

Rikuo Sasaki, Michimasa Inaba

机构 * The University of Electro-Communications(电子通信大学)

专题命中 Agent评测 :agent(abstract,comments);AI agent(abstract);分类 cs.AI

Comments 23 pages, 19 figures. International Conference on Human-Agent Interaction (HAI 2025), November 10-13, 2025, Yokohama, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏