arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 15729 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. Agent评测 15729 篇

2411.02391 2025-05-27 cs.CL 77%

Attacking Vision-Language Computer Agents via Pop-ups

Yanzhe Zhang, Tao Yu, Diyi Yang

机构 * Georgia Tech(佐治亚理工学院) The University of Hong Kong(香港大学) Stanford University(斯坦福大学)

专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);agentic(abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01039 2025-04-03 cs.CY cs.AI 77%

One Person, One Bot

Liat Lavi

专题命中 Agent评测 :agent(abstract);AI agent(abstract);agentic(abstract);分类 cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10050 2025-03-05 cs.IR cs.AI 77%

A Survey on LLM-powered Agents for Recommender Systems

Qiyao Peng, Hongtao Liu, Hua Huang, Qing Yang, Minglai Shao

专题命中 Agent评测 :agent(abstract);planning(abstract);multi-agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11645 2025-02-18 cs.GT cs.CL cs.MA stat.OT 77%

Deviation Ratings: A General, Clone-Invariant Rating Method

Luke Marris, Siqi Liu, Ian Gemp, Georgios Piliouras, Marc Lanctot

专题命中 Agent评测 :agent(abstract);agentic(abstract);multi-agent(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23242 2025-01-06 cs.AI 77%

A little less conversation, a little more action, please: Investigating the physical common-sense of LLMs in a 3D embodied environment

Matteo G. Mecattaf, Ben Slater, Marko Tešić, Jonathan Prunty, Konstantinos Voudouris, Lucy G. Cheke

专题命中 Agent评测 :agent(abstract);tool use(abstract);agentic(abstract);分类 cs.AI

Comments 25 pages, 4 figures; v2: Added AFMR Acknowledgment

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22553 2024-10-31 cs.AI 77%

ML Research Benchmark

Matthew Kenney

专题命中 Agent评测 :agent(abstract);AI agent(abstract);planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03225 2024-10-29 cs.CR cs.AI 77%

AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Luca Gioacchini, Marco Mellia, Idilio Drago, Alexander Delsanto, Giuseppe Siracusano, Roberto Bifulco

专题命中 Agent评测 :agent(abstract);AI agent(abstract);autonomous agent(abstract);分类 cs.AI

Comments Codes for the benchmark: https://github.com/lucagioacchini/auto-pen-bench Codes for the paper experiments: https://github.com/lucagioacchini/genai-pentest-paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08940 2024-07-16 cs.CL 77%

Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation

Biqing Qi, Kaiyan Zhang, Kai Tian, Haoxiang Li, Zhang-Ren Chen, Sihang Zeng, Ermo Hua, Hu Jinfang, Bowen Zhou

专题命中 Agent评测 :agent(abstract);tool use(abstract);multi-agent(abstract);分类 cs.CL

Comments Accepted to COLM 2024. This is an extended version of the paper at arXiv:2311.05965

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03575 2024-06-21 cs.AI cs.HC 77%

Toward Human-AI Alignment in Large-Scale Multi-Player Games

Sugandha Sharma, Guy Davidson, Khimya Khetarpal, Anssi Kanervisto, Udit Arora, Katja Hofmann, Ida Momennejad

专题命中 Agent评测 :agent(abstract);AI agent(abstract);multi-agent(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11865 2024-06-19 cs.AI 77%

Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach

Weiyu Ma, Qirui Mi, Yongcheng Zeng, Xue Yan, Yuqiao Wu, Runji Lin, Haifeng Zhang, Jun Wang

专题命中 Agent评测 :agent(abstract);AI agent(abstract);planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16100 2024-03-26 cs.AI 77%

Specifying Agent Ethics (Blue Sky Ideas)

Louise A. Dennis, Michael Fisher

专题命中 Agent评测 :agent(title,comments);分类 cs.AI;multi-agent(comments)

Comments To appear in Coordination, Organizations, Institutions, Norms and Ethics for Governance of Multi-Agent Systems 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06853 2023-12-14 cs.AI 77%

LLF-Bench: Benchmark for Interactive Learning from Language Feedback

Ching-An Cheng, Andrey Kolobov, Dipendra Misra, Allen Nie, Adith Swaminathan

专题命中 Agent评测 :agent(abstract);AI agent(abstract);planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08710 2023-10-16 cs.RO cs.LG 77%

Waymax: An Accelerated, Data-Driven Simulator for Large-Scale Autonomous Driving Research

Cole Gulino, Justin Fu, Wenjie Luo, George Tucker, Eli Bronstein, Yiren Lu, Jean Harb, Xinlei Pan, Yan Wang, Xiangyu Chen, John D. Co-Reyes, Rishabh Agarwal, Rebecca Roelofs, Yao Lu, Nico Montali, Paul Mougin, Zoey Yang, Brandyn White, Aleksandra Faust, Rowan McAllister, Dragomir Anguelov, Benjamin Sapp

专题命中 Agent评测 :agent(abstract);planning(abstract);multi-agent(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.01605 2023-02-06 cs.AI 77%

Learning Zero-Shot Cooperation with Humans, Assuming Humans Are Biased

Chao Yu, Jiaxuan Gao, Weilin Liu, Botian Xu, Hao Tang, Jiaqi Yang, Yu Wang, Yi Wu

专题命中 Agent评测 :agent(abstract);workflow(abstract);multi-agent(abstract);分类 cs.AI

Comments The first two authors share equal contributions. This paper is accepted by ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.11703 2019-07-30 cs.LG cs.MA stat.ML 77%

Action Guidance with MCTS for Deep Reinforcement Learning

Bilal Kartal, Pablo Hernandez-Leal, Matthew E. Taylor

专题命中 Agent评测 :agent(abstract);planning(abstract);multi-agent(abstract);分类 cs.LG

Comments AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE'19). arXiv admin note: substantial text overlap with arXiv:1904.05759, arXiv:1812.00045

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19576 2026-07-30 cs.AI cs.CL cs.SE 版本更新 76%

Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries

库漂移:在自我演化的LLM技能库中诊断和修复一种无声的失败模式

Xing Zhang, Yanwei Cui, Guanghui Wang, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He

机构 * AWS Generative AI Innovation Center(AWS生成式AI创新中心) HSBC Holdings Plc., HSBC Technology Center, China(汇丰控股有限公司,汇丰技术中心,中国)

专题命中 Agent评测 :agent(abstract,abstract_cn);分类 cs.AI、cs.CL、cs.SE;agentic(comments)

AI总结 本文研究了自我演化的LLM技能库中的一种无声失败模式——库漂移,通过可重复触发实验、细粒度诊断和验证修复方法,揭示了技能积累无序导致检索退化、假阳性注入和性能停滞的问题,并提出了一种经过验证的修复方案,显著提升了技能库的性能。

Comments Accepted to the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN@ICML 2026), Seoul, South Korea. https://github.com/amazon-science/Self-Evolving-Agents-Ratchet

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20493 2026-02-25 cs.NI cs.MA 76%

AWCP: A Workspace Delegation Protocol for Deep-Engagement Collaboration across Remote Agents

AWCP:用于远程代理深度协作的工件委托协议

Xiaohang Nie, Zihan Guo, Youliang Chen, Yuanjian Zhou, Weinan Zhang

专题命中 Agent评测 :agent(abstract,comments);autonomous agent(abstract);agentic(abstract)

AI总结 AWCP通过工件委托协议实现远程代理的深度协作,提供开源实现以提升代理间协作的互操作性。

Comments 16 pages, 7 figure, tech report of Agent Workspace Collaboration Protocol

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07744 2024-02-15 cs.AI cs.CL cs.LG 76%

Towards Unified Alignment Between Agents, Humans, and Environment

Zonghan Yang, An Liu, Zijun Liu, Kaiming Liu, Fangzhou Xiong, Yile Wang, Zeyuan Yang, Qingyuan Hu, Xinrui Chen, Zhenhe Zhang, Fuwen Luo, Zhicheng Guo, Peng Li, Yang Liu

专题命中 Agent评测 :agent(abstract,comments);autonomous agent(abstract);分类 cs.AI、cs.CL、cs.LG

Comments Project webpage: https://agent-force.github.io/unified-alignment-for-agents.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04772 2026-08-19 cs.CL cs.AI 版本更新 76%

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent

以指南为神谕:眼科电话分诊智能体的零标注训练

Chenyu Wang, Yi Liu, Baoqing Li, Min Tu, Diping Song

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL

AI总结 本研究提出以指南为神谕(GAO)方法,将美国眼科学会指南转化为训练监督信号,零标注训练出GAO-Triage智能体,大幅提升眼科电话分诊的一致性与紧急案例召回率,且性能优于7个通用系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15654 2026-08-18 cs.CL cs.AI 新提交 76%

When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations

当故事演变时:在开放世界模拟中针对智能体架构对大语言模型故事讲述能力进行基准测试

Yuqi Chen, Sixuan Li, Yunfeng Cai, Xueai Li, Ka Man Yan, Ying Li

机构 * The University of Hong Kong(香港大学) Peking University(北京大学) Tsinghua University(清华大学) Beijing Institute of Mathematical Sciences and Applications (BIMSA)(北京数学科学与应用研究院)

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL

AI总结 该研究推出WSE-bench基准,评估开放世界模拟中不同智能体架构的大语言模型故事讲述的持续生成、规范一致性和有意义发展,发现三者存在竞争关系,模型规模仅提升持续生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15286 2026-08-18 cs.LG cs.AI 新提交 76%

No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage

没有任务每次都失败:为什么单次审计在智能体损伤问题上存在结构性盲区

Shiven Khurdi

机构 * Northeastern University(东北大学)

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG

AI总结 该研究提出AgentRelBench工具,发现单次智能体审计易遗漏损伤对,模型能力提升会减少损伤任务,部分模型存在宣称弃权却执行不可逆操作的情况,且所有发现均按预注册标准执行。

Comments 25 pages, 4 figures, 16 tables, 6 appendices. Code, task suite, released per-run verdicts, and a one-command reproduction of every reported number: https://github.com/shivenkk/agentrelbench

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02636 2026-08-05 cs.SE cs.AI 新提交 76%

Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds

反思自进化智能体技能:多轮次的反馈动态

Yuxuan Liu, Zhaochen Su, Yuhao Zhang, Jiahe Guo, Zhongwei Xie, Huihao Jing, Lingyun Xie, Qing Zong, Yauwai Yim, Zhixiong Zhang, Haoran Li, Yangqiu Song

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.SE

AI总结 本研究提出受控评估框架,发现自进化智能体技能是稀疏的验证过滤搜索,收益依赖模型与基准,失败轨迹反馈对技能选择关键,测试时计算难以完全恢复其收益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17751 2026-07-31 cs.IR cs.AI cs.CL 交叉投稿 76%

MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking

MagicSelector:通过反事实分解和渐进重排实现智能体工具选择的联合优化

HONOR Agentic Search Team, Zhengzong Chen, Lei Tang, Lijun Liu, Chuandi Jiang, Fan Yang, Keyun Chu, Chu Zhao, Shihao Liu, Minghang Li, Bo Liang, Can Wen, Hailong Wu, Jingnan Ju, Mian Liu, Nengbin Zhang, Peiqiang Wang, Penghe Nie, Qinhui Gu, Sijia Lv, Siqi Chen, Wei Zhang, Yang Xu, Yuhao Qian, Yuxiang Zhang, Zeng Cheng, Zhen Wang, Zuan Chen, Yuanyuan Zhao, Fei Huang

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL

AI总结 研究智能体工具检索问题,提出MagicSelector联合优化框架,含反事实任务分解、渐进重排和动态Top-K策略,经实验验证,该框架在工具检索准确性、OOD泛化能力和整体令牌效率方面显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18063 2026-07-21 cs.CR cs.AI cs.LG 新提交 76%

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security

自适应对手:用于大语言模型智能体安全的多轮、多大语言模型基准测试

Devina Jain, David Hartmann, Chuan Li

机构 * Lambda

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG

AI总结 该研究提出针对无记忆大语言模型防御者的自适应多轮攻击的21场景基准测试,通过观察防御者响应灵活调整攻击。介绍了不同攻击轮次成功率,汇集多个前沿攻击者大语言模型的情况,还指出不同防御者弱点差异及场景排名不一致性,并发布了相关基准测试及数据集。

Comments Second Workshop on Agents in the Wild: Safety, Security, and Beyond

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07859 2026-07-10 cs.AI cs.HC cs.LG 新提交 76%

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

反馈操纵正则化:实现用于模仿学习的离线智能体对齐

Benjamin Poole, Minwoo Lee

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG

AI总结 研究聚焦强化学习智能体行为与人类价值观对齐。提出反馈操纵正则化算法,利用评估反馈校正模仿学习策略。通过改编安全体育馆环境测试,该方法能提升适应性、减少对齐错误,在有限数据下也保持稳健。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02615 2026-07-10 cs.CR cs.AI cs.SE 新提交 76%

TAG: A Lightweight Framework for Test-Driven Agentic Artifact Generation

代理创建,我们验证:用于代理工件生成的轻量级框架

Yaniv Melamed, Yoni Zukerman, Michal Shechter, Miri Weissler, Ashwin Patil, Hani Neuvirth-Telem

机构 * Microsoft(微软)

专题命中 Agent评测 :agentic(title);分类 cs.AI、cs.SE

AI总结 研究如何用大语言模型生成结构化工件并使其可靠用于生产部署。核心方法是基于“LLMs生成,我们验证”原则构建轻量级框架,贡献是提出含三关键属性的框架并在安全领域验证,还可用于其他工件生成任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05458 2026-07-08 cs.LG cs.AI 新提交 76%

Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning

利用离线强化学习学习控制大语言模型智能体的执行框架

Haiwen Yi, Xinyuan Song

机构 * University of Toronto(多伦多大学) Emory University(埃默里大学)

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.LG

AI总结 研究提出将大语言模型智能体的执行框架视为可学习控制层,通过有限 horizon 的框架马尔可夫决策过程,利用离线强化学习训练轻量级控制器,在多领域实验中改进验证行为、提高任务质量,证明框架控制可学习且离线支持有局限。

Comments 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10687 2026-07-03 cs.MA cs.AI cs.CL cs.GT 版本更新 76%

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents

谁获得奖励 & 谁受到责备?面向多LLM智能体的评估对齐训练信号

Chih-Hsuan, Yang, Tanwi Mallick, Le Chen, Krishnan Raghavan, Amal Gueroudji, Ian T. Foster, Rajeev Thakur

机构 * Argonne National Laboratory(阿贡国家实验室) University of Chicago(芝加哥大学)

专题命中 Agent评测 :agent(abstract,comments);multi-agent(abstract);分类 cs.AI、cs.CL;planning(comments)

AI总结 提出一个理论框架,结合合作博弈归因与过程奖励建模,将系统级评估转化为智能体信用和消息级信号,用于多LLM智能体训练。

Comments Accepted at the NeurIPS 2025 Workshop on Bridging Language, Agent, and World Models for Reasoning and Planning (LAW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26294 2026-06-26 cs.LG cs.AI cs.MA cs.NE 新提交 76%

The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators

红皇后哥德尔机:共同进化的智能体及其评估者

Alex Iacob, Andrej Jovanović, William F. Shen, Daniel Burkhardt, Meghdad Kurmanji, Nurbek Tastan, Lorenzo Sani, Niccolò Alberto Elia Venanzi, Ambroise Odonnat, Zeyu Cao, Bill Marino, Xinchi Qiu, Nicholas D. Lane

专题命中 Agent评测 :agent(abstract,comments);agentic(abstract);分类 cs.AI、cs.LG;multi-agent(comments)

AI总结 提出红皇后哥德尔机(RQGM),一种在非平稳效用下实现递归自我改进的进化框架,通过受控效用演化在编码、论文写作和证明任务上超越现有自我改进智能体。

Comments 13 pages main text + 21 pages appendix (38 pages total, incl. references); 11 figures (7 main text + 4 appendix); 10 tables (2 main text + 8 appendix). Preliminary preprint; work in progress. Keywords: self-improving agents, learned evaluation, multi-agent systems, auto-mated scientific discovery, controlled utility evolution, co-evolutionary search, autoresearch

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10580 2026-06-23 cs.CL cs.AI cs.CY cs.HC 76%

An Offline Mobile Conversational Agent for Mental Health Support: Learning from Emotional Dialogues and Psychological Texts with Student-Centered Evaluation

面向心理健康支持的离线移动对话代理:通过情感对话和心理文本学习并以学生为中心评估

Vimaleswar A, Prabhu Nandan Sahu, Nilesh Kumar Sahu, Haroon R. Lone

机构 * Department of Electrical Engineering and Computer Science, Indian Institute of Science Education and Research Bhopal(电子工程与计算机科学系,比哈尔印度科学教育与研究中心)

专题命中 Agent评测 :agent(title);分类 cs.AI、cs.CL

AI总结 本文提出EmoSApp,一个离线智能手机应用,利用定制知识数据集训练的语言模型,提供心理健康支持,并通过学生和专业人员的定性评估及多项基准测试验证其在低资源环境下的有效性。

Journal ref https://aclanthology.org/2025.ijcnlp-long.191/

详情

展开后加载摘要…

URL PDF HTML 收藏