Explanation as Question Answering based on a Task Model of the Agent's Design
专题命中 Agent评测 :agent(title,abstract);AI agent(abstract);分类 cs.AI、cs.LG
Comments 7 Pages, 10 Figures, IJCAI Explainable AI Workshop
AI 大模型
智能体、工具调用、规划、工作流、多智能体和自主任务执行。
专题命中 Agent评测 :agent(title,abstract);AI agent(abstract);分类 cs.AI、cs.LG
Comments 7 Pages, 10 Figures, IJCAI Explainable AI Workshop
专题命中 Agent评测 :agent(title,abstract);AI agent(abstract);分类 cs.AI、cs.LG
Comments 11 Pages, 9 Figures
专题命中 Agent评测 :autonomous agent(title,abstract);agent(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :autonomous agent(title,abstract);agent(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(title,abstract);planning(abstract);分类 cs.AI、cs.LG
Comments ICML 2021, 12 pages, 7 figures
专题命中 Agent评测 :autonomous agent(title,abstract);agent(abstract);分类 cs.AI、cs.LG
Journal ref IEEE Transactions on Games ( Volume: 11 , Issue: 2 , June 2019 )
VAmoS Bench:语音智能体仿真基准
机构 * Veris AI(维瑞斯人工智能公司)
专题命中 Agent评测 :agent(title,abstract);分类 cs.AI
AI总结 该研究推出 VAmoS Bench 语音智能体仿真基准,针对现有基准无法评估语音智能体独立处理电话呼叫能力的空白,在金融服务场景中端到端评估完整语音智能体系统,支持动态排行榜。
Comments 12 pages, 3 figures, 2 tables. Agent implementations: https://github.com/veris-ai/riley-agent
ProofCouncil:用于解决开放性数学问题的语言模型智能体
机构 * ETH Zurich(苏黎世联邦理工学院) ; Aarhus University(奥胡斯大学) ; Leiden University(莱顿大学)
专题命中 Agent评测 :agent(title,abstract);agentic(abstract);分类 cs.AI
AI总结 研究旨在提升大型语言模型解决数学开放性问题的性能,介绍了ProofCouncil智能体,它采用作者-评论家架构,在FirstProof挑战及其他问题测试中表现出色,还介绍了其开发及相关构建库并开源。
Comments 25 pages, 7 figures. ProofCouncil appears as System A (IMProofBench ProofCouncil) in the official FirstProof second-batch report (arXiv:2606.18119). Code and agent-building library: https://github.com/eth-sri/proof-council
Anchor:缓解智能体基准生成中的工件漂移
机构 * Agentic Labs
专题命中 Agent评测 :agent(title,abstract);AI agent(abstract);分类 cs.AI;agentic(comments)
AI总结 提出Anchor管道,通过约束优化程序联合生成指令、环境、真值解和验证器,解决基准生成中的工件漂移问题,并构建ERP-Bench基准评估前沿模型性能。
Comments Accepted to RLEval '26 (Workshop at ACM Conference on AI and Agentic Systems 2026)
AnomalyClaw:通过工具引导的反驳实现通用视觉异常检测代理
机构 * Department of Computer Science and Engineering, Southern University of Science and Technology (SUSTech), Shenzhen, China(南方科技大学计算机科学与工程系,深圳,中国) ; School of EEE, Nanyang Technological University (NTU), Singapore(南洋理工大学电子工程学院,新加坡) ; CFAR, Agency for Science, Technology and Research (A*STAR), Singapore(科技研究局(A*STAR)的CFAR,新加坡)
专题命中 Agent评测 :agent(title,abstract);agentic(abstract);分类 cs.AI
AI总结 本文提出AnomalyClaw,一种无需训练的视觉异常检测代理,通过多轮反驳过程提升异常判断。在CrossDomainVAD-12基准上,AnomalyClaw在多个模型上均取得显著提升,且引入自演化扩展提升模型性能。
Comments We release the agent, the benchmark, and the analysis artifacts at https://github.com/jam-cc/AnomalyClaw
机构 * Department of Data Science and AI, IIT Madras, India(数据科学与人工智能系,印度理工学院马德拉斯学院) ; Department of Engineering Design, IIT Madras, India(工程设计系,印度理工学院马德拉斯学院) ; LoveForm Health Technologies, India(LoveForm健康科技公司,印度) ; Department of Radiology and Imaging Sciences, Sri Ramachandra Institute of Higher Education and Research, India(放射学与成像科学系, Sri Ramachandra高等教育与研究学院,印度) ; Department of Neuro and Interventional Radiology, Sri Ramachandra Institute of Higher Education and Research, India(神经放射学与介入放射学系,Sri Ramachandra高等教育与研究学院,印度)
专题命中 Agent评测 :agent(title,abstract);agentic(abstract,comments);分类 cs.AI
Comments Paper published at "Agentic AI for Medicine" Workshop, MICCAI 2025
Journal ref Lecture Notes in Computer Science, vol 16147, 2025. Springer, Cham
机构 * Tongyi Lab(通义实验室) ; Alibaba Group(阿里巴巴集团)
专题命中 Agent评测 :agentic(title,abstract);agent(abstract,comments);分类 cs.CL
Comments https://tongyi-agent.github.io/blog/introducing-tongyi-deep-research/
机构 * Centre for Digital Governance, Hertie School(数字治理中心,赫尔姆斯学校) ; Oxford Internet Institute, University of Oxford(牛津互联网研究所,牛津大学) ; Weizenbaum Institute Berlin, Germany(贝伦贝格Weizenbaum研究所) ; Technical University Munich, Germany(慕尼黑技术大学)
专题命中 Agent评测 :agentic(title,abstract);agent(abstract,journal_ref);分类 cs.AI
Comments To appear at REALM@ACL2025
Journal ref Proceedings of the 1st Workshop for Research on Agent Language Models (REALM 2025), pages 298-308, Vienna, Austria. Association for Computational Linguistics
专题命中 Agent评测 :agent(title,abstract);agentic(abstract);分类 cs.AI
Comments The project can be found at https://github.com/metauto-ai/agent-as-a-judge. The dataset is released at https://huggingface.co/DEVAI-benchmark
专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.CL、cs.LG;autonomous agent(journal_ref)
Journal ref International Conference on Autonomous Agents and Multiagent Systems 2024
长视频中的条件多事件时间定位
机构 * University of Central Florida(中佛罗里达大学) ; Qualcomm AI Research(高通人工智能研究院)
专题命中 Agent评测 :agent(summary_cn,abstract);agentic(abstract)
AI总结 提出CoMET-Bench基准和CoMET-Agent框架,解决长视频中基于组合时空条件定位所有事件的任务,F1@0.5提升6.1%。
基于PRISM的自主代理行为建模——一个案例研究
机构 * University of Glasgow(格拉斯哥大学) ; University of Sheffield(谢菲尔德大学)
专题命中 Agent评测 :agent(title);autonomous agent(title)
AI总结 本文提出了一种抽象的自主性定义,用于建模自主场景,并通过小规模仿真模型推断定量数据,以验证无人飞行器在自主场景中的行为。
通过技能级评估与诊断解构多模态语言模型的具身能力
机构 * Northeastern University, Boston, MA, USA ; The Chinese University of Hong Kong, Hong Kong, China ; Peking University, Beijing, China ; Westlake University, Hangzhou, China ; Harvard University, Cambridge, MA, USA ; Purdue University, West Lafayette, IN, USA ; University of Oxford, Oxford, United Kingdom
专题命中 Agent评测 :agent(summary_cn,abstract);planning(abstract)
AI总结 本文提出BEAR基准,通过分解具身任务为14个原子技能进行细粒度评估,发现感知能力是推理失败的主要瓶颈,并提出BEAR-Agent多模态对话代理,显著提升具身技能性能。
Comments Accepted to ICML 2026
你的驾驶世界模型是全能选手吗?
专题命中 Agent评测 :agent(summary_cn,abstract);planning(abstract)
AI总结 本文提出WorldLens基准测试,评估驾驶世界模型在视觉和行为真实性方面的综合表现,揭示现有模型在不同维度上的不足,并引入WorldLens-26K和WorldLens-Agent提升评估的可解释性。
Comments CVPR 2026 VideoWorldModel Workshop; Project Page at https://worldbench.github.io/worldlens GitHub at https://github.com/worldbench/WorldLens
E2EDev:端到端软件开发任务中大语言模型的基准测试
机构 * College of Computer Science, Sichuan University(四川大学计算机学院) ; Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所) ; CHAT NLP Group, Singapore Management University(新加坡国立管理大学CHAT NLP组) ; Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education, China(教育部长机器学习与产业智能工程研究中心)
专题命中 Agent评测 :agent(abstract,abstract_cn);multi-agent(abstract,abstract_cn);分类 cs.AI、cs.CL、cs.SE
AI总结 本文提出E2EDev基准,基于行为驱动开发原则,通过模拟真实用户交互评估端到端软件开发框架的能力,揭示现有方法在解决复杂任务中的不足。
Comments Accepted to ACL 2026 main
一般均衡模型中无限寿命 agent 与重叠世代模型的关系及其一些应用
专题命中 Agent评测 :agent(title_cn,summary_cn)
AI总结 本文研究了无限寿命 agent 一般均衡模型与重叠世代模型的关系,证明了两者在均衡条件下的相互转化,并探讨了其在经济应用中的意义。
利用低功耗自主代理控制的纳米卫星进行太空碎片清除
机构 * Institut für Datentechnik, TU Braunschweig(数据技术研究所, Braunschweig 技术大学) ; Institut für Raumfahrtsysteme, TU Braunschweig(航天系统研究所, Braunschweig 技术大学)
专题命中 Agent评测 :autonomous agent(title,abstract);agent(abstract,comments);multi-agent(comments)
AI总结 本文提出利用低功耗自主代理控制的纳米卫星群,用于安全清除太空碎片,通过实验验证其可行性与能耗效率。
Comments This is an open-access, author-archived version of a manuscript published in European Conference on Multi-Agent Systems 2024
专题命中 Agent评测 :agent(title);multi-agent(title)
专题命中 Agent评测 :agent(title);multi-agent(title)
专题命中 Agent评测 :agent(title,abstract);multi-agent(abstract,comments)
Comments Presented at the International Workshop on Multi-Agent Systems and Agent-Based Simulation (MABS@AAMAS) 2021, 12 pages, 8 figures
专题命中 Agent评测 :agent(title);multi-agent(title)
专题命中 Agent评测 :agent(title);multi-agent(title)
Comments IJCAI 2021 Reinforcement Learning for Intelligent Transportation Systems (RL4ITS) Workshop
专题命中 Agent评测 :agent(title,abstract);multi-agent(abstract,journal_ref)
Journal ref Presented at the 2019 ICML Workshop on AI in Finance: Applications and Infrastructure for Multi-Agent Learning. Long Beach, CA
专题命中 Agent评测 :agent(title,abstract);分类 cs.AI、cs.SE;autonomous agent(comments);multi-agent(comments)
Comments This is a preprint of an article with the same title, accepted in 9th International Workshop on Engineering Multi-Agent Systems (EMAS 2021) which was held as a part of 20th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2021)
TimeSage-EV:面向动态环境下智能体时间序列分析的实时基准测试集
机构 * Eindhoven University of Technology(埃因霍温理工大学) ; University of Oxford(牛津大学) ; VulpiVox Intelligence(VulpiVox智能公司) ; Squirrel Ai Learning(Squirrel AI学习公司) ; Griffith University(格里菲斯大学)
专题命中 Agent评测 :agentic(title,abstract);agent(abstract);分类 cs.AI
AI总结 针对现有时间序列基准未评估时间有效性等问题,推出动态环境智能体时间序列分析基准TimeSage-EV,含多领域真实场景,实验揭示模型性能差距及失败模式,提供研究资源。