RecSim NG: Toward Principled Uncertainty Modeling for Recommender Ecosystems
专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG
AI 大模型
智能体、工具调用、规划、工作流、多智能体和自主任务执行。
专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG
Comments Presented at ICLR 2021
专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.LG
Comments AAAI 2020
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG
Comments 16 pages, 5 figures, Code available at: https://github.com/salmanmaq/segmentationNetworks, Dataset available at: https://www.kaggle.com/salmanmaq/m2caiseg
专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG
Comments AAAI 2021
专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG
Comments 3 figures
Journal ref Sci. Adv. 6, eabb6987 (2020)
专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG
Comments PhD thesis, Aerospace Engineering, Texas A&M (2020). For more information, see https://vggoecks.com/
专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG
Comments Submitted to IEEE transaction on sustainable computing. arXiv admin note: text overlap with arXiv:1906.05735
专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG
Comments Published as a conference paper at ICLR 2020
Journal ref International Conference on Learning Representations 2020
专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG
Comments 27 pages, 16 figures
专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG
Comments AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE) 2019
专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.LG
Comments 8 pages, for: Conference on Games (CoG), London, 2019. Index Terms: game learning, general game playing, AI, temporal difference learning, board games, n-tuple systems
专题命中 Agent评测 :agent(abstract);planning(abstract);分类 cs.AI、cs.LG
Comments Preprint submitted to ASME JMD
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG
专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.LG
Comments To be published in Machine Learning Journal
Journal ref Machine Learning 2017
专题命中 Agent评测 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG
Comments 23rd International Joint Conference on Artificial Intelligence (IJCAI 2013), Extended version with proofs, 10 pages
主动检索生成式人工智能何时应进行检索?效用、校准和成本的预算感知评估
机构 * Carnegie Mellon University(卡内基梅隆大学) ; University of Glasgow(格拉斯哥大学) ; Georgia Institute of Technology(佐治亚理工学院)
专题命中 Agent评测 :agentic(abstract,abstract_cn);分类 cs.LG
AI总结 研究主动检索生成式人工智能何时检索,通过将其重述为效用估计进行预算感知评估,并分离出相关三个问题,利用多种方法实现,在多数据集和模型中验证,强调评估应报告多方面指标。
Comments Accepted at the ACM SIGKDD KDD 2026 Workshop on Evaluation and Trustworthiness of Agentic AI; 7 pages, 1 figure, and 4 tables
柏拉图生物:通过时间再发现和结构基准进行验证优先的生物新奇性筛选
机构 * Praxa Labs(普拉克斯实验室)
专题命中 Agent评测 :agent(abstract,comments);workflow(abstract);分类 cs.AI
AI总结 研究开发柏拉图生物,扩展开放架构并结合多种功能,修复评估缺陷。通过Python套件等验证其有效性,评估两个用例,如历史再发现任务及蛋白质结构比较,提供可重复软件契约和筛选基准,为生物研究提供支持。
Comments 16 pages, 6 figures, 3 tables. Companion code and data: https://github.com/Eldergenix/Plato-Scientific-Research-Autonomous-Agent. This fork-specific validation study cites, but does not duplicate, arXiv:2510.26887
弗拉索夫方程平均场推导的形式化:作为策略游戏的人工智能辅助精益形式化
机构 * Stanford University(斯坦福大学) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 Agent评测 :agent(abstract,comments);AI agent(abstract);分类 cs.AI
AI总结 该研究以数学家指导AI在Lean 4中形式化研究成果为案例,将其构建为形式化游戏。通过此方式对非线性弗拉索夫方程适定性完整形式化,展示了开发过程及成果,还介绍了最优传输机制的分离情况及开发时间等,为形式化研究提供了新方法。
Comments 26 pages, 4 figures. Lean 4 development, blueprint site, and agent logs: https://github.com/Hydrodynamical/Vlasov_Meanfield_Formalization
当智能体与自身意见相左:测量基于LLM的智能体的行为一致性
机构 * Aman Mehta
专题命中 Agent评测 :agentic(abstract,comments);agent(abstract);分类 cs.AI
AI总结 研究发现基于LLM的智能体在相同任务上运行结果不一致,且这种不一致与任务成功率密切相关,通过监控行为一致性可提升智能体可靠性。
Comments Accepted at the ICML 2026 Workshop on Statistical Frameworks for Uncertainty in Agentic Systems. 12 pages, 9 figures
开源LLM代理能否取代静态应用安全测试工具?一项实证评估
机构 * College of Engineering and Science, Florida Institute of Technology(工程学院与科学学院,佛罗里达理工学院)
专题命中 Agent评测 :agentic(abstract,comments);agent(abstract);分类 cs.AI
AI总结 评估基于开源LLM的代理在静态应用安全测试中的性能,与SAST工具Bandit对比,发现当前不适合实际应用。
Comments Keywords: Agentic AI, Cybersecurity, Large Language Models, Static Application Security Testing, Model performance evaluation
机构 * The University of Electro-Communications(电子通信大学)
专题命中 Agent评测 :agent(abstract,comments);AI agent(abstract);分类 cs.AI
Comments 23 pages, 19 figures. International Conference on Human-Agent Interaction (HAI 2025), November 10-13, 2025, Yokohama, Japan