TalkTive: A Conversational Agent Using Backchannels to Engage Older Adults in Neurocognitive Disorders Screening
专题命中 Agent评测 :agent(title);分类 cs.AI
Comments Accepted by CHI2022
AI 大模型
智能体、工具调用、规划、工作流、多智能体和自主任务执行。
专题命中 Agent评测 :agent(title);分类 cs.AI
Comments Accepted by CHI2022
专题命中 Agent评测 :agent(title);分类 cs.AI
Comments 9 pages, 3 figures
专题命中 Agent评测 :agent(title);分类 cs.AI
专题命中 Agent评测 :agent(abstract,comments);multi-agent(abstract,comments);分类 cs.LG
Comments MPhil Thesis, 76 pages, Reinforcement Learning, Multi-agent, multi-task
专题命中 Agent评测 :agent(title);分类 cs.AI
专题命中 Agent评测 :agent(title);分类 cs.AI
Comments 8 pages. arXiv admin note: text overlap with arXiv:1308.5032, arXiv:1005.1516, arXiv:1309.7407, arXiv:0911.2390, arXiv:0811.2551, arXiv:1310.0522
Journal ref (2011). In A. Goel, F. Harrell, B. Magerko, & J. Prophet (Eds.), Proceedings of the 8th ACM Conference on Cognition & Creativity (pp. 299-306). New York: Association for Computing Machinery (ACM)
专题命中 Agent评测 :agent(title);分类 cs.AI
专题命中 Agent评测 :agent(title);分类 cs.AI
Comments AAAI-19 Workshop on Games and Simulations for Artificial Intelligence
专题命中 Agent评测 :agent(title);分类 cs.AI
Comments arXiv admin note: text overlap with arXiv:1708.04782 by other authors
专题命中 Agent评测 :agent(title);分类 cs.AI
专题命中 Agent评测 :agent(title);分类 cs.AI
Comments IJCSI International Journal of Computer Science Issues, Vol. 9, Issue 5, No 3, September 2012
专题命中 Agent评测 :agent(title);分类 cs.CL
Comments 10pages, 11 figures; ISSN (Online): 1694-0814
Journal ref IJCSI International Journal of Computer Science Issues, Vol. 9, Issue 1, No 1, January 2012
专题命中 Agent评测 :agent(title);分类 cs.AI
Comments 5 pages
专题命中 Agent评测 :agent(title);分类 cs.AI
Journal ref Dans Proceedings of the World Congress on Engineering - Proceedings of the World Congress on Engineering, London : Royaume-Uni (2007)
专题命中 Agent评测 :agent(title);分类 cs.AI
Comments 12 pages, in MICAI 2000: Advances in Artificial Intelligence. Lecture Notes in Artificial Intelligence 1793, pp. 634-648. Springer-Verlag
Journal ref # MICAI 2000: Advances in Artificial Intelligence. Lecture Notes in Artificial Intelligence 1793, pp. 634-648. Springer-Verlag
专题命中 Agent评测 :agent(abstract,comments);autonomous agent(abstract);分类 cs.AI
Comments 13 pages, 9 figures, in 1999 international conference of Intelligent Agent Technology. Nominated for the best paper award
Journal ref in Jiming Liu and Ning Zhong (Eds.), Intelligent Agent Technology: Systems, Methodologies, and Tools, page 110-120, The World Scientific Publishing Co. Pte, Ltd., Nov. 1999
专题命中 Agent评测 :agent(abstract,comments);multi-agent(abstract,comments);分类 cs.LG
Comments 12 pages, 6 figures, 1 table, accepted in Second Symposium on Adaptive Agents and Multi-Agent Systems (AAMAS-II), 2002
专题命中 Agent评测 :agent(title);autonomous agent(comments)
Comments An extended abstract of this paper has been accepted for the Eighteenth International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 2019
专题命中 Agent评测 :agent(title,journal_ref)
Comments Won Best Paper Award. This work is sponsored by the Assistant Secretary of Defense for Research & Engineering under Air Force Contract #FA8721-05-C-0002. Opinions, interpretations, conclusions and recommendations are those of the author and are not necessarily endorsed by the United States Government
Journal ref Bernstein, G. and O'Brien, K. 'Stochastic Agent-Based Simulations of Social Networks.' Proceedings of 46th Annual Simulation Symposium, San Diego, 7-10 April 2013. 33-40. Print
DeltaML-Bench:基于真实研究仓库评估机器学习智能体
机构 * Algorithmic Research Group(算法研究组)
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG
AI总结 该研究推出DeltaML-Bench基准,评估发现基于搜索的ARG框架可提升GPT-5在机器学习实验任务中的成功率,且能避免规范博弈,为自主ML智能体部署提供了关键考量
Comments 18 pages, 3 figures, 12 tables. Code and benchmark: https://github.com/AlgorithmicResearchGroup/deltaml-bench-public
超越回忆:行为规范作为AI个性化的解释层
专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.CL
AI总结 提出行为规范作为解释层,通过压缩用户数据为解释性模式,显著提升AI代理对用户意图的表示准确性,减少模型规避,并在解释型问题上优于原始语料和商业记忆系统。
Comments 142 pages, 4 figures. Code, data, judge prompts, and reproduction instructions: github.com/agulaya24/beyond-recall 8/19 Replacement: Pages updated to 142 from 134. Added Table of Contents. Updated table/graph formatting. Updated Section 8: Data, Code, and Reproducibility to include updated links, corrected author names
工程推理与指令(ERI)基准:一个大规模的基于分类体系的数据集用于基础模型和代理
专题命中 Agent评测 :tool-use(abstract);agentic(abstract);分类 cs.AI、cs.SE
AI总结 ERI基准通过大规模分类数据集评估工程能力,揭示LLM在不同难度问题上的性能差异,并采用收敛验证协议降低幻觉风险。
毒苹果效应:通过AI代理技术扩展战略操纵中介市场
机构 * Faculty of Data and Decision Sciences, Technion – Israel Institute of Technology, Haifa, Israel(以色列理工学院数据与决策科学学院,海法,以色列)
专题命中 Agent评测 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.CL
AI总结 研究探讨了AI代理技术扩展对博弈论场景中经济影响,发现技术扩展可显著改变均衡收益与监管结果,提出'毒苹果'效应通过释放未被使用的新技术操纵监管设计,损害对手与监管公平性。
SimulCost: 一个用于自动化物理模拟的代价感知基准与工具包
机构 * University of California San Diego(加州大学圣地亚哥分校) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Peking University(北京大学) ; University of California, Los Angeles(加州大学洛杉矶分校) ; California Institute of Technology(加州理工学院) ; ETH Zurich(苏黎世联邦理工学院)
专题命中 Agent评测 :tool-use(abstract);agentic(abstract);分类 cs.AI、cs.LG
AI总结 针对现有LLM评估忽略工具使用代价的问题,提出SimulCost基准,通过单轮和多轮参数调优任务比较LLM与传统扫描方法在准确性和计算代价上的表现,发现LLM在高精度任务中初始猜测不可靠且多轮模式效率更低。
Comments post conference revision version at ICML; update: removed CGYRO due to bug in cases search. Will add back soon; Make the title consistent w/ pdf
绝非数字:将答案作为事实使用的AI系统的结构弃权
机构 * Apple Inc.(苹果公司)
专题命中 Agent评测 :agent(abstract);agentic(abstract);分类 cs.AI、cs.CL
AI总结 针对LLM文本转SQL系统返回错误答案且无法区分的问题,提出带生成式外壳的可信内核架构,即结构弃权,在生产案例中验证其可靠性优于两种生成式替代方案。
Comments 26 pages, 5 figures, 5 tables. Technical report. Describes architecture and design principles only; contains no code, schemas, datasets, or performance metrics
DiG-bench:游戏中的发现
专题命中 Agent评测 :AI agent(abstract);agentic(abstract);分类 cs.AI、cs.LG
AI总结 本研究发布DiG-bench基准,含70个需智能体通过交互实验发现规则的游戏,设7个难度级,部分公开部分私有,用于评估智能体在目标未知环境中发现新知识的能力。
迈向具身智能体的管控
专题命中 Agent评测 :agent(abstract);agentic(abstract);分类 cs.AI、cs.LG
AI总结 本文提出名为Thea的具身智能体管控系统,通过场景图和退出码评估弥合物理世界与智能体的差距,实现长程任务完成。
Comments Project page: https://eit-hai.github.io/thea
LakeQuest:用于跨数据湖的有基础问答的三领域基准测试
机构 * University of Waterloo(滑铁卢大学)
专题命中 Agent评测 :tool-use(abstract);agentic(abstract);分类 cs.AI、cs.CL
AI总结 介绍LakeQuest基准测试,用于评估跨数据湖的有基础问答。它跨越三个领域,含9846个QA对及证据指针,能暴露系统故障模式。通过基线评估发现高质量检索不能保证正确推理,凸显未来智能QA系统需强大发现和跨文件组合机制。
Comments 24 pages, 4 figures, 18 tables. Accepted at the Conference on Language Modeling (COLM) 2026
利用自主LLM研究循环优化专家设计的晶体图网络用于带隙预测
机构 * Department of Materials Science and NanoEngineering(材料科学与纳米工程系)
专题命中 Agent评测 :agent(abstract);autonomous agent(abstract);分类 cs.AI、cs.LG
AI总结 提出一个自主LLM研究循环,在MatBench带隙基准上构建了无需外部预训练的最准确模型,超越了所有17个专家设计模型,通过实现元素对特征和空间群嵌入等已知方法。
DuplexWorld:语音智能体能帮你度过一天吗?
机构 * Centific Global Solutions Inc.(森蒂菲克全球解决方案公司) ; University of Maryland(马里兰大学)
专题命中 Agent评测 :agent(abstract);agentic(abstract);分类 cs.AI、cs.CL
AI总结 DuplexWorld针对现有语音智能体评估基准的不足,构建含六大领域的156个场景开展评估,发现现有最优语音智能体在多维度仍有较大改进空间,并分析了相关性能与失败模式。