LOA: Logical Optimal Actions for Text-based Interaction Games
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL、cs.LG
Comments ACL-IJCNLP 2021 (demo paper)
AI 大模型
智能体、工具调用、规划、工作流、多智能体和自主任务执行。
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL、cs.LG
Comments ACL-IJCNLP 2021 (demo paper)
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL、cs.LG
Comments Contains main article and supplementaries
Journal ref Neurips 2021
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.LG、cs.SE
Comments 15 pages, 7 figures
专题命中 软件智能体 :agent(abstract);AI agent(abstract)
Comments Code is available at: https://github.com/amazon-research/progressive-coordinate-transforms
专题命中 软件智能体 :agent(abstract);autonomous agent(abstract)
Comments SAIAD CVPR21 Workshop
专题命中 软件智能体 :agent(abstract);multi-agent(abstract)
Comments 15 pages, 9 figures, journal
专题命中 软件智能体 :agent(abstract);AI agent(abstract)
Comments arXiv admin note: text overlap with arXiv:2003.10286
专题命中 软件智能体 :agent(abstract);autonomous agent(abstract)
Journal ref IEEE Transactions on Robotics, 2020
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL、cs.LG
Comments Accepted at WACV 2019. Also at NeurIPS 2017 workshop on Visually-Grounded Interaction and Language (ViGIL)
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL、cs.LG
Comments Published in Advances in Neural Information Processing Systems (NIPS) 30, December 2017
Journal ref Rothe, A., Lake, B. M., and Gureckis, T. M. (2017). Question asking as program generation. Advances in Neural Information Processing Systems 30
专题命中 软件智能体 :agent(abstract);autonomous agent(abstract)
专题命中 软件智能体 :agent(abstract);multi-agent(abstract)
Journal ref The World of Computer Science and Information Technology Journal (WSCIT). 2014, Volume 4, Issue 2. pp. 18.25
专题命中 软件智能体 :agent(abstract);multi-agent(abstract)
Comments 6 pages
Journal ref IJCSI International Journal of Computer Science Issues, Vol. 10, Issue 2, No 3, March 2013
专题命中 软件智能体 :agent(abstract);planning(abstract)
Comments 12 Pages, 3 Tables, 3 Figures
Journal ref Proc. European Spreadsheet Risks Int. Grp. (EuSpRIG) 2003 147-159 ISBN 1 86166 199 1
跨模型大语言模型代码审查:应该用Claude审查Codex还是相反?
机构 * University of California, Davis(加州大学戴维斯分校) ; Johns Hopkins University(约翰霍普金斯大学) ; California State University, Long Beach(长滩加州州立大学)
专题命中 软件智能体 :workflow(abstract);分类 cs.AI、cs.SE;agentic(comments)
AI总结 研究开发者同时使用Claude和Codex进行代码审查的成本、时间及配对顺序问题,通过对116个任务的六种条件实验发现,Claude审查Codex草稿效果好,反向则不佳,有用的配对是不对称的,应Claude审查Codex。
Comments This paper had been accepted by Agentic SE @ KDD'26
用强化学习建模客户轨迹以获得实际零售洞察
机构 * McGill University(麦吉尔大学) ; Mila - Quebec AI Institute(魁北克人工智能研究所)
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)
AI总结 本文提出了一种基于智能体的建模框架,将客户轨迹预测转化为最大熵强化学习问题,以更准确地反映具有有限理性的客户行为,从而提供更精确的冲动购买率和货架交通密度估计。
Comments Proceeding of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026)
Banana100: 通过100次迭代图像复制打破NR-IQA度量标准
机构 * University of California, Santa Barbara(加州大学圣塔芭芭拉分校)
专题命中 软件智能体 :agentic(abstract,comments);分类 cs.AI、cs.LG
AI总结 Banana100通过100次迭代编辑生成28000张退化图像,揭示多轮编辑中图像质量退化问题,发现现有NR-IQA度量标准无法检测严重退化图像,威胁未来模型训练稳定性。
Comments Accepted to CVPR 2026 Workshop on Agentic AI for Visual Media
基于反射的可信代码代理控制
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.SE;agentic(comments)
AI总结 本文提出反射驱动控制方法,通过内部反思循环提升代码生成的安全性和合规性,实现自主、安全且可审计的AI代码代理。
Comments Accepted to AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)
机构 * LinkedIn(领英)
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL;agentic(comments)
Comments 11 pages, 8 figures, Workshop on Agentic AI for Enterprise at KDD '25
专题命中 软件智能体 :agent(abstract,comments);分类 cs.AI、cs.CL
Comments Code, data, and over 24k agent trajectories are released at https://github.com/Leezekun/SOPBench
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.LG;autonomous agent(comments)
Comments To appear in 18th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2019) as a full paper. arXiv admin note: substantial text overlap with arXiv:1806.08055
超越满意:从安慰性到可操作性解释以提升可理解性
机构 * The University of Tulsa(图兰大学)
专题命中 软件智能体 :agent(abstract,journal_ref);分类 cs.AI;multi-agent(journal_ref)
AI总结 本文探讨了可解释性在提升系统可理解性中的作用,通过实验发现可操作性解释在任务表现上优于安慰性解释,但用户满意度评分相同,强调需结合客观指标与主观评估来衡量解释质量。
Comments 21 pages, 7 figures, 6 tables. EXTRAAMAS 2025 submission. Preprint version
Journal ref In: Calvaresi, D., et al. Explainable, Trustworthy, and Responsible AI and Multi-Agent Systems. EXTRAAMAS 2025. Lecture Notes in Computer Science. Springer, Cham
机构 * Center for Modeling Social Systems(社会科学建模中心) ; NORCE Norwegian Research Center AS(挪威NORCE研究机构) ; Kristiansand, Norway(挪威克里斯蒂安桑)
专题命中 软件智能体 :agent(abstract,comments);分类 cs.AI;multi-agent(comments)
Comments 13 pages, 3 figures, 23rd International Conference on Practical applications of Agents and Multi-Agent Systems (PAAMS 2025)
个性化技能对编码智能体有帮助吗?开发者交互历史的实证研究
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.SE
AI总结 本研究通过对13位开发者的206个真实会话实验,发现从交互历史提炼的开发者个性化技能改进有限,而汇集自所有开发者的通用技能增益最大且最一致,为编码智能体的个性化策略提供了实证依据。
Comments 15 pages, 10 figures
进化代理群体
机构 * National University of Singapore(新加坡国立大学)
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.LG
AI总结 本文提出EvE框架,通过进化编码代理群体实现算法发现,解决了传统方法的局限,展示了自修正群体在复杂代码库中的优势。
构建可靠的编码智能体:评估与运行模型周边系统
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.SE
AI总结 本研究针对AI编码智能体的可靠性问题,整合多源证据构建系统级评估与运行框架,区分模型与基础设施效应,提出可靠性记录目录及相关方法以提升智能体系统可靠性。
Comments Technical review and engineering monograph, 314 pages, 30 figures. Includes an evidence audit, a companion research artifact with 206 reliability records, and runnable protocols for evaluating and operating AI coding agents. August 2026. Source, companion, and reusable protocols: https://github.com/sjarmak/engineering-reliable-coding-agents
微调Qwen3-27B实现C到Rust代码翻译:预训练、调试感知监督微调与任务特定监督微调的三阶段课程
专题命中 软件智能体 :agentic(abstract);分类 cs.AI、cs.SE
AI总结 本研究针对C到Rust代码翻译任务,对Qwen3-27B采用三阶段微调课程,结合SACTOR框架评估后,其性能优于基线模型及其他LLM。
语言服务器能为编码智能体节省 token 吗?一种测量方法与初步研究
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.CL
AI总结 本文通过五组消融实验等方法,对比 LSP 与 grep 检索的 token 效率,发现 LSP 通常不节省 token,仅对最弱模型有 token 节省效果,需根据任务、模型等选择检索工具。
Comments 13 pages, 6 figures. Code and data: https://github.com/Poytr1/lsp-vs-grep-token-study
从代码审查到代码批判:大规模AI生成代码差异的意图、漂移与焦点
专题命中 软件智能体 :agent(abstract);分类 cs.AI、cs.SE
AI总结 针对AI生成代码超出传统审查能力且现有工具忽视核心问题的缺陷,提出ARCTIC系统,经实验验证其在意图预测、漂移检测等方面表现优异,可降低代码不一致性并获高认可。
一种面向可重复、低代码定位的配置优先框架
机构 * Jožef Stefan Institute(乔塞夫·斯蒂芬研究所)
专题命中 软件智能体 :workflow(abstract);分类 cs.LG、cs.SE
AI总结 本文提出了一种低代码、配置优先的框架,通过版本化配置和自动化流程提升定位任务的可重复性和效率。
Comments 16 pages, 7 figures