arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 3200 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 软件智能体 3200 篇

2503.11237 2025-08-01 cs.AI cs.CL cs.SE 75%

Collaboration is all you need: LLM Assisted Safe Code Translation

Rabimba Karanjai, Sam Blackshear, Lei Xu, Weidong Shi

机构 * University Of Houston(休斯顿大学) Mysten Labs(Mysten实验室) Kent State University(肯特州立大学)

专题命中 软件智能体 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.CL、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02003 2025-06-30 cs.SE cs.AI cs.LG cs.PL 75%

L2MAC: Large Language Model Automatic Computer for Extensive Code Generation

Samuel Holt, Max Ruiz Luyten, Mihaela van der Schaar

机构 * University of Cambridge(剑桥大学)

专题命中 软件智能体 :agent(abstract);multi-agent(abstract);分类 cs.AI、cs.LG、cs.SE

Comments Published in The Twelfth International Conference on Learning Representations (ICLR), 2024. Copyright 2023 by the author(s)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11107 2025-04-17 cs.SE cs.AI cs.CL cs.DC cs.NI 75%

ChaosEater: Fully Automating Chaos Engineering with Large Language Models

Daisuke Kikuta, Hiroki Ikeuchi, Kengo Tajiri

专题命中 软件智能体 :workflow(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.SE

Comments 114 pages (7 main), 11 figures. Project page: https://ntt-dkiku.github.io/chaos-eater

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11915 2024-10-11 cs.SE cs.AI cs.LG 75%

miniCodeProps: a Minimal Benchmark for Proving Code Properties

Evan Lohn, Sean Welleck

专题命中 软件智能体 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.LG、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08335 2024-08-19 cs.SE cs.AI cs.CL 75%

Plan with Code: Comparing approaches for robust NL to DSL generation

Nastaran Bassamzadeh, Chhaya Methani

专题命中 软件智能体 :planning(abstract);workflow(abstract);分类 cs.AI、cs.CL、cs.SE

Comments 9 pages, 1 figure, 5 tables. arXiv admin note: substantial text overlap with arXiv:2407.02742

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.07511 2021-08-05 cs.AI cs.CL cs.CV cs.IR cs.LG 75%

Ensemble of MRR and NDCG models for Visual Dialog

Idan Schwartz

专题命中 软件智能体 :agent(abstract);AI agent(abstract);分类 cs.AI、cs.CL、cs.LG

Comments Accepted to NAACL2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30353 2026-05-29 cs.AI astro-ph.CO cs.HC cs.SE 74%

Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software

物理学就是一切?物理学家监督人工智能开发科学软件的案例研究

Nhat-Minh Nguyen

机构 * Kavli IPMU (WPI), UTIAS, The University of Tokyo(Kavli研究所(WPI)、UTIAS、东京大学) Center for Data-Driven Discovery(数据驱动发现中心) Institute For Interdisciplinary Research in Science(科学跨学科研究中心)

专题命中 软件智能体 :AI agent(abstract,comments);agent(abstract);分类 cs.AI、cs.SE

AI总结 通过一个物理学家监督AI编码代理开发可微扰动理论模块的案例,研究AI代理在科学软件开发中的可靠性,发现监督设计比模型能力更能决定输出可信度。

Comments 10 pages, 2 figures, 2 tables, 1 physicist and a few AI agents. Accepted by ICML 2026 AI for Science Workshop. Code and development log are available at this repo: https://github.com/MinhMPA/clax-pt

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09388 2026-04-28 cs.SE cs.AI 74%

The AI Codebase Maturity Model: From Assisted Coding to Fully Autonomous Systems

AI代码库成熟度模型:从辅助编码到完全自主系统

Andy Anderson

机构 * IBM Research — Hybrid Cloud and AI Platform(IBM混合云和人工智能平台研究院)

专题命中 软件智能体 :agent(abstract,comments);multi-agent(abstract);分类 cs.AI、cs.SE

AI总结 本文提出AI代码库成熟度模型(ACMM),通过100天经验报告和Hive系统验证,展示代码库从基础AI辅助到完全自主的六级进化框架,强调测试覆盖率和反馈机制是关键因素。

Comments 30 pages, 7 tables. v2: Extended to 6 levels. Added Level 6 (Fully Autonomous), Hive reference implementation, Beads for agent memory continuity, throughput acceleration data. Metrics updated to 100 days. Source: https://github.com/kubestellar/console and https://github.com/kubestellar/hive

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22758 2026-08-11 cs.AI 版本更新 74%

AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts

AutoRefine: 从轨迹到可重用的专业知识为连续LLM代理优化

Libin Qiu, Zhirong Gao, Junfu Chen, Yuhang Ye, Liangyu Li, Weizhi Huang, Xiaobo Xue, Wenkai Qiu, Shuo Tang

机构 * alibaba(阿里巴巴集团)

专题命中 软件智能体 :agent(title);分类 cs.AI

AI总结 AutoRefine通过提取和维护双重形式的经验模式,提升连续LLM代理的知识积累与维护效率,实现98.4%的准确率提升。

Comments 7 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07696 2026-07-09 cs.DB cs.AI 新提交 74%

Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass

打破数据库锁定:用于数据库绕过的高性能存储读取器的智能再生

Victor Giannakouris, Immanuel Trummer

机构 * Cornell University(康奈尔大学)

专题命中 软件智能体 :agentic(title);分类 cs.AI

AI总结 研究分析工作负载面临的数据访问瓶颈问题,提出Jailbreak方法,利用大语言模型辅助代码合成绕过数据库引擎直接读取存储文件,在PostgreSQL和MySQL上评估,性能显著提升,证明该方法可打破数据库系统数据锁定。

Comments To be presented at AIDB 2026 (co-located with VLDB)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05677 2026-07-08 cs.SE cs.HC 新提交 74%

From Conversation to Contribution: Characterizing Coding Agent in Open-Source Software

从对话到贡献:开源软件中编码代理的特征分析

Zihan Fang, Yueke Zhang, Ningzhi Tang, Collin McMillan, Toby Jia-Jun Li, Yu Huang

专题命中 软件智能体 :agent(title);分类 cs.SE

AI总结 研究开源软件中开发者与人工智能编码代理聊天交互和后续开发协作的关系,通过收集大量会话及关联仓库历史等进行分析,发现小仓库AI使用多,采用后贡献者有变化,还揭示了开发者对AI代码的看法及相关担忧,为开源开发实践提供见解。

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17855 2026-06-09 cs.AI cs.RO 版本更新 74%

QuickLAP: Quick Language-Action Preference Learning for Semi-Autonomous Agents

QuickLAP: 为半自主代理快速语言-动作偏好学习

Jordan Abi Nader, David Lee, Nathaniel Dennler, Andreea Bobu

专题命中 软件智能体 :autonomous agent(title);分类 cs.AI

AI总结 本研究提出QuickLAP,一种融合物理和语言反馈的贝叶斯框架,用于实时推断奖励函数,通过大规模语言模型提取奖励特征注意力掩码和偏好偏移,从而在半自主驾驶模拟器中将奖励学习误差降低70%,并通过用户研究验证其可理解性和协作性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11012 2026-06-09 cs.SE 版本更新 74%

Beyond Accuracy: Behavioral Dynamics of Agentic Multi-Hunk Repair

超越准确率:智能体多块修复的行为动力学

Noor Nashid, Daniel Ding, Keheliya Gallaba, Ahmed E. Hassan, Ali Mesbah

专题命中 软件智能体 :agentic(title);分类 cs.SE

AI总结 研究LLM驱动的编码智能体在多块缺陷修复中的行为,通过404个多块bug的1616条修复轨迹分析定位、修复准确率、回归行为及操作动力学,发现修复准确率随bug分散度和复杂度下降,且失败修复消耗更多资源。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09268 2026-05-26 cs.SE 74%

Decoding the Configuration of AI Coding Agents: Insights from Claude Code Projects

解码AI编码代理的配置:来自Claude Code项目的洞察

Helio Victor F. Santos, Vitor Costa, Joao Eduardo Montandon, Marco Tulio Valente

专题命中 软件智能体 :agent(abstract,journal_ref);agentic(abstract,journal_ref);分类 cs.SE

AI总结 通过对328个Claude Code配置文件进行实证研究,揭示了代理编码系统的配置文件中指定的软件工程关注点及其共现模式,强调了定义架构约束的重要性。

Journal ref Accepted at 1st International Workshop on Agentic Engineering (AGENT 2026, colocated with ICSE), pages 63-67

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13874 2026-05-15 cs.NE cs.AI 74%

GEAR: Genetic AutoResearch for Agentic Code Evolution

GEAR:基于遗传算法的自主研究代理

Ahmadreza Jeddi, Minh Ngoc Le, Hakki C. Karaimer, Konstantinos G. Derpanis, Babak Taati

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) AI Center-Toronto, Samsung Electronics(多伦多AI中心,三星电子) York University(约克大学)

专题命中 软件智能体 :agentic(title);分类 cs.AI

AI总结 GEAR通过多路径搜索提升自主研究代理效果,采用群体搜索策略保持多个 promising 方向,通过突变和交叉探索新想法,优于传统方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06464 2026-05-12 cs.SE 74%

To What Extent Does Agent-generated Code Require Maintenance? An Empirical Study

代理生成的代码需要多少维护?一项实证研究

Shota Sawada, Tatsuya Shirai, Yutaro Kashiwa, Ken'ichi Yamaguchi, Hiroshi Iwata, Hajimu Iida

专题命中 软件智能体 :agent(title);分类 cs.SE

AI总结 研究通过分析AI生成与人工编写代码的维护情况,发现AI代码维护频率较低,主要修改类型为功能扩展,而人工代码则以修复bug为主,人类开发者主导了大部分维护工作。

Comments 6 pages, 2 figures, 2 tables. Accepted at the 30th International Conference on Evaluation and Assessment in Software Engineering (EASE 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18413 2026-04-22 cs.SE 74%

TypeScript Repository Indexing for Code Agent Retrieval

TypeScript仓库索引用于代码代理检索

Junsong Pu, Yichen Li, Zhuangbin Chen

专题命中 软件智能体 :agent(title);分类 cs.SE

AI总结 本文提出abcoder-ts-parser,通过TypeScript编译器API直接生成可靠的代码索引,提升大型TypeScript仓库中代码代理的上下文检索效率。

Comments This is a tool demonstration paper. 4 tables and 1 listing

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18341 2026-04-09 cs.SE 74%

Agentic Much? Adoption of Coding Agents on GitHub

代理那么多?GitHub上编码代理的采用情况

Romain Robbes, Théo Matricon, Thomas Degueule, Andre Hora, Stefano Zacchiroli

专题命中 软件智能体 :agentic(title);分类 cs.SE

AI总结 研究探讨了编码代理在GitHub上的采用情况,发现其采用率高达22.20%-28.66%,且在项目成熟度、组织类型和编程语言上均广泛适用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21083 2026-02-10 cs.AI 74%

OpenSec: Measuring Incident Response Agent Calibration Under Adversarial Evidence

OpenSec: 在对抗性证据下评估事件响应代理的校准

Jarrod Barnes

机构 * Jarrod Barnes

专题命中 软件智能体 :agent(title);分类 cs.AI

AI总结 OpenSec通过双控制强化学习环境评估事件响应代理在对抗性证据下的校准能力,发现前沿模型存在过度触发问题,校准差距体现在克制而非检测。

Comments 7 pages, 3 figures, 3 tables. Code: https://github.com/jbarnes850/opensec-env. Dataset: https://huggingface.co/datasets/Jarrodbarnes/opensec-seeds

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20230 2026-01-30 cs.CL cs.HC 74%

Unit-Based Agent for Semi-Cascaded Full-Duplex Dialogue Systems

基于单元的代理用于半级联全双工对话系统

Haoyuan Yu, Yuxuan Chen, Minjie Cai

机构 * Hunan University(湖南大学) Gongdao Technology(公道科技) Jilin University(吉林大学)

专题命中 软件智能体 :agent(title);分类 cs.CL

AI总结 本文提出基于单元的半级联全双工对话系统,利用多模态大语言模型和辅助模块实现高效对话处理,实验显示其在挑战赛中表现优异。

Comments ICASSP 2026 (Grant Challenge). https://github.com/yu-haoyuan/fd-badcat

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11421 2026-01-19 cs.RO cs.AI 74%

The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents

伟大的百日行进100:100个注重细节的任务用于评估具身体验人工智能代理

Ziyu Wang, Chenyuan Liu, Yushun Xiang, Runhao Zhang, Qingbo Hao, Hongliang Lu, Houyu Chen, Zhizhong Feng, Kaiyue Zheng, Dehao Ye, Xianchao Zeng, Xinyu Zhou, Boran Wen, Jiaxin Li, Mingyu Zhang, Kecheng Zheng, Qian Zhu, Ran Cheng, Yong-Lu Li

机构 * SJTU(上海交通大学) SII(上海人工智能研究院) Robbyant

专题命中 软件智能体 :AI agent(title);分类 cs.AI

AI总结 GM-100通过100个精心设计的任务全面评估具身体验人工智能代理的能力,推动机器人数据集任务设计的多样性和复杂性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01234 2025-12-03 cs.HC cs.AI 74%

Proactive Agentic Whiteboards: Enhancing Diagrammatic Learning

主动代理白板:增强图示学习

Suveen Ellawela, Sashenka Gamage, Dinithi Dissanayake

专题命中 软件智能体 :agentic(title);分类 cs.AI

AI总结 DrawDash通过实时语音驱动的视觉辅助,主动完成和优化教育图表,以减少教师认知负担并提升图示教学效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01720 2025-11-04 cs.CL 74%

Efficient Tool-Calling Multi-Expert NPC Agent for Commonsense Persona-Grounded Dialogue

Mahammad Nuriyev

机构 * Université Paris-Saclay(巴黎-萨克雷大学)

专题命中 软件智能体 :agent(title);分类 cs.CL

Comments 10 pages, 1 figure, 2 tables. Technical report for the Commonsense Persona-Grounded Dialogue Challenge (CPDC) 2025, part of the Wordplay 2025 Workshop @ EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21143 2025-10-29 cs.AI 74%

PanicToCalm: A Proactive Counseling Agent for Panic Attacks

Jihyun Lee, Yejin Min, San Kim, Yejin Jeon, SungJun Yang, Hyounghun Kim, Gary Geunbae Lee

机构 * Graduate School of Artificial Intelligence, POSTECH(人工智能研究生院,POSTECH) Department of Computer Science and Engineering, POSTECH(计算机科学与工程系,POSTECH)

专题命中 软件智能体 :agent(title);分类 cs.AI

Comments Accepted in EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13081 2025-10-16 cs.MA cs.AI 74%

Agentic Discovery: Closing the Loop with Cooperative Agents

J. Gregory Pauloski, Kyle Chard, Ian T. Foster

机构 * University of Chicago(芝加哥大学) Argonne National Laboratory(阿贡国家实验室)

专题命中 软件智能体 :agentic(title);分类 cs.AI

Comments Published in IEEE Computer Volume 58 Issue 10

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05739 2025-06-09 cs.CR cs.AI 74%

To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt

Zhilong Wang, Neha Nagaraja, Lan Zhang, Hayretdin Bahsi, Pawan Patil, Peng Liu

专题命中 软件智能体 :agent(title);分类 cs.AI

Comments To appear in the Industry Track of the 55th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21069 2025-05-28 cs.SE 74%

CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building

Zhengmin Yu, Yuan Zhang, Ming Wen, Yinan Nie, Wenhui Zhang, Min Yang

专题命中 软件智能体 :agent(title);分类 cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07835 2025-05-15 cs.NI cs.AI cs.MA 74%

Intelligent Product 3.0: Decentralised AI Agents and Web3 Intelligence Standards

Alex C. Y. Wong, Duncan McFarlane, C. Ellarby, M. Lee, M. Kuok

机构 * University of Cambridge(剑桥大学) RedBite Solutions Ltd(RedBite解决方案有限公司)

专题命中 软件智能体 :AI agent(title);分类 cs.AI

Comments 18 pages, 1 Figure, 3 Tables; Corrected typo in Section 3.4 heading

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23948 2025-04-01 cs.AI 74%

AI2Agent: An End-to-End Framework for Deploying AI Projects as Autonomous Agents

Jiaxiang Chen, Jingwei Shi, Lei Gan, Jiale Zhang, Qingyu Zhang, Dongqian Zhang, Xin Pang, Zhucong Li, Yinghui Xu

专题命中 软件智能体 :autonomous agent(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05292 2024-10-11 cs.HC cs.SE 74%

How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging

Qianou Ma, Hua Shen, Kenneth Koedinger, Tongshuang Wu

专题命中 软件智能体 :agent(title);分类 cs.SE

Comments 14 pages, 6 figures

Journal ref AIED 2024, LNAI 14829, pp. 1-16

详情

展开后加载摘要…

URL PDF HTML 收藏