arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 3188 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 软件智能体 3188 篇

2507.18755 2025-07-28 cs.SE cs.AI cs.PL 84%

Agentic Program Repair from Test Failures at Scale: A Neuro-symbolic approach with static analysis and test execution feedback

Chandra Maddila, Adam Tait, Claire Chang, Daniel Cheng, Nauman Ahmad, Vijayaraghavan Murali, Marshall Roch, Arnaud Avondet, Aaron Meltzer, Victor Montalvao, Michael Hopko, Chris Waterson, Parth Thakkar, Renuka Fernandez, Kristian Kristensen, Sivan Barzily, Sherry Chen, Rui Abreu, Nachiappan Nagappan, Payam Shodjai, Killian Murphy, James Everingham, Aparna Ramani, Peter C. Rigby

机构 * Meta

专题命中 软件智能体 :agentic(title,abstract);agent(abstract);分类 cs.AI、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16938 2025-07-23 cs.AI cs.CL cs.CV 84%

InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification

InternAgent Team, Bo Zhang, Shiyang Feng, Xiangchao Yan, Jiakang Yuan, Runmin Ma, Yusong Hu, Zhiyin Yu, Xiaohan He, Songtao Huang, Shaowei Hou, Zheng Nie, Zhilong Wang, Jinyao Liu, Tianshuo Peng, Peng Ye, Dongzhan Zhou, Shufei Zhang, Xiaosong Wang, Yilan Zhang, Meng Li, Zhongying Tu, Xiangyu Yue, Wangli Ouyang, Bowen Zhou, Lei Bai

机构 * InternAgent Team(InternAgent团队) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 软件智能体 :agent(title,abstract);multi-agent(abstract);分类 cs.AI、cs.CL

Comments Code: https://github.com/Alpha-Innovator/InternAgent, HomePage: https://alpha-innovator.github.io/InternAgent-project-page

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16650 2025-06-23 cs.SE cs.AI cs.MA 84%

SemAgent: A Semantics Aware Program Repair Agent

Anvith Pabba, Alex Mathai, Anindya Chakraborty, Baishakhi Ray

专题命中 软件智能体 :agent(title);workflow(abstract);agentic(abstract);分类 cs.AI、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20115 2025-05-27 cs.SE cs.AI 84%

AutoP2C: An LLM-Based Agent Framework for Code Repository Generation from Multimodal Content in Academic Papers

Zijie Lin, Yiqing Shen, Qilin Cai, He Sun, Jinrui Zhou, Mingjun Xiao

机构 * University of Science and Technology of China(中国科学技术大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 软件智能体 :agent(title,abstract);multi-agent(abstract);分类 cs.AI、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05006 2025-02-13 cs.SE cs.AI 84%

COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis

Weiqing Yang, Hanbin Wang, Zhenghao Liu, Xinze Li, Yukun Yan, Shuo Wang, Yu Gu, Minghe Yu, Zhiyuan Liu, Ge Yu

专题命中 软件智能体 :agent(title,abstract);multi-agent(abstract);分类 cs.AI、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12275 2024-09-24 cs.AI cs.CL 84%

WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment

Hao Tang, Darren Key, Kevin Ellis

专题命中 软件智能体 :agent(title,abstract);planning(abstract);分类 cs.AI、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12165 2024-08-01 cs.SE cs.AI cs.DC 84%

Building AI Agents for Autonomous Clouds: Challenges and Design Principles

Manish Shetty, Yinfang Chen, Gagan Somashekar, Minghua Ma, Yogesh Simmhan, Xuchao Zhang, Jonathan Mace, Dax Vandevoorde, Pedro Las-Casas, Shachee Mishra Gupta, Suman Nath, Chetan Bansal, Saravan Rajmohan

专题命中 软件智能体 :AI agent(title,abstract);agent(abstract);分类 cs.AI、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14690 2026-07-02 cs.SE 版本更新 84%

Harness Engineering for Agentic AI Coding Tools: An Exploratory Study

配置代理AI编码工具:探索性研究

Matthias Galster, Seyedmoein Mohsenimofidi, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes

专题命中 软件智能体 :agentic(title,abstract);agent(abstract);分类 cs.SE

AI总结 本文探讨了代理AI编码工具的配置机制,分析了多种工具的配置方法,发现Context Files主导配置,并提出AGENTS.md作为通用标准。

Comments 10 pages, 7 figures, 3 tables, based on "Configuring Agentic AI Coding Tools: An Exploratory Study" published in Proceedings of the 3rd ACM/IEEE International Conference on AI-powered Software (AIware 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16572 2026-06-02 cs.CR cs.AI 84%

Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem

上下文很重要:基于仓库感知的代理技能生态系统安全分析

Florian Holzbauer, David Schmidt, Gabriel Gegenhuber, Sebastian Schrittwieser, Johanna Ullrich

机构 * Interdisciplinary Transformation University (IT:U)(交叉学科转化大学) University of Vienna(维也纳大学) CDL AsTra Faculty of Computer Science(计算机科学学院CDL AsTra系)

专题命中 软件智能体 :agent(title,abstract);AI agent(abstract);分类 cs.AI;agentic(comments)

AI总结 通过仓库上下文感知分析,发现现有扫描器高估了恶意技能比例(从46.8%降至0.52%),并识别出废弃仓库劫持等新攻击向量。

Comments AgentSkills '26 Workshop: ACM Conference on AI and Agentic Systems (CAIS), Best Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20690 2026-05-27 cs.AI 84%

Declarative Data Services: Structured Agentic Discovery for Composing Data Systems

声明式数据服务:用于组合数据系统的结构化智能体发现

Shanshan Ye, Duo Lu

机构 * Northeastern University(东北大学) Brown University(布朗大学)

专题命中 软件智能体 :agentic(title,abstract);agent(abstract);分类 cs.AI;AI agent(comments)

AI总结 提出声明式数据服务(DDS)架构,通过分层类型契约将全局搜索分解为有界子搜索,解决无界智能体发现无法稳定收敛的问题,并在交易后端工作负载上验证其有效性。

Comments Accepted at AI Agents for Discovery in the Wild (AID-Wild), Workshop at ACM CAIS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11839 2026-04-15 cs.CR cs.AI 84%

Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents

超越静态沙箱:为自主AI代理的学得能力治理

Bronislav Sidik, Lior Rokach

机构 * Institute for Applied AI Research(应用人工智能研究所) Faculty of Computer and Information Science(计算机与信息科学学院) Ben-Gurion University of the Negev(贝内-约尔大学)

专题命中 软件智能体 :AI agent(title,abstract);agent(abstract,comments);分类 cs.AI

AI总结 本文提出Aethelgard框架,通过学得策略实现AI代理的最小必要能力集,解决能力过度配置问题。

Comments 17 pages (9 content pages), 2 figures, 7 tables. Submitted to NeurIPS 2026 Agent Safety Workshop. Code and dataset available at https://github.com/sidikbro/aethelgard

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01483 2026-04-03 cs.LO cs.AI cs.CR 84%

Type-Checked Compliance: Deterministic Guardrails for Agentic Financial Systems Using Lean 4 Theorem Proving

类型检查合规性:使用Lean 4定理证明的确定性守卫机制用于代理金融系统

Devakh Rashie, Veda Rashi

机构 * Independent Researcher(独立研究员) Thomas Jefferson High School for Science and Technology(托马斯·杰斐逊科技高中)

专题命中 软件智能体 :agentic(title,abstract);agent(abstract,comments);分类 cs.AI

AI总结 本文提出Lean-Agent协议,通过Lean 4定理证明技术,将机构政策自动形式化为代码,确保代理金融系统在微秒延迟内满足监管要求。

Comments 8 pages, 1 table. Code and live demo available at https://github.com/arkanemystic/lean-agent-protocol and https://axiom.devrashie.space

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15793 2024-11-13 cs.SE cs.AI cs.CL cs.HC cs.LG 84%

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, Ofir Press

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

Comments Code, data, and demo available at https://swe-agent.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11386 2026-08-13 cs.SE 新提交 83%

The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior

关键在于接口:评估工具架构如何影响编码智能体行为

Xiangzhe Xu, Hamidreza Saghir, Qianhui Wu, Marc-Alexandre Côté, Tong Wang, Kiran Lakkaraju, Kexin Pei, Xiangyu Zhang

专题命中 软件智能体 :agent(title,abstract);agentic(abstract);分类 cs.SE

AI总结 该研究通过编码智能体的对照实验,发现工具架构会显著影响智能体行为,不同工具架构在一致性、文件访问量、步骤数等方面表现出不同效果,其中Python CodeAct风格接口性能优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11241 2026-08-13 cs.AI cs.IR 新提交 83%

RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle

RecSys Factory:限制LLM智能体在工业推荐生命周期内的自主决策点

Dongyang Ao, Kaixiang Fang, Shijie Xu

机构 * Tencent(腾讯) FiT(腾讯FiT)

专题命中 软件智能体 :agent(title,abstract);workflow(abstract);分类 cs.AI

AI总结 RecSys Factory 是限制 LLM 智能体自主到决策点的平台,解决工业推荐系统运营的三角困境,在腾讯三业务线部署 78 天,CLI 调度成功率达 78.6%。

Comments 21 pages, 6 figures, 9 tables. Reports a 78-day deployment across three heterogeneous industrial recommender business lines (1,624 CLI-tool dispatches). Companion paper: AutoResearch (P3b), which instantiates the same substrate for autonomous research

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09290 2026-08-12 cs.SE 版本更新 83%

OpenCodeReview: Determinism over Non-Determinism for Cost-Effective Agent-Based Code Review

OpenCodeReview:面向成本效益型智能体代码评审的非确定性转确定性方案

Zhengfeng Li, Lei Zhang, Xianwei Wu, Zhengqi Zhuang, Yingjie Xu, Boge Wang, Shaofei Zhu, Chuan Wang, Peng Zhao, Xinyu Zheng, Guoping Rong

专题命中 软件智能体 :agent(title,abstract);tool use(abstract);分类 cs.SE

AI总结 OpenCodeReview针对LLM代码评审智能体的非确定性与上下文局部性缺陷,通过在三个流水线节点注入确定性,在AACR-Bench上实现SEM-F1提升且大幅降低令牌消耗,性能优于主流编码智能体。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27380 2026-08-11 cs.CV cs.AI 版本更新 83%

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

VideoCoCo:基于智能体双引擎系统的、以代码为思维链(CoT)的物理一致性视频生成方法

Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng

机构 * CUHK(香港中文大学) USTC(中国科学技术大学) SCUT(华南理工大学) HKU(香港大学) NTU(南洋理工大学) PKU(北京大学) SJTU(上海交通大学) CMU(卡内基梅隆大学) THU(清华大学) HUST(华中科技大学) MBZUAI(穆罕默德·本·扎耶德人工智能大学)

专题命中 软件智能体 :agentic(title,abstract);agent(abstract);分类 cs.AI

AI总结 VideoCoCo作为智能体双引擎框架,以可执行Blender代码为过程级思维链,构建专属数据集提升视频编辑器适配性,在两个基准上显著优于基线,实现了高质量物理一致性视频生成。

Comments 15 pages, 3 figures, and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07147 2026-08-10 cs.AI 新提交 83%

DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training

DiDPO:用于编码智能体训练的差异内差异策略优化

Xucong Wang, Zhe Zhao, Liheng Yu, Di Wu, Xiaofeng Cao, Pengkun Wang

机构 * University of Science and Technology of China (USTC)(中国科学技术大学) Stanford University(斯坦福大学) Suzhou Institute for Advanced Research, USTC(中国科学技术大学苏州高等研究院) Tongji University(同济大学)

专题命中 软件智能体 :agent(title,abstract);agentic(abstract);分类 cs.AI

AI总结 针对编码智能体训练的细粒度信用分配难题,提出DiDPO方法,在Qwen2.5-7B-Coder上性能超可比方法10%以上,开源了支持多种RL方法的verl-code代码库。

Comments 16 pages, 6 figures, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05959 2026-08-07 cs.SE 新提交 83%

AgentExecutor: Partial Code Execution via Agentic Context Generation

AgentExecutor:基于智能体上下文生成的部分代码执行工具

Junkai Chen, Chengran Yang, Xing Hu, Zhenhao Li, Xin Xia, David Lo

专题命中 软件智能体 :agentic(title);agent(abstract);multi-agent(abstract);分类 cs.SE

AI总结 本文提出多智能体框架AgentExecutor,通过三阶段设计和自适应优化策略提升部分代码执行效果,在两个数据集上的覆盖度、执行时间和成本均优于现有最优方法Treefix。

Comments ASE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04968 2026-08-06 cs.LG 新提交 83%

EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement

EvolveNet:智能体自我改进的协作式工具链进化

Jun Nie, Yonggang Zhang, Qianshu Cai, Yiu-ming Cheung, Xinmei Tian, Bo Han

机构 * Hong Kong Baptist University(香港浸会大学) University of Science and Technology of China(中国科学技术大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 软件智能体 :agent(title,abstract);agentic(abstract);分类 cs.LG

AI总结 EvolveNet提出协作式工具链进化范式,通过数据本地智能体的程序适配组合实现共享工具链改进,在五类任务场景均获增益,异构工作负载下效果更显著。

Comments 20 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03585 2026-08-05 cs.AI cs.CY 新提交 83%

From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities

从社交编码到智能体编码:开源社区中的生产力与关系重构

Mengying Zhou, Yongjie Yin, Yang Chen

专题命中 软件智能体 :agentic(title);agent(abstract);multi-agent(abstract);分类 cs.AI

AI总结 该研究基于含1084名开发者的GitHub数据模拟发现,编码智能体(CA)可提升开源社区生产力,但采用率低且收益集中于活跃开发者,还会减少人类直接交互、降低公共知识实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26275 2026-08-05 cs.CL 版本更新 83%

SPEAR: Code-Augmented Agentic Prompt Optimization

SPEAR: 代码增强的智能体提示优化

Mengyin Lu, Cong Feng, Huimin Han, Guangming Lu, Yu Sun, Xiaonan Ding, Shihui Long, Fengyi Li, Tanvi Motwani

机构 * LinkedIn Corporation(领英公司)

专题命中 软件智能体 :agentic(title,abstract);agent(abstract);分类 cs.CL

AI总结 提出SPEAR方法,将代码执行作为智能体工具进行提示优化,通过Python沙箱实现结构错误分析,并在工业任务和基准测试中取得显著提升。

Comments 19 pages, 3 figures, EMNLP 2026 submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11589 2026-08-05 cs.SE 83%

A Study of Library Usage in Agent-Authored Pull Requests

对代理生成的拉取请求中图书馆使用情况的研究

Lukas Twist, Jie M. Zhang

专题命中 软件智能体 :agent(title,abstract);agentic(abstract);分类 cs.SE

AI总结 研究探讨了代理生成的拉取请求中图书馆使用情况,发现代理频繁导入图书馆但很少引入新依赖项,且在引入时遵循严格的版本控制实践。

Comments 5 pages, 3 tables. Accepted at MSR '26 (Challenge Track)

Journal ref 23rd International Conference on Mining Software Repositories: MSR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27309 2026-07-31 cs.SE 新提交 83%

SIGIL: Compiling Agent Skills into Typed Harnesses

SIGIL:将智能体技能编译为类型化工具框架

Jayanaka Dantanarayana, Savini Kashmira, Lingjia Tang, Jason Mars

专题命中 软件智能体 :agent(title,abstract);agentic(abstract);分类 cs.SE

AI总结 SIGIL通过Skill Compilation将散文式智能体技能编译为类型化工具框架,提升了智能体执行技能要求步骤的比例、完整过程的频率并减少token使用,且效果与模型无关。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26563 2026-07-29 cs.SE 版本更新 83%

TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems

TrajAudit:智能体编码系统的自动化故障诊断

Minxing Wang, Xiaofei Xie, Yintong Huo

专题命中 软件智能体 :agentic(title,abstract);agent(abstract);分类 cs.SE

AI总结 针对仓库级编码轨迹中噪声多、上下文长的问题,提出TrajAudit框架,通过模式匹配过滤和测试报告先验知识辅助,实现高效故障诊断,在RootSE基准上定位准确率提升24.4个百分点,令牌消耗降低18%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09833 2026-07-28 cs.HC cs.AI cs.CY 版本更新 83%

CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

CollabSkill: 评估真实世界任务中的人机协作

Yijia Shao, Zora Zhiruo Wang, Neel Ahuja, Yicheng Wang, Bowen Liu, Diyi Yang

机构 * Stanford University(斯坦福大学) Carnegie Mellon University(卡内基梅隆大学) University of Rochester(罗切斯特大学) Individual Researcher(独立研究者)

专题命中 软件智能体 :agent(title,abstract);AI agent(abstract);分类 cs.AI

AI总结 提出CollabSkill框架,通过配对真实工人与AI代理执行职业任务,利用贝叶斯技能评级系统量化人机贡献,揭示Claude Code排名第一且实践经验是协作技能的主要驱动力。

Comments 10 pages of main paper, COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22114 2026-07-28 cs.SE 版本更新 83%

Automated Lemma Discovery in Agentic Program Verification

代理程序验证中的引理发现

Huan Zhao, Haoxin Tu, Zhengyao Liu, Martin C. Rinard, Abhik Roychoudhury

专题命中 软件智能体 :agentic(title,abstract);agent(abstract);分类 cs.SE

AI总结 本文提出LemmaNet代理,通过分析源代码和规格,发现辅助引理以提升程序验证的正确性保证。

Comments 13 pages. To appear in ASE 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18057 2026-07-21 cs.SE 新提交 83%

Test Coverage Analysis of Agentic Pull Requests

智能拉取请求的测试覆盖分析

Atish Kumar Dipongkor, Talank Baral, Wing Lam, Kevin Moran

专题命中 软件智能体 :agentic(title,abstract);agent(abstract);分类 cs.SE

AI总结 研究人工智能编码代理生成的拉取请求的测试覆盖情况,分析4882个相关请求,发现代理包含测试更改比例低,现有测试覆盖不足,代理编写测试仅在少数请求中提升覆盖,错误处理结构测试严重不足,为此提出相关改进措施。

Comments 12 pages, to appear 42nd International Conference on Software Maintenance and Evolution

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17937 2026-07-21 cs.SE 新提交 83%

When and How Context Rot Appears in Coding Agents: A White-Box Study of Agent Skills in Code Auditing

在长上下文情况下代理技能如何失效:代码审计中的白盒研究

Yue Xue

专题命中 软件智能体 :agent(title,abstract);workflow(abstract);分类 cs.SE

AI总结 研究长上下文下代码审计中代理技能失效问题,通过改变上下文分类故障位置,对比不同条件下Codex等运行情况,发现长上下文影响大,不同任务表现有别,外部检查表效果好,编码代理支架有帮助,给出故障分类和实证案例。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16740 2026-07-21 cs.SE 新提交 83%

Agentic Code Review in the Terminal: A Trajectory-Level Analysis of Behavior, Cost, and Human-Alignment

终端中的智能代码审查:行为、成本和人工对齐的轨迹级分析

Wachiraphan Charoenwet, Kla Tantithamthavorn, Patanamon Thongtanunam, Hong Yi Lin, Minwoo Jeong, Ming Wu

专题命中 软件智能体 :agentic(title,abstract);planning(abstract);分类 cs.SE

AI总结 研究终端环境中智能代码审查的行为、成本等,基于审查者轨迹分析,发现智能审查者审查精度高但有探索等开销,成功审查与规划等有关,凸显轨迹感知和成本敏感评估对未来此类系统的好处。

详情

展开后加载摘要…

URL PDF HTML 收藏