arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 3188 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 软件智能体 3188 篇

2605.25438 2026-07-08 econ.GN q-fin.EC 版本更新 82%

Agentic Delegation and the Language Frontier of Software Developers: A Model and Evidence from Claude Code on GitHub

超越训练范围编码:Claude Code 与软件开发者的技术前沿

Alexander Quispe, Kevin Xu

专题命中 软件智能体 :agentic(title,abstract);agent(abstract)

AI总结 利用双重稳健估计器分析 Claude Code 的逐步推出,发现 AI 编码助手显著增加了开发者的月度提交数、贡献仓库数、使用语言数及语言熵,且累积语言效应随时间增长,表明 AI 降低了技术切换障碍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03105 2026-07-07 quant-ph 新提交 82%

ORBIT-Q: Dual-axis benchmarking of autonomous agents in scientific quantum programming

ORBIT-Q:科学量子编程中自主代理的双轴基准测试

Shi-Xin Zhang, Yu-Qin Chen

专题命中 软件智能体 :autonomous agent(title,abstract);agent(abstract)

AI总结 为解决自主编码代理在科学计算中缺乏严格验证范式问题,引入ORBIT-Q。其核心是精心策划的量子工作流程套件,结合严格验证管道进行双轴比较,评估结果为相关研究提供基准。

Comments 7.5 pages, 4 figures with references and supplementary information

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15657 2026-07-07 cs.AR 82%

Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification

理解代理硬件验证中推理时间标记分配与覆盖限制

Vihaan Patel, Vidya Chhabria, Aman Arora

专题命中 软件智能体 :agentic(title,abstract);agent(abstract)

AI总结 本文研究了代理硬件验证中推理时间标记分配与覆盖限制,通过两个层级的代理框架,分析覆盖孔径并改进效率,展示领域专业化如何提升覆盖效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19501 2026-07-01 cs.AI cs.CL cs.LG q-fin.RM 新提交 82%

DeXposure-Claw: An Agentic System for DeFi Risk Supervision

DeXposure-Claw: 一个用于DeFi风险监管的智能体系统

Aijie Shu, Bowei Chen, Wenbin Wu, Cathy Yi-Hsuan Chen, Fengxiang He

机构 * University of Edinburgh(爱丁堡大学) University of Glasgow(格拉斯哥大学) University of Cambridge(剑桥大学)

专题命中 软件智能体 :agentic(title,abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 针对DeFi监管中LLM智能体易误报的问题,提出DeXposure-Claw系统,通过图时间序列基础模型预测风险网络,结合确定性监控和置信度门控生成可审计监管票据,并构建六轴评估基准DeXposure-Bench,实验验证有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13692 2026-07-01 cs.OS cs.MA 版本更新 82%

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System

ThunderAgent: 一个简单、快速且程序感知的智能推理系统

Hao Kang, Ziyang Li, Weili Xu, Xinyu Yang, Yinfang Chen, Junxiong Wang, Beidi Chen, Tushar Krishna, Chenfeng Xu, Simran Arora

专题命中 软件智能体 :agentic(title,abstract);workflow(abstract)

AI总结 提出ThunderAgent系统,通过LLM程序抽象统一管理KV缓存和工具资源,实现程序感知调度,在编码、路由和科学发现任务中吞吐量提升1.5-3.9倍,磁盘内存节省高达4.2倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22484 2026-06-23 cs.HC 新提交 82%

Governed AI-Assisted Engineering: Graduated Human Oversight for Agentic Code Generation in Regulated Domains

受治理的AI辅助工程:受监管领域中代理代码生成的渐进式人工监督

Richard Kang

专题命中 软件智能体 :agentic(title,abstract);autonomous agent(abstract)

AI总结 针对受监管行业中代理AI编码系统的治理挑战,提出GAIE框架,通过三级渐进式人工监督模型(人机协同、人机监控、自动监控)平衡编码速度与合规性,在保留91%编码速度的同时确保监管证据覆盖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21804 2026-06-23 cs.SE cs.AI cs.CL 新提交 82%

Is Agent Code Less Maintainable Than Human Code?

智能体代码比人类代码更不易维护吗?

Shaswat Patel, Betty Li Hou, Arun Purohit, Kai Xu, Jane Pan, He He, Valerie Chen

机构 * New York University(纽约大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI、cs.CL、cs.SE

AI总结 研究通过CodeThread框架对比智能体与人类代码在维护场景下的表现,发现基于智能体代码解决任务的成功率下降高达13.1%,传统可维护性指标无法解释差异,而输入验证、错误处理等行为差异是关键因素。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06530 2026-06-23 cs.AR 新提交 82%

RTLScout: Joint Agentic Code and Synthesis Optimization for Efficient Digital Circuits

RTLScout:面向高效数字电路的联合智能体代码与综合优化

Felix Arnold, Ryan Amaudruz, Dimitrios Tsaras, Renzo Andri, Lukas Cavigelli

专题命中 软件智能体 :agentic(title,abstract);agent(abstract)

AI总结 提出RTLScout系统,结合LLM驱动智能体设计与电路级综合优化及算术架构扫描,通过多轮精英池框架迭代优化RTL设计,在浮点乘法器上实现面积减少35%、延迟减少45%。

Comments Updated Appendix Table 4: corrected three Verilog baseline values. Conclusions unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16287 2026-06-17 cs.CR 新提交 82%

Dynamic Malicious Skills in Agentic AI

智能体AI中的动态恶意技能

Tianhao Chen, Zhengyuan Jiang, Yuepeng Hu, Yebei Gou, Neil Zhenqiang Gong

专题命中 软件智能体 :agentic(title,abstract);agent(abstract)

AI总结 研究智能体AI中通过自然语言文档注入恶意指令实现动态恶意技能的攻击方法,并提出基于操作系统内核只读挂载的系统级防御。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07889 2026-06-09 cs.LG cs.AI cs.CL 新提交 82%

Strained Coherence: A Pre-Failure Signal in Coding Agent Execution Trajectories

应变连贯性:编码代理执行轨迹中的故障前信号

Marut Pandya, Kasey Zhang, Baiqing Lyu

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 提出“应变连贯性”模式,即编码代理识别到问题但仍按原计划行动,通过构建Claude Sonnet 4.6检测器在44条轨迹上实现94%故障预测精度,优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00866 2026-06-02 cs.OS 82%

Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI

空闲是相对的:利用工具调用空闲窗口在MORI代理系统中进行卸载

Tian Xia, Hanchen Li, Zhifei Li, Xiaokun Chen, Hao Kang, Yifan Qiao, Yi Xu, Ion Stoica

专题命中 软件智能体 :agentic(title,abstract);agent(abstract)

AI总结 提出MORI系统,通过将代理程序按空闲程度连续排序并动态划分内存层级,解决了LLM代理工作负载中KV缓存卸载效率低下的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18451 2026-05-19 cs.CV cs.GR 82%

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis

Code-as-Room: 通过代理代码合成从俯视图图像生成3D房间

Yixuan Yang, Zhen Luo, Wanshui Gan, Jinkun Hao, Junru Lu, Jinghao Yan, Zhaoyang Lyu, Xudong Xu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) Southern University of Science and Technology(南方科技大学) University of Warwick(沃里克大学)

专题命中 软件智能体 :agentic(title,abstract);agent(abstract)

AI总结 本文提出Code-as-Room框架,通过结构化执行 harness 生成3D房间,利用Blender代码表示房间,并引入专门的代码基3D房间合成基准进行评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08112 2026-05-12 cs.SE cs.AI cs.CE cs.LG cs.LO 82%

Context-Augmented Code Generation: How Product Context Improves AI Coding Agent Decision Compliance by 49%

上下文增强的代码生成:产品上下文如何通过49%提高AI编码代理的决策合规性

Drew Dillon, Kasyap Varanasi

机构 * Brief(简述)

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI、cs.LG、cs.SE

AI总结 本文提出一个衡量决策合规性的基准,通过产品上下文检索系统提升AI编码代理的决策合规性,实验结果显示合规率从46%提升至95%。

Comments 16 pages, 3 figures, 16 tables. Benchmark repository: https://github.com/brief-hq/dcbench

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06158 2026-05-08 cs.CR 82%

Stateful Agent Backdoor

具备状态的代理后门

Zhengchunmin Dai, Jiaxiong Tang, Liantao Wu, Peng Sun, Honglong Chen

专题命中 软件智能体 :agent(title,abstract)

AI总结 本文提出一种具备状态的代理后门,通过持久化组件跨多个会话维持状态,实现自主递增执行。该方法基于Mealy机建模,并通过分解框架提升攻击成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16520 2026-04-21 cs.HC 82%

AgentClick: A Skill-Based Human-in-the-Loop Review Layer for Terminal AI Agents

AgentClick: 一个基于技能的人机协作审查层用于终端AI代理

Haomin Zhuang, Hanwen Xing, Xiangliang Zhang

专题命中 软件智能体 :AI agent(title,abstract);agent(abstract)

AI总结 AgentClick通过浏览器界面提供结构化交互,降低非专家用户与AI代理协作的门槛,提升效率和质量。

Comments Accepted to ACM CAIS 2026 System Demonstrations. Conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14234 2026-04-16 cs.CV 82%

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

ViBES:一个具有行为智能的3D虚拟身体对话代理

Juze Zhang, Changan Chen, Xin Chen, Heng Yu, Tiange Xiang, Ali Sartaz Khan, Shrinidhi K. Lakshmikanth, Ehsan Adeli

机构 * Stanford University(斯坦福大学) ByteDance(字节跳动)

专题命中 软件智能体 :agent(title,abstract);agentic(abstract)

AI总结 ViBES通过联合规划语言和运动,实现对话条件下的身体动作生成,提升了多轮对话中的社会交互能力。

Comments Project page: https://ai.stanford.edu/~juze/ViBES/. Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10800 2026-04-14 cs.SE cs.AI cs.CR cs.LG cs.PL 82%

Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis

在修复前验证:面向可信跨语言代码分析的代理执行 grounding

Jugal Gajjar

专题命中 软件智能体 :agentic(title,abstract);分类 cs.AI、cs.LG、cs.SE

AI总结 本文提出统一的跨语言漏洞生命周期框架,通过LLM驱动的三个阶段:混合结构-语义检测、执行 grounding 的代理验证和验证-aware 的迭代修复,确保修复前有执行验证。框架利用uAST和图神经网络融合,实现高准确率的漏洞检测与跨语言修复。

Comments 20 pages (13 main + 7 appendices), 9 figures, 10 tables. Submitted to NeurIPS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04238 2026-04-07 cs.PL 82%

Agentic Code Optimization via Compiler-LLM Cooperation

通过编译器-LLM协作实现代理代码优化

Benjamin Mikek, Danylo Vashchilenko, Bryan Lu, Panpan Xu

专题命中 软件智能体 :agentic(title);agent(abstract);multi-agent(abstract)

AI总结 本文提出编译器-LLM协作方法,结合传统编译优化与LLM生成,提升代码性能,实验显示比现有方法快1.25倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29664 2026-04-01 cs.CV 82%

CutClaw: Agentic Hours-Long Video Editing via Music Synchronization

CutClaw:通过音乐同步实现代理长时间视频编辑

Shifang Zhao, Yihan Hu, Ying Shan, Yunchao Wei, Xiaodong Cun

机构 * Beijing Jiaotong University(北京交通大学) GVC Lab, Great Bay University(大湾区大学GVC实验室) ARC Lab, Tencent(腾讯ARC实验室)

专题命中 软件智能体 :agentic(title);agent(abstract);multi-agent(abstract)

AI总结 本文提出CutClaw,一种多代理框架,利用多模态语言模型自动编辑长时间原始素材为有意义的短视频,通过音乐同步和视觉优化提升视频质量。

Comments Project Code: https://github.com/GVCLab/CutClaw

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26990 2026-03-31 hep-ph 82%

Agentic Diagrammatica: Towards Autonomous Symbolic Computation in High Energy Physics

代理图示学:迈向高能物理中自主符号计算

Tony Menzo, Alexander Roman, George T. Fleming, Sergei Gleyzer, Konstantin T. Matchev, Stephen Mrenna

专题命中 软件智能体 :agentic(title,abstract);agent(abstract)

AI总结 本文提出Diagrammatica,通过HEPTAPOD框架扩展实现LLM代理的多步骤理论计算。通过工具约束计算和知识接地方法,实现符号计算的精确性与可验证性,验证了在高能物理中的应用效果。

Comments 54 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14417 2026-03-17 cs.CY cs.AI cs.CL cs.LG 82%

Questionnaire Responses Do not Capture the Safety of AI Agents

问卷回复无法捕捉AI代理的安全性

Max Hellrigel-Holderbaum, Edward James Young

专题命中 软件智能体 :AI agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文指出问卷式评估无法准确衡量AI代理的安全性,因LLM在情景中的表现与实际代理行为存在差异,导致评估方法缺乏有效性。

Comments 31 pages, 11 pages main text

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05621 2026-03-12 cs.RO cs.AI cs.CL cs.LG cs.MA 82%

RACAS: Controlling Diverse Robots With a Single Agentic System

RACAS:通过单一代理系统控制多样化机器人

Dylan R. Ashley, Jan Przepióra, Yimeng Chen, Ali Abualsaud, Nurzhan Yesmagambet, Shinkyu Park, Eric Feron, Jürgen Schmidhuber

机构 * Center of Excellence in Generative AI, King Abdullah University of Science and Technology (KAUST), Saudi Arabia(沙特王国科学与技术大学生成人工智能卓越中心) Dalle Molle Institute for Artificial Intelligence Research (IDSIA), Switzerland(人工智能研究达勒莫利 institute) Università della Svizzera italiana (USI), Switzerland(瑞士意大利大学) Scuola universitaria professionale della Svizzera italiana (SUPSI), Switzerland(瑞士意大利专业大学) Robotics, Intelligent Systems, and Control Lab, King Abdullah University of Science and Technology (KAUST), Saudi Arabia(机器人、智能系统与控制实验室,沙特王国科学与技术大学(KAUST)) Department of Process Control, AGH University of Krakow, Poland(波兰克拉科夫AGH大学过程控制系)

专题命中 软件智能体 :agentic(title,abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 RACAS通过单一代理系统实现跨平台机器人控制,无需修改代码或模型,有效降低机器人原型开发难度。

Comments 7 pages in main text + 1 page of appendices + 1 page of references, 5 figures in main text + 1 figure in appendices, 2 tables in main text; source code available at https://github.com/janprz11/robot-agnostic-control

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01918 2026-02-10 cs.HC 82%

When Workout Buddies Are Virtual: AI Agents and Human Peers in a Longitudinal Physical Activity Study

当训练伙伴是虚拟的:AI代理与人类同伴在长期身体活动研究中的应用

Alessandro Silacci, Mauro Cherubini, Arianna Boldi, Amon Rapp, Maurizio Caon

专题命中 软件智能体 :AI agent(title,abstract);agent(abstract)

AI总结 本文研究了AI代理与人类同伴在长期身体活动中的影响,发现AI能提供更稳定的鼓励,而人类同伴则增强社会存在感,两者互补促进持续锻炼。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20211 2025-10-24 cs.SE cs.AI cs.LG 82%

Automated Cloud Infrastructure-as-Code Reconciliation with AI Agents

Zhenning Yang, Hui Guan, Victor Nicolet, Brandon Paulsen, Joey Dodds, Daniel Kroening, Ang Chen

机构 * University of Michigan(密歇根大学) Amazon Web Services(亚马逊网络服务)

专题命中 软件智能体 :AI agent(title);agentic(abstract);分类 cs.AI、cs.LG、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09740 2025-09-15 q-bio.QM cs.AI cs.CL cs.LG 82%

HypoGeneAgent: A Hypothesis Language Agent for Gene-Set Cluster Resolution Selection Using Perturb-seq Datasets

Ying Yuan, Xing-Yue Monica Ge, Aaron Archer Waterman, Tommaso Biancalani, David Richmond, Yogesh Pandit, Avtar Singh, Russell Littman, Jin Liu, Jan-Christian Huetter, Vladimir Ermakov

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13905 2025-09-10 cs.AR 82%

Spec2RTL-Agent: Automated Hardware Code Generation from Complex Specifications Using LLM Agent Systems

Zhongzhi Yu, Mingjie Liu, Michael Zimmer, Yingyan Celine Lin, Yong Liu, Haoxing Ren

专题命中 软件智能体 :agent(title,abstract);multi-agent(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04227 2025-06-18 cs.HC cs.AI cs.CL cs.LG 82%

Agent Laboratory: Using LLM Agents as Research Assistants

Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, Emad Barsoum

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15625 2025-06-02 cs.LG cs.AI cs.CL cs.DC 82%

Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces

Anjiang Wei, Allen Nie, Thiago S. F. X. Teixeira, Rohan Yadav, Wonchan Lee, Ke Wang, Alex Aiken

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11243 2025-04-16 cs.AI cs.CL cs.LG 82%

Towards Automated Safety Requirements Derivation Using Agent-based RAG

Balahari Vignesh Balu, Florian Geissler, Francesco Carella, Joao-Vitor Zacchi, Josef Jiru, Nuria Mata, Reinhard Stolle

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI、cs.CL、cs.LG

Comments 9 pages, 3 figures

Journal ref Proceedings of the AAAI-make Spring Symposium, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17703 2025-04-07 cs.RO 82%

RAIDER: Tool-Equipped Large Language Model Agent for Robotic Action Issue Detection, Explanation and Recovery

Silvia Izquierdo-Badiola, Carlos Rizzo, Guillem Alenyà

专题命中 软件智能体 :agent(title,abstract);agentic(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏