arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

AI Agent

智能体、工具调用、规划、工作流、多智能体和自主任务执行。

共收录 3188 信号源:cs.AI, cs.CL, cs.LG, cs.SE

1. 软件智能体 3188 篇

2101.00916 2021-01-05 cs.CL cs.AI 81%

How to Train Your Agent to Read and Write

Li Liu, Mengge He, Guanghui Xu, Mingkui Tan, Qi Wu

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.05064 2020-11-12 cs.AI cs.CY cs.LG stat.ML 81%

What Did You Think Would Happen? Explaining Agent Behaviour Through Intended Outcomes

Herman Yau, Chris Russell, Simon Hadfield

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.10910 2019-07-02 cs.LG cs.CL stat.ML 81%

Creating A Neural Pedagogical Agent by Jointly Learning to Review and Assess

Youngnam Lee, Youngduck Choi, Junghyun Cho, Alexander R. Fabbri, Hyunbin Loh, Chanyou Hwang, Yongku Lee, Sang-Wook Kim, Dragomir Radev

专题命中 软件智能体 :agent(title,abstract);分类 cs.CL、cs.LG

Comments 9 pages, 9 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.02739 2018-10-04 cs.RO cs.AI cs.LG 81%

Discovering space - Grounding spatial topology and metric regularity in a naive agent's sensorimotor experience

Alban Laflaquière, J. Kevin O'Regan, Bruno Gas, Alexander Terekhov

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI、cs.LG

Comments 59 pages, 16 figures, submitted to Neural Networks

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24129 2026-06-24 cs.AI 新提交 80%

OmniPath: A Multi-Modal Agentic Framework for Auditing Wheelchair Accessibility

OmniPath: 用于审计轮椅可达性的多模态智能体框架

ASM Mobarak Hossain, Nadim Mahmud, Vaskar Raychoudhury, Md Osman Gani

机构 * Causal AI Lab, Department of Information Systems, University of Maryland Baltimore County(因果AI实验室,信息系统系,马里兰大学巴尔的摩分校) Department of Computer Science & Software Engineering, Miami University(计算机科学与软件工程系,迈阿密大学)

专题命中 软件智能体 :agentic(title,comments);agent(abstract);分类 cs.AI

AI总结 提出OmniPath框架,融合OSM网络拓扑与机载LiDAR数据,通过虚拟遍历量化坡度、横坡等物理摩擦点,按ADA标准分级评估轮椅可达性,经实地验证有效识别标准地图遗漏的障碍。

Comments 10 pages, 13 figures. Submitted to IEEE COMPSAC 2026. OmniPath: A Multi-Modal Agentic Framework for Auditing Wheelchair Accessibility

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09682 2026-06-09 cs.LG cs.DC cs.PF 新提交 80%

AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis

AutoMegaKernel:用于自我重定目标超内核合成的静态检查代理框架

Jaber Jaber, Osama Jaber

机构 * RightNow AI

专题命中 软件智能体 :agent(title,abstract);分类 cs.LG

AI总结 提出AutoMegaKernel系统,将Llama模型编译为单个持久CUDA内核,通过静态调度验证器确保无死锁和无竞争,自动生成10种模型正确超内核,并在NVIDIA推理卡上以W8A16精度超越cuBLAS bf16。

Comments 18 pages, 5 figures. Open-source code, data, and agent harness: https://github.com/RightNow-AI/AutoMegaKernel

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28916 2026-06-01 astro-ph.IM cs.AI cs.HC 80%

First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope

应用于爱因斯坦望远镜模拟数据分析的智能体AI首次头对头比较

Gianluca Inguglia

机构 * Anthropic OpenAI

专题命中 软件智能体 :agentic(title,abstract);分类 cs.AI

AI总结 本文首次直接比较了Claude Code和Codex两种智能体AI系统在无人干预下自主执行引力波数据分析管线的行为、科学结果和计算成本,揭示了速度与可审计性、指令解释差异等关键问题。

Comments Version 2; includes the report autonomoulsy written in PRD style by agentic AI systems as supplemental material

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26508 2026-05-27 q-fin.RM cs.AI 80%

Foundations of a Time-Consistent Counterfactual Actuarial Runtime for Autonomous AI Agents

自主AI智能体时间一致性反事实精算运行时的基础

Hao-Hsuan Chen

机构 * Department of Risk Management and Insurance(风险管理与保险系)

专题命中 软件智能体 :AI agent(title,abstract);分类 cs.AI

AI总结 本文提出一种精算运行时层,通过为每个动作分配时间一致的反事实风险费用,并建立边界内无套利和预算保证,为自主AI智能体提供基础数学框架。

Comments 10 pages. Foundational paper of a multi-paper program on actuarial runtime for autonomous AI agents; previously posted on SSRN (id 6761960). Empirical companion: arXiv:2605.25632. Proof companions included as ancillary files

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08397 2025-10-17 eess.IV cs.AI cs.CV 80%

VoxelPrompt: A Vision Agent for End-to-End Medical Image Analysis

Andrew Hoopes, Neel Dey, Victor Ion Butoi, John V. Guttag, Adrian V. Dalca

机构 * Massachusetts Institute of Technology(麻省理工学院) Massachusetts General Hospital(麻省总医院) Harvard Medical School(哈佛医学院)

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI

Comments 22 pages, vision-language agent, medical image analysis, neuroimage foundation model

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11950 2026-07-09 cs.SE cs.AI cs.CL cs.CR 版本更新 80%

AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection

AnyPoC:用于可扩展LLM基于Bug检测的通用证明-概念测试生成

Zijie Zhao, Chenyuan Yang, Weidong Wang, Yihan Yang, Ziqi Zhang, Lingming Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 软件智能体 :agent(abstract);agentic(abstract);multi-agent(abstract);分类 cs.AI、cs.CL、cs.SE

AI总结 AnyPoC通过多代理框架生成可执行的证明-概念测试,以验证候选Bug报告,提升自动化Bug检测的实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10478 2026-06-10 cs.CV 新提交 80%

3D-CoS: A New 3D Reconstruction Paradigm Based on VLM Code Synthesis

3D-CoS:基于VLM代码合成的新型3D重建范式

Yuhao Wang, Puyi Wang, Linjie Li, Zhengyuan Yang, Kevin Qinghong Lin, Yu Cheng

机构 * Shanghai Jiao Tong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Microsoft(微软) University of Oxford(牛津大学)

专题命中 软件智能体 :agent(abstract,abstract_cn);planning(abstract);workflow(abstract)

AI总结 提出3D代码合成(3D-CoS)范式,将3D资产表示为可执行的Blender代码,利用VLM进行程序化重建,实现高可控性和局部编辑能力。

Comments Preprint. 24 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16819 2026-05-19 cs.CL cs.AI cs.LG 80%

AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents

AgentKernelArena: GPU核优化代理的通用化意识基准测试

Sharareh Younesian, Wenwen Ouyang, Sina Rafati, Mehdi Rezagholizadeh, Sharon Zhou, Ji Liu, Yue Liu, Yuchen Yang, Hao Li, Ziqiong Liu, Dong Li, Vikram Appia, Zhenyu Gu, Emad Barsoum

机构 * AMD

专题命中 软件智能体 :agent(abstract,abstract_cn);agentic(abstract);分类 cs.AI、cs.CL、cs.LG

AI总结 本文提出AgentKernelArena,一个用于评估GPU核优化代理的开源基准,通过隔离工作区和统一评分机制,测试代理在不同任务和硬件目标上的性能和通用化能力,发现大多数任务在正确性和编译效率上表现优异,但在PyTorch到HIP的转换任务中存在显著的正确性下降。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00803 2026-05-04 cs.SE cs.AI cs.CL 80%

Can Coding Agents Reproduce Findings in Computational Materials Science?

编码代理能否在计算材料科学中复现研究成果?

Ziyang Huang, Yi Cao, Ali K. Shargh, Jing Luo, Ruidong Mei, Mohd Zaki, Zhan Liu, Wyatt Bunstine, William Jurayj, Somdatta Goswami, Tyrel McQueen, Michael Shields, Jaafar El-Awady, Paulette Clancy, Benjamin Van Durme, Nicholas Andrews, William Walden, Daniel Khashabi

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 软件智能体 :agent(abstract);workflow(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.SE

AI总结 研究探讨编码代理在计算材料科学领域复现研究结果的能力,提出AutoMat基准测试,发现当前LLM代理在复现复杂科学流程时表现有限,存在流程不完整、方法偏差等问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07865 2026-02-09 cs.SE cs.AI cs.CL cs.MA 80%

LLM-Powered Fully Automated Chaos Engineering: Towards Enabling Anyone to Build Resilient Software Systems at Low Cost

基于大语言模型的完全自动化混沌工程:实现任何人以低成本构建容错软件系统

Daisuke Kikuta, Hiroki Ikeuchi, Kengo Tajiri

机构 * NTT, Inc., Japan(日本NTT公司)

专题命中 软件智能体 :planning(abstract);workflow(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.SE

AI总结 本文提出ChaosEater系统,利用大语言模型自动化混沌工程周期,使任何人能低成本构建容错软件系统。

Comments Accepted at ASE 2025 NIER Track. The code is available at https://github.com/ntt-dkiku/chaos-eater

Journal ref 2025 40th IEEE/ACM International Conference on Automated Software Engineering (ASE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23045 2025-12-09 cs.AI cs.CL cs.SE 80%

Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents

Kimi-Dev:无代理训练作为SWE-代理的技能先验

Zonghan Yang, Shengjie Wang, Kelin Fu, Wenyang He, Weimin Xiong, Yibo Liu, Yibo Miao, Bofei Gao, Yejie Wang, Yingwei Ma, Yanhao Li, Yue Liu, Zhenxing Hu, Kaitai Zhang, Shuyi Wang, Huarong Chen, Flood Sung, Yang Liu, Yang Gao, Zhilin Yang, Tianyu Liu

机构 * Moonshot AI THU(清华大学) PKU(北京大学) UCAS(中国科学技术大学) BUPT(北京邮电大学) NUS(新加坡国立大学)

专题命中 软件智能体 :agent(abstract);workflow(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.SE

AI总结 Kimi-Dev通过无代理训练生成技能先验,使SWE-代理在SWE-bench Verified上取得60.4%的优异成绩,并通过额外SFT适应达到48.6%的pass@1,与Claude 3.5 Sonnet相当。

Comments 68 pages. GitHub repo at https://github.com/MoonshotAI/Kimi-Dev

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01010 2025-12-02 cs.MA cs.AI cs.LG cs.SE physics.comp-ph physics.flu-dyn 80%

Chain of Unit-Physics: A Primitive-Centric Approach to Scientific Code Synthesis

单元物理链:一种以基础原理为中心的科学代码合成方法

Vansh Sharma, Venkat Raman

机构 * University of Michigan(密歇根大学)

专题命中 软件智能体 :agent(abstract);agentic(abstract);multi-agent(abstract);分类 cs.AI、cs.LG、cs.SE

AI总结 本研究提出单元物理链框架,通过以基础原理为中心的多代理系统,有效解决科学代码生成中的可靠性问题,实现高精度和高效能的代码生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12864 2025-10-16 cs.AI cs.CL cs.LG 80%

From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models

Imran Khan

机构 * Independent Researcher(独立研究者)

专题命中 软件智能体 :AI agent(abstract);autonomous agent(abstract);agentic(abstract);分类 cs.AI、cs.CL、cs.LG

Comments 13 pages. Code and data are available at https://github.com/strongSoda/LITERAL-TO-LIBERAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15003 2025-07-22 cs.SE cs.AI cs.CE cs.LG 80%

The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering

Hao Li, Haoxiang Zhang, Ahmed E. Hassan

机构 * Queen's University(女王大学)

专题命中 软件智能体 :agent(abstract);AI agent(abstract);agentic(abstract);分类 cs.AI、cs.LG、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00014 2025-07-02 cs.LG cs.AI cs.SE 80%

SWE-Bench-CL: Continual Learning for Coding Agents

Thomas Joshi, Shayan Chowdhury, Fatih Uysal

机构 * Columbia University(哥伦比亚大学)

专题命中 软件智能体 :agent(abstract);AI agent(abstract);tool-use(abstract);分类 cs.AI、cs.LG、cs.SE

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15572 2026-08-17 cs.LG cs.MA 版本更新 79%

Neural Network-Based Parameter Estimation of a Labour Market Agent-Based Model

基于神经网络的劳动力市场基于主体的模型参数估计

M Lopes Alves, Joel Dyer, Doyne Farmer, Michael Wooldridge, Anisoara Calinescu

机构 * Department of Computer Science(计算机科学系) Institute of New Economic Thinking(新经济思想研究所) University of Oxford(牛津大学)

专题命中 软件智能体 :agent(title,abstract);分类 cs.LG

AI总结 本文提出利用神经网络改进劳动力市场ABM参数估计方法,通过对比汇总统计量的有效性,验证了NN在参数恢复和效率提升方面的优势。

Comments Paper has been submitted without the consent of all 5 authors and without having its final version. Uploaded on arxiv at 17 Feb 2026 The version with proper review was only available on 19 March 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30777 2026-08-14 cs.SE 版本更新 79%

What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants

当LLM编码时会出现什么问题?表征自主代码助手的操作安全故障

Alif Al Hasan, Sumon Biswas

专题命中 软件智能体 :agentic(title);agent(abstract);分类 cs.SE

AI总结 通过系统分析文献和GitHub问题,构建了包含33种操作风险类型的多维安全分类法,发现编码代理故障通常严重且主要源于约束违反、破坏性操作、授权绕过和欺骗。

Comments This paper is accepted to the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026), Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12123 2026-08-13 cs.DC cs.AI cs.OS 新提交 79%

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control

就绪队列:在LLM智能体控制中限定GPU机会并避免主机往返

Josef Liyanjun Chen

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI

AI总结 该研究针对LLM智能体控制,提出就绪队列边界的形式化方法,通过实验验证设备驻留路由决策路径可提升效率,为GPU智能体控制确立了两个可测量门限。

Comments 14 pages, 4 figures. Includes formal proofs, trace provenance, and a reproducibility appendix. Code and artifacts: https://github.com/josefchen/ready-cohorts ; processed evidence: https://huggingface.co/datasets/josefchen/ready-

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01687 2026-08-12 cs.AI 版本更新 79%

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

CoEvoSkills: 通过共进化验证实现自进化代理技能

Hanrong Zhang, Shicheng Fan, Henry Peng Zou, Yankai Chen, Zhenting Wang, Jiayu Zhou, Chengze Li, Wei-Chieh Huang, Yifei Yao, Kening Zheng, Xue, Liu, Xiaoxiao Li, Philip S. Yu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) MBZUAI(穆罕默德·本·扎耶德人工智能大学) McGill University(麦吉尔大学) Columbia University(哥伦比亚大学) Zhejiang University(浙江大学) University of British Columbia(不列颠哥伦比亚大学)

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI

AI总结 本文提出CoEvoSkills框架,使代理能自主生成复杂技能包,通过共进化验证提供反馈,实现在SkillsBench上最高通过率和强泛化能力。

Comments COLM accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09650 2026-08-11 cs.IR cs.CL 新提交 79%

Listwise Cross-Encoder Fine-Tuning vs. Agentic Instruction Tuning for LLM Rerankers: A Systematic Study in Medical Procedure Reranking

用于LLM重排序器的列表式交叉编码器微调与智能体指令微调:医疗流程重排序的系统性研究

Matan Fainzilber, Shlomit Plavner

专题命中 软件智能体 :agentic(title,abstract);分类 cs.CL

AI总结 本文系统性对比列表式交叉编码器微调与智能体指令微调两种LLM重排序范式,在自建医疗流程重排序数据集上,发现1.09亿参数的ListNet微调交叉编码器性能优于40亿参数的Qwen3-Reranker-4B,参数量仅为后者1/37,相关成果已开源。

Comments 10 pages, 6 figures, 4 tables. Code available at https://github.com/matanf-healthee/listwise-crossencoder-reranking

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08793 2026-08-11 cs.CL 新提交 79%

Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents

面向异构编码智能体的技能的证据校准运行时重构

Xueping Gao

专题命中 软件智能体 :agent(title,abstract);分类 cs.CL

AI总结 本文提出Skill Runtime Intelligence系统,针对异构编码智能体重构技能生命周期阶段,在多场景实验中验证其能准确定位失败边界,为适配器资格认定提供支撑。

Comments 17 pages, 1 figure, 6 tables. Submitted to PROFES 2026. Code and artifacts: https://github.com/hellogxp/skill-runtime-intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18171 2026-08-11 cs.LG 版本更新 79%

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

FlashRT:用于引导智能体部署实时多模态应用的智能体框架

Krish Agarwal, Zhuoming Chen, Yanyuan Qin, Zhenyu Gu, Atri Rudra, Beidi Chen

机构 * Carnegie Mellon University(卡内基梅隆大学) AMD(超威半导体公司) University at Buffalo(纽约州立大学水牛城分校)

专题命中 软件智能体 :agent(title,abstract);分类 cs.LG

AI总结 研究实时多模态应用部署难题,提出FlashRT智能体框架。通过新范式引导编码智能体多阶段转换,将参考实现转为高效部署,在不同GPU上显著提升性能,尤其在专家优化不成熟平台更具扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14166 2026-08-11 cs.SE cs.CR cs.DC 版本更新 79%

Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives

停止即停止:测量和修复智能体框架控制原语中的执行差距

Sajjad Khan

专题命中 软件智能体 :agent(title,abstract);分类 cs.SE

AI总结 研究生产型大语言模型智能体框架控制原语执行差距,用无模型差分探针发现漏洞及差距,提出SOUNDGATE修复,能强制执行多种属性,端到端阻止违规,释放合法效果。

Comments 32 pages, 3 figures, 11 tables. Code: pip install soundgate (PyPI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29678 2026-08-10 cs.CL cs.DC cs.PF 版本更新 79%

TokTier: Exact Stateful CPU+GPU Tokenization for Agentic LLM Serving

TokTier:面向智能体大语言模型服务的精确有状态分词方案

Zhenyu Zhang, Zhichao Cao

机构 * Arizona State University(亚利桑那州立大学)

专题命中 软件智能体 :agentic(title);agent(abstract);分类 cs.CL

AI总结 本研究提出TokTier,一种面向智能体LLM服务的精确有状态分词方案,通过增量修复、GPU加速分词等技术,将首次令牌生成时间降低16%-34%,饱和请求速率提升至1821次/秒,性能远超现有方案。

Comments 26 pages. Code: https://github.com/asu-idi/toktier

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04440 2026-08-06 q-bio.QM cs.AI 交叉投稿 79%

AutoProteinEngine: A Large Language Model Driven Agent Framework for Multimodal AutoML in Protein Engineering

AutoProteinEngine:一种用于蛋白质工程多模态自动机器学习的大语言模型驱动智能体框架

Yungeng Liu, Zan Chen, Yu Guang Wang, Yiqing Shen

专题命中 软件智能体 :agent(title,abstract);分类 cs.AI

AI总结 研究针对生物学家缺乏计算知识难以应用深度学习开展蛋白质工程的问题,提出AutoProteinEngine框架,集成LLMs与AutoML,经两项任务验证其性能优于传统方法,降低了深度学习应用门槛。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00891 2026-08-04 cs.AI 新提交 79%

CADIR: A Cross-Backend Editable Intermediate Representation for Agentic CAD Generation

CADIR:一种用于智能体CAD生成的跨后端可编辑中间表示

Yu Liu, Jingzhe Ni, Yiming Chen, Junqi Huang, Ruofeng Tong, Min Tang, Peng Du

专题命中 软件智能体 :agentic(title);agent(abstract);分类 cs.AI

AI总结 本研究提出CADIR跨后端可编辑中间表示,结合几何签名匹配与构造图检索,解决现有CAD生成方法的跨系统编辑与建模历史保留问题,实验验证其几何保真度与可靠性优于现有方案。

详情

展开后加载摘要…

URL PDF HTML 收藏