arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

代码大模型 / AI 编程

代码生成、软件工程智能体、程序修复、测试生成和开发者工具。

2026-04-30 至 2026-04-30 共收录 6 信号源:cs.SE, cs.CL, cs.AI, cs.LG, cs.PL

1. 代码评测 6 篇

2604.26275 2026-04-30 cs.SE 70%

Agentic AI in the Software Development Lifecycle: Architecture, Empirical Evidence, and the Reshaping of Software Engineering

软件开发生命周期中的代理AI:架构、实证证据与软件工程的重塑

Happy Bhati

专题命中 代码评测 :code generation(abstract);repository(abstract);分类 cs.SE

AI总结 本文探讨代理AI在软件开发生命周期中的应用,提出六层参考架构,分析代理SDLC与传统SDLC的差异,并通过实证数据展示其在性能、生产力和劳动力市场的影响。

Comments 9 pages, 5 figures, 2 tables, survey paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26243 2026-04-30 cs.CL cs.AI 62%

StratMem-Bench: Evaluating Strategic Memory Use in Virtual Character Conversation Beyond Factual Recall

StratMem-Bench:评估虚拟角色对话中的战略记忆使用超越事实回忆

Yerong Wu, Tianxing Wu, Minghao Zhu, Hangyu Sha, Haofen Wang

机构 * School of Computer Science and Engineering, Southeast University, China(东南大学计算机科学与工程学院,中国) Key Laboratory of New Generation Artificial Intelligence Technology and its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用国家重点实验室(东南大学),中华人民共和国教育部,中国) College of Design and Innovation, Tongji University, China(同济大学设计与创新学院,中国)

专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI

AI总结 StratMem-Bench旨在评估虚拟角色对话中战略记忆的使用,通过异构记忆池测试角色在不同记忆类型间的决策能力,提出多种评估指标以衡量记忆整合与利用效果。

Comments 20 pages, accepted by ACL 2026 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25982 2026-04-30 cs.LG cs.AI cs.CY cs.ET 62%

Open Problems in Frontier AI Risk Management

前沿人工智能风险管理中的开放问题

Marta Ziosi, Miro Plueckebaum, Stephen Casper, Henry Papadatos, Ze Shen Chin, Peter Slattery, James Gealy, Tim G. J. Rudner, Brian Tse, Ariel Gil, Patricia Paskov, Maximilian Negele, Rokas Gipiškis, Nada Madkour, Vera Lummis, Rupal Jain, Luise Eder, Kristina Fort, Malou C. van Draanen Glismann, Inès Belhadj, Amin Oueslati, Anna K. Wisakanto, Richard Mallah, Koen Holtman, Ranj Zuhdi, Daniel S. Schiff, Jessica Newman, Malcolm Murray, Robert Trager

机构 * Oxford Martin AI Governance Initiative, University of Oxford(牛津大学人工智能治理倡议) MIT Computer Science and Artificial Intelligence Laboratory, MIT(麻省理工学院计算机科学与人工智能实验室) MIT Future Tech(麻省理工学院未来技术) Stanford University(斯坦福大学) Governance and Responsible AI Lab, Purdue University(普渡大学治理与负责任的人工智能实验室) University of Toronto(多伦多大学) Mercatus Center, George Mason University(乔治·马歇尔大学麦卡锡中心) Vilnius University(维尔纽斯大学) Vijil SaferAI AI Standards Lab(人工智能标准实验室) The Future Society(未来社会) Concordia AI(康科德人工智能) Pivotal Research Center for AI Risk Management & Alignment(人工智能风险管理和对齐中心) UC Berkeley Center for Long-Term Cybersecurity(伯克利大学长期网络安全中心) Independent(独立)

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨前沿人工智能风险管理中的核心问题,通过文献综述识别未解决的挑战,并分类问题类型以指导未来研究与治理。

Comments 81 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21444 2026-04-30 cs.CV 50%

Benchmarking Deep Learning and Vision Foundation Models for Atypical vs. Normal Mitosis Classification with Cross-Dataset Evaluation

基于跨数据集评估的深度学习与视觉基础模型在异常有丝分裂分类中的基准测试

Sweta Banerjee, Viktoria Weiss, Taryn A. Donovan, Rutger H. J. Fick, Thomas Conrad, Jonas Ammeling, Nils Porsche, Robert Klopfleisch, Christopher Kaltenecker, Katharina Breininger, Marc Aubreville, Christof A. Bertram

机构 * Flensburg University of Applied Sciences(弗劳恩霍夫应用科技大学) University of Veterinary Medicine, Vienna(维也纳兽医大学) The Schwarzman Animal Medical Center(施瓦茨曼动物医疗中心) Diffusely(扩散地) Freie Universität Berlin(柏林自由大学) Technische Hochschule Ingolstadt(因戈尔施塔特技术大学) Medical University of Vienna(维也纳医学大学) Julius-Maximilians-Universität Würzburg(维尔茨堡约瑟夫-马克斯姆-大学)

专题命中 代码评测 :repository(abstract)

AI总结 本文通过对比深度学习方法,评估了自动识别异常有丝分裂图的性能,提出两种新数据集并展示在不同领域中的分类准确率。

Comments Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://melba-journal.org/2026:006

Journal ref Machine.Learning.for.Biomedical.Imaging. 2026 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26495 2026-04-30 cs.CR 50%

Beyond Code Reasoning: A Specification-Anchored Audit Framework for Expert-Augmented Security Verification

超越代码推理:一种基于规范的审计框架用于专家增强的安全验证

Masato Kamba, Hirotake Murakami, Akiyoshi Sannai

专题命中 代码评测 :repository(abstract)

AI总结 本文提出SPECA框架,通过自然语言规范生成类型化安全属性,实现基于规范的审计,解决代码层面无法检测规范依赖漏洞的问题,提升安全验证的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13421 2026-04-30 cs.SD eess.AS 50%

Explainable Detection of Machine Generated Music and Early Systematic Evaluation

可解释性地检测机器生成音乐及早期系统性评估

Yupei Li, Qiyang Sun, Hanqian Li, Lucia Specia, Björn W. Schuller

机构 * Imperial College London, Computing(帝国理工学院伦敦分校,计算机系) Shandong University(山东大学)

专题命中 代码评测 :repository(abstract)

AI总结 本文通过多种模型评估机器生成音乐检测,揭示ResNet18在领域内和领域外测试中的最佳性能,提出未来研究方向以提升检测方法的鲁棒性和有效性。

Comments Accepted at Scientific report

Journal ref Sci Rep 16, 13757 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏