arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

代码大模型 / AI 编程

代码生成、软件工程智能体、程序修复、测试生成和开发者工具。

2026-04-06 至 2026-04-06 共收录 6 信号源:cs.SE, cs.CL, cs.AI, cs.LG, cs.PL

1. 代码评测 6 篇

2604.02729 2026-04-06 cs.SE cs.AI cs.CL 82%

IndustryCode: A Benchmark for Industry Code Generation

IndustryCode: 一个用于行业代码生成的基准

Puyu Zeng, Zhaoxi Wang, Zhixu Duan, Liang Feng, Shaobo Wang, Cunxiang Wang, Jinghang Wang, Bing Zhao, Hu Wei, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Alibaba Group(阿里巴巴集团) Tsinghua University(清华大学)

专题命中 代码评测 :code generation(title,abstract);分类 cs.SE、cs.CL、cs.AI

AI总结 本文提出IndustryCode基准,涵盖125个工业挑战和579个子问题,支持多种编程语言,评估大模型在工业场景中的代码生成能力。

Comments 37 pages, 28 figures, 4 tables. Includes appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02648 2026-04-06 cs.SE cs.AI 62%

GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers

GBQA:用于评估LLM作为质量保证工程师的博弈基准

Shufan Jiang, Chios Chen, Zhiyang Chen

机构 * The University of Hong Kong(香港大学) Independent Researcher(独立研究者) Westlake University(西湖大学) Datawhale Org(Datawhale组织)

专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.AI

AI总结 本文提出GBQA基准,通过30款游戏和124个经人类验证的bug测试LLM自主发现软件缺陷的能力,实验表明最佳模型仅能识别48.39%的bug。

Comments Accepted as a workshop paper at the Fourteenth International Conference on Learning Representations (ICLR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02382 2026-04-06 cs.SE cs.AI 62%

Ambig-IaC: Multi-level Disambiguation for Interactive Cloud Infrastructure-as-Code Synthesis

Ambig-IaC:交互式云基础设施即代码合成的多级消歧

Zhenning Yang, Kaden Gruizenga, Tongyuan Miao, Patrick Tser Jern Kon, Hui Guan, Ang Chen

机构 * University of Michigan(密歇根大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.AI

AI总结 本文提出Ambig-IaC框架,通过多级消歧方法提升IaC配置生成的准确性,针对IaC配置的不可逆性,设计了基于结构分歧的评估框架,在结构和属性评估上分别提升18.4%和25.4%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01021 2026-04-06 cs.LG cs.AI 62%

Transfer learning for nonparametric Bayesian networks

非参数贝叶斯网络的迁移学习

Rafael Sojo, Pedro Larrañaga, Concha Bielza

机构 * Aingura IIoT Universidad Politécnica de Madrid, Departamento de Inteligencia Artificial(马德里理工大学人工智能系)

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

AI总结 本文提出两种非参数贝叶斯网络迁移学习方法,PC-stable-transfer学习和hill climbing transfer学习,通过约束和评分方法提升模型性能,同时引入特定指标解决负迁移问题,最终通过统计检验证明方法有效性。

Comments An earlier version was previously posted on SSRN. This version includes improvements in experiments and evaluation metrics following reviewer comments. Revision submitted to Knowledge-Based Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02520 2026-04-06 physics.data-an cs.LG 57%

Neural posterior estimation for scalable and accurate inverse parameter inference in Li-ion batteries

用于锂离子电池可扩展和准确逆参数推断的神经后验估计

Malik Hassanaly, Corey R. Randall, Peter J. Weddle, Paul J. Gasper, Conlain Kelly, Tanvir R. Tanim, Kandler Smith

机构 * Computational Science Center, National Laboratory of the Rockies(落基山国家实验室计算科学中心) Energy Conversion and Storage Systems Center, National Laboratory of the Rockies(落基山国家实验室能量转换与存储系统中心) Energy Storage Research and Analysis Department, Idaho National Laboratory(爱达荷国家实验室储能研究与分析部)

专题命中 代码评测 :repository(abstract);分类 cs.LG

AI总结 本文提出神经后验估计方法,用于高效估计锂离子电池参数,尽管在高维情况下计算成本较高,但能提供更准确的参数校准,并在电压预测中存在误差,同时具有更高的可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07743 2026-04-06 eess.SY cs.SY 50%

The Reliability of Remotely Piloted Aircraft System Performance under Aeronautical Communication Uncertainties

远程飞行器系统性能在航空通信不确定性下的可靠性

Yutian Pang, Andrew Paul Kendall, John-Paul Clarke

专题命中 代码评测 :repository(abstract)

AI总结 研究在随机通信条件下,高机动远程飞行器系统任务完成性能的可靠性,通过蒙特卡洛模拟评估通信延迟和信号丢失对系统稳定性的影响,并提出新的可靠性指标communicability。

详情

展开后加载摘要…

URL PDF HTML 收藏