arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

代码大模型 / AI 编程

代码生成、软件工程智能体、程序修复、测试生成和开发者工具。

共收录 1722 信号源:cs.SE, cs.CL, cs.AI, cs.LG, cs.PL

1. 代码评测 1722 篇

2107.06217 2021-07-14 cs.LG cs.AI 62%

What classifiers know what they don't?

Mohamed Ishmael Belghazi, David Lopez-Paz

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

Comments 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.00456 2021-06-02 stat.ME cs.AI cs.CR cs.LG 62%

Federated Estimation of Causal Effects from Observational Data

Thanh Vinh Vo, Trong Nghia Hoang, Young Lee, Tze-Yun Leong

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08755 2021-06-01 cs.CL cs.AI cs.IR 62%

DCH-2: A Parallel Customer-Helpdesk Dialogue Corpus with Distributions of Annotators' Labels

Zhaohao Zeng, Tetsuya Sakai

专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.16490 2021-03-31 cs.LG cs.AI 62%

Human Activity Analysis and Recognition from Smartphones using Machine Learning Techniques

Jakaria Rabbi, Md. Tahmid Hasan Fuad, Md. Abdul Awal

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

Comments Submitted to the 10th International Conference on Informatics, Electronics & Vision (ICIEV), 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.08785 2021-03-10 cs.NE cs.AI cs.LG 62%

Unsupervised Anomaly Detection in Stream Data with Online Evolving Spiking Neural Networks

Piotr S. Maciąg, Marzena Kryszkiewicz, Robert Bembenik, Jesus L. Lobo, Javier Del Ser

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

Comments 52 pages

Journal ref Neural Networks, Volume 139, 2021, Pages 118-139

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.01341 2021-02-03 cs.LG cs.AI cs.AR 62%

Benchmarking Quantized Neural Networks on FPGAs with FINN

Quentin Ducasse, Pascal Cotret, Loïc Lagadec, Robert Stewart

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

Comments Presented at DATE Friday Workshop on System-level Design Methods for Deep Learning on Heterogeneous Architectures (SLOHA 2021) (arXiv:2102.00818)

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02609 2020-12-09 cs.CL cs.LG 62%

Relevance Transformer: Generating Concise Code Snippets with Relevance Feedback

Carlos Gemmell, Federico Rossetto, Jeffrey Dalton

专题命中 代码评测 :code generation(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.06603 2020-09-23 q-bio.QM cs.CL cs.LG 62%

Transformer-CNN: Fast and Reliable tool for QSAR

Pavel Karpov, Guillaume Godin, Igor V. Tetko

专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.01075 2020-08-25 stat.ML cs.AI cs.LG 62%

Learning Neural Causal Models from Unknown Interventions

Nan Rosemary Ke, Olexa Bilaniuk, Anirudh Goyal, Stefan Bauer, Hugo Larochelle, Bernhard Schölkopf, Michael C. Mozer, Chris Pal, Yoshua Bengio

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.04766 2020-06-09 cs.AI cs.LG 62%

A Heuristically Self-Organised Linguistic Attribute Deep Learning in Edge Computing For IoT Intelligence

Hongmei He, Zhenhuan Zhu

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.01345 2020-03-04 cs.CL cs.LG 62%

Benchmark Performance of Machine And Deep Learning Based Methodologies for Urdu Text Document Classification

Muhammad Nabeel Asim, Muhammad Usman Ghani, Muhammad Ali Ibrahim, Sheraz Ahmad, Waqar Mahmood, Andreas Dengel

专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.11558 2019-11-27 cs.LG cs.CL stat.ML 62%

FairyTED: A Fair Rating Predictor for TED Talk Data

Rupam Acharyya, Shouman Das, Ankani Chattoraj, Md. Iftekhar Tanveer

专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.LG

Comments 9 pages, 4 figures, 3 tables. Accepted as a conference paper to be presented at AAAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.12271 2019-09-27 cs.RO cs.AI cs.CV cs.LG 62%

RLBench: The Robot Learning Benchmark & Learning Environment

Stephen James, Zicong Ma, David Rovick Arrojo, Andrew J. Davison

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

Comments Videos and code: https://sites.google.com/view/rlbench

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.08250 2019-09-19 cs.AI cs.CL 62%

Natural Language Generation for Non-Expert Users

Van Duc Nguyen, Tran Cao Son, Enrico Pontelli

专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI

Comments In Proceedings ICLP 2019, arXiv:1909.07646

Journal ref EPTCS 306, 2019, pp. 280-294

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.07805 2019-07-12 cs.LG cs.AI stat.ML 62%

Constrained Policy Improvement for Safe and Efficient Reinforcement Learning

Elad Sarafian, Aviv Tamar, Sarit Kraus

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.06312 2019-04-15 cs.LG cs.AI stat.ML 62%

Let's Play Again: Variability of Deep Reinforcement Learning Agents in Atari Environments

Kaleigh Clary, Emma Tosch, John Foley, David Jensen

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2018 Critiquing and Correcting Trends Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.02312 2018-07-30 cs.SE cs.CV cs.LG 62%

Machine Learning-Based Prototyping of Graphical User Interfaces for Mobile Apps

Kevin Moran, Carlos Bernal-Cárdenas, Michael Curcio, Richard Bonett, Denys Poshyvanyk

专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.LG

Comments Accepted to IEEE Transactions on Software Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.06378 2018-01-22 stat.ML cs.LG cs.SE 62%

Introducing ReQuEST: an Open Platform for Reproducible and Quality-Efficient Systems-ML Tournaments

Thierry Moreau, Anton Lokhmotov, Grigori Fursin

专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.LG

Comments ReQuEST tournament website: http://cKnowledge.org/request

详情

展开后加载摘要…

URL PDF HTML 收藏
1409.7165 2016-11-15 cs.LG cs.IR cs.SE 62%

Heterogeneous Metric Learning with Content-based Regularization for Software Artifact Retrieval

Liang Wu, Hui Xiong, Liang Du, Bo Liu, Guandong Xu, Yong Ge, Yanjie Fu, Yuanchun Zhou, Jianhui Li

专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.LG

Comments to appear in IEEE International Conference on Data Mining (ICDM), Shen Zhen, China, December 2014

详情

展开后加载摘要…

URL PDF HTML 收藏
1506.02465 2016-04-07 cs.AI cs.LG 62%

ASlib: A Benchmark Library for Algorithm Selection

Bernd Bischl, Pascal Kerschke, Lars Kotthoff, Marius Lindauer, Yuri Malitsky, Alexandre Frechette, Holger Hoos, Frank Hutter, Kevin Leyton-Brown, Kevin Tierney, Joaquin Vanschoren

专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG

Comments Accepted to be published in Artificial Intelligence Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
1506.06256 2015-06-23 cs.SE cs.LG cs.PF 62%

Collective Mind, Part II: Towards Performance- and Cost-Aware Software Engineering as a Natural Science

Grigori Fursin, Abdul Memon, Christophe Guillon, Anton Lokhmotov

专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.LG

Comments Presented at the 18th International Workshop on Compilers for Parallel Computing (CPC'15), London, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06413 2026-07-08 stat.ME cs.AI 新提交 61%

An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery

一种评估智能体人工智能自主模型发现的实验设计方法

Hao He, Xueying Liu, Chris J. Kuhlman, Xinwei Deng

机构 * Department of Statistics, Virginia Tech(统计学系,弗吉尼亚理工学院) Department of Statistical Science, Baylor University(统计科学系,贝勒大学) Advanced Research Computing, Virginia Tech(高级研究计算,弗吉尼亚理工学院)

专题命中 代码评测 :coding agent(abstract);分类 cs.AI;repository(comments)

AI总结 研究大型语言模型编码智能体自主模型发现行为,提出实验设计与分析框架,将智能体视为随机模型发现算子,在多种受控因素下研究Codex和Claude Code两个算子,进行回归模型和推理,开发规范分解,通过网络造词游戏测试平台得出相关深刻发现。

Comments 39 pages, 11 figures, 6 tables. Data and code available at the GitHub repository listed in the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.06085 2023-09-20 cs.CL 61%

BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models

Wei Qi Leong, Jian Gang Ngui, Yosephine Susanto, Hamsawardhini Rengarajan, Kengatharaiyer Sarveswaran, William Chandra Tjhi

专题命中 代码评测 :repository(abstract,comments);分类 cs.CL

Comments 86 pages, 7 figures, added link to repository in abstract, minor formatting changes and typo corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13928 2026-08-17 cs.CR cs.SE 新提交 57%

CoSA: Context-Aware Severity Assessment via Context Analysis with Large Language Models

CoSA:基于大语言模型的上下文分析实现上下文感知的漏洞严重程度评估

Jinfeng Jiang, Yikun Li, Chengran Yang, Ting Zhang, Wen Bin Leow, Yide Yin, Eng Lieh Ouh, Lwin Khin Shar, David Lo

专题命中 代码评测 :repository(abstract);分类 cs.SE

AI总结 CoSA是一种基于大语言模型的上下文感知漏洞严重程度评估方法,通过两阶段仓库剪枝策略与Transformer预测器,在6816个CVSS标注实例上较最优基线提升了14.4%准确率与15.3% Macro-F1

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13077 2026-08-14 cs.SE 新提交 57%

How Powerful are LLMs in Generating Formal Program Specifications?

大型语言模型在生成形式化程序规约方面的能力有多强?

Fanpeng Yang, Xing Li, Shuling Wang, Jie An, Zeyu Sun, Shenghua Feng, Wenhan Wang, Weiyi Wang, Naijun Zhan, Fanjiang Xu

专题命中 代码评测 :code generation(abstract);分类 cs.SE

AI总结 该研究引入基于Rocq的Coins评估框架,在HumanEval数据集上开展大规模研究,发现LLMs生成形式化程序规约仍具挑战,准确的规约评估是理解其能力的核心。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12805 2026-08-14 cs.LG 新提交 57%

CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

CoMedBench:合成医疗数据保真度与下游效用的多源基准

Akanta Das, Al Amin Farhad, Mrinmoy Sarkar Anto, David Rehkopf, Ayin Vala, Tanmoy Sarkar Pias

专题命中 代码评测 :repository(abstract);分类 cs.LG

AI总结 CoMedBench是涵盖多源合成医疗数据的可复现基准,评估多生成器在静态表格和时序ICU任务上的保真度与下游效用,实验显示合成数据可保留多数下游信号,不同生成器表现存在差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12262 2026-08-13 cs.CV cs.AI 新提交 57%

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

Diagram-MMU:面向科学图表的多模态基准测试

Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu, Yongke Yao, Jinhao Du, Wei He, Kai Zou, Zechao Li, Jingdong Wang

机构 * Nanjing University of Science and Technology(南京理工大学) Baidu Inc(百度公司) AIML, Adelaide University(阿德莱德大学AIML) SUTD(新加坡科技设计大学) Southeast University(东南大学) East China Normal University(华东师范大学) University of Oxford(牛津大学)

专题命中 代码评测 :code generation(abstract);分类 cs.AI

AI总结 本文构建了多模态基准测试Diagram-MMU,评估MLLMs的科学图表解析与理解能力,发现图表转代码任务更具挑战,Claude-4.6 Opus在智能体场景下表现最优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11216 2026-08-13 cs.AI 新提交 57%

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

AutoWorldModel-Bench:面向自动化世界模型研究的以状态为中心的基准

Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri

机构 * Electronic Arts(美国艺电公司) Simon Fraser University(西蒙菲莎大学)

专题命中 代码评测 :coding agent(abstract);分类 cs.AI

AI总结 AutoWorldModel-Bench是面向AI编码智能体的闭环基准,涵盖8个游戏环境,采用结构化状态表征,64次会话中多数智能体通过研究式修改改进了世界模型,可评估智能体的开放式研究能力。

Comments Project page: https://electronicarts.github.io/AutoWorldModelBench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10145 2026-08-12 cs.LG 新提交 57%

The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom

评估协议决定结果:在TwoRoom上独立复现LeWorldModel

Joyjeet Singh

专题命中 代码评测 :repository(abstract);分类 cs.LG

AI总结 本研究独立复现LeWorldModel在TwoRoom上的结果,发现评估协议(如目标偏移量)会显著影响性能,还揭示单步预测准确率无法预测长程规划成功、批量归一化层会夸大验证损失等关键结论。

Comments Independent reproduction of arXiv:2603.19312 - https://github.com/joyjeet-singh/tinylab

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15237 2026-08-12 stat.ME cs.LG stat.ML 交叉投稿 57%

Optimized Sequential Testing for Binary Ensemble Classifiers

二元集成分类器的优化序贯测试

Joseph Kalman, Amit Moscovich

专题命中 代码评测 :repository(abstract);分类 cs.LG

AI总结 提出一种序贯测试方法,通过提前停止基模型评估来降低二元集成分类器的计算成本,同时控制与完整集成的不一致率,并利用线性规划求解最优停止策略。

Comments 33 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏