Large Language Models Illuminate a Progressive Pathway to Artificial Healthcare Assistant: A Review
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 24 pages, 1 figure, 3 tables
AI 大模型
代码生成、软件工程智能体、程序修复、测试生成和开发者工具。
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 24 pages, 1 figure, 3 tables
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.CL、cs.LG
Comments Project site with code and data: https://intercode-benchmark.github.io
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Published as workshop paper at Practical ML for Developing Countries Workshop @ ICLR 2020
专题命中 代码评测 :code generation(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted by ACL 2023 Findings. The first three authors contributed equally
专题命中 代码评测 :program synthesis(abstract);分类 cs.CL、cs.AI、cs.LG
Comments humaneval results, clarity
专题命中 代码评测 :code model(abstract);分类 cs.CL、cs.LG、cs.PL
Comments Accepted to ICML 2023, Code and data release: https://github.com/google-research/babelcode
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.AI、cs.PL
Comments In the proceedings of Advances in Neural Information Processing Systems, 2022
专题命中 代码评测 :code model(abstract);分类 cs.SE、cs.AI、cs.PL
Comments The 37th IEEE/ACM International Conference on Automated Software Engineering
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Title and relevant changes are made
专题命中 代码评测 :program repair(abstract);分类 cs.SE、cs.AI、cs.PL
Comments 3 pages, 3 tables, 1 GitHub url: https://github.com/pmorvalho/C-Pack-IPAs
专题命中 代码评测 :program synthesis(abstract);分类 cs.SE、cs.AI、cs.LG
Comments Accepted to the Technical Track of ICSE 2022
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted by AAAI2021
机构 * Microsoft(微软公司) ; Hong Kong Baptist University(香港 Baptist 大学)
专题命中 代码评测 :code generation(abstract,comments);分类 cs.CL、cs.AI
Comments Large Language model, Code Generation, Code LLMs.This paper has been accepted to ICLR 2024. Please cite the ICLR version
Journal ref The Twelfth International Conference on Learning Representations (ICLR 2024)
专题命中 代码评测 :program repair(abstract,comments);分类 cs.SE、cs.AI
Comments Accepted to the 6th International Workshop on Automated Program Repair (APR 2025)
专题命中 代码评测 :code model(abstract,comments);分类 cs.SE、cs.LG
Comments Accepted to IEEE Transactions on Software Engineering. Extension of our previous paper "What do pre-trained code models know about code?" (ASE 2021, arXiv:2108.11308). 21 pages
H-VAEP与H-xT:通过概率估计评估手球进攻中持球动作的价值
机构 * Paderborn University(帕德博恩大学) ; SG Flensburg-Handewitt(弗伦斯堡-汉德维特体育俱乐部)
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
AI总结 本文将足球的xT与VAEP框架适配至手球,开发H-xT与H-VAEP模型,利用手球德甲数据验证其有效性并发布代码,实现了手球球员的合理评估。
Comments 13 pages, 6 figures, 1 table. Accepted at the 13th Workshop on Machine Learning and Data Mining for Sports Analytics (MLSA 2026), co-located with ECML PKDD 2026
Aftab:并行Q网络中CNN编码器与高级价值函数的综合基准
机构 * University of Padua(帕多瓦大学)
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
AI总结 本研究针对并行Q网络,设计评估8种CNN拓扑并集成多种Q学习扩展,提出复合架构Aftab,在Atari-57与Procgen Hard基准上均优于基线,已开源。
迈向具身智能体的管控
专题命中 代码评测 :coding agent(abstract);分类 cs.AI、cs.LG
AI总结 本文提出名为Thea的具身智能体管控系统,通过场景图和退出码评估弥合物理世界与智能体的差距,实现长程任务完成。
Comments Project page: https://eit-hai.github.io/thea
利用自主LLM研究循环优化专家设计的晶体图网络用于带隙预测
机构 * Department of Materials Science and NanoEngineering(材料科学与纳米工程系)
专题命中 代码评测 :coding agent(abstract);分类 cs.AI、cs.LG
AI总结 提出一个自主LLM研究循环,在MatBench带隙基准上构建了无需外部预训练的最准确模型,超越了所有17个专家设计模型,通过实现元素对特征和空间群嵌入等已知方法。
基于大语言模型的因果智能体
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI
AI总结 该研究针对LLM难以处理因果问题的挑战,提出Causal Agent框架,构建CausalTQA基准,实验显示其在多层级因果问题及真实数据集上性能优于SOTA。
放射学视觉语言模型基准的法证可重复性审计:从预期协议到发布工件
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI
AI总结 对胸部X光视觉语言模型试点进行法证可重复性审计,追踪提示绑定等多方面情况,发现存在图像渲染、数据分割等问题,重建队列改变统计值,撤回原声明并指定机器可验证控制。
Comments Withdrawn by the author. On further review, the archived artifacts underlying this audit are too incomplete to support the reported statistics, and the paper's conclusions do not follow from the available evidence. The work is withdrawn in full; earlier versions should not be cited
InfiniteScienceGym:一个无界的、程序生成的科学分析基准
机构 * Kahlert School of Computing(卡勒特计算学院)
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI
AI总结 本文提出InfiniteScienceGym,通过程序生成科学仓库和可验证问答任务,评估证据推理、回避和工具分析能力,发现现有模型在回答不可答问题上存在显著缺陷。
Comments 31 pages, 5 figures, 8 tables. Accepted to COLM 2026. See https://infinitesciencegym.github.io/
AI评估应衡量验证成本,而非仅正确性
机构 * Hitachi, Ltd.(日立制作所) ; Hitachi Rail(日立铁路)
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.AI
AI总结 该研究指出AI评估仅靠正确性不足,提出需纳入验证成本,定义了验证成本错误,通过代码生成等证据说明高基准准确率可能掩盖高验证成本,主张评估应考虑现实资源约束下的错误可检测性。
Comments 20 pages, 1 table
ATLAS: 大规模软件生态系统的智能体分类体系
机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; Nanyang Technological University(南洋理工大学) ; Nankai University(南开大学) ; Sinosoft Company Limited(中软国际有限公司)
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.CL
AI总结 提出ATLAS框架,结合LLM全局知识与实际仓库分布,通过智能体协作和自修正循环自动构建软件仓库的层次化分类体系,在54,387个GitHub仓库上评估,分类质量F1达83.13%,优于基线15个百分点。
Comments Accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)
超越问题:评估大型语言模型(实际)知道什么
机构 * ScaDS.AI Dresden/Leipzig & TU Dresden, Germany(ScaDS.AI 德尔布兰德/莱比锡及德累斯顿技术大学,德国)
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI
AI总结 提出开放知识评估新范式,通过开放式提示(如“告诉我关于M.L. King的一切”)评估模型自然表达的知识,并构建BeQu基准测试10,000个实体。
AgentMeter: 评估基于CLI的本地任务求解智能体的模型-CLI匹配
机构 * Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) ; Hangzhou Institute for Advanced Study, Chinese Academy of Sciences(中国科学院杭州高等研究院)
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.AI
AI总结 提出AgentMeter基准,通过成功锚定、成本感知的AMS指标评估模型与CLI的匹配,发现模型和CLI选择不可解耦,需作为部署单元评估。
Comments 13 pages, 4 figures, 12 tables; includes supplementary material
WebCoderBench: 用全面且可解释的评估指标对Web应用生成进行基准测试
专题命中 代码评测 :code generation(abstract);分类 cs.SE、cs.AI
AI总结 本文提出WebCoderBench,首个基于真实用户需求的Web应用生成基准,包含1572个真实用户需求及24个细粒度评估指标,结合规则和LLM评估方法,提供可解释的评估结果,实验显示无单一最优模型,为模型优化提供方向。
谁来评估评估者?用于自我改进的语言模型代理的协同进化评估指标和技能
机构 * Amazon(亚马逊)
专题命中 代码评测 :code generation(abstract);分类 cs.CL、cs.AI
AI总结 研究自我进化智能体系统中评估指标缺失问题,提出指标可进化,通过“双棘轮”协同进化指标与技能循环,在多任务中保留提升效果,还阐述安全性源于锚定规则与外部审核,为无可靠自动验证器场景提供架构。
Comments Code: https://github.com/amazon-science/Self-Evolving-Agents-Double-Ratchet
图作为验证器:用于过程间漏洞检测的智能体强化学习
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.AI
AI总结 针对现有漏洞检测孤立处理函数的问题,提出基于CPG的智能体强化学习框架VulAgentRL,通过图验证设计可靠奖励并预热初始化策略,在仓库级划分下的严格指标及多场景中均优于基线。
在何处进行干预?对差分隐私合成表格数据上的公平感知学习进行基准测试
机构 * ÉTS Montréal(蒙特利尔高等商学院) ; Inria Grenoble(格勒诺布尔计算机科学及自动化研究所)
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
AI总结 研究在差分隐私合成表格数据上公平干预的效果,以自适应迭代机制为基准,在多数据集、指标及策略下评估,比较四种管道配置,发现仅DP会降效,公平干预可部分恢复公平,后处理方法权衡更优,还开源相关内容以支持研究。
Comments Paper accepted at PETS 2026. Code is available at https://github.com/vinicius-verona/dp-fair-intervention-benchmark