What classifiers know what they don't?
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
Comments 27 pages
AI 大模型
代码生成、软件工程智能体、程序修复、测试生成和开发者工具。
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
Comments 27 pages
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
Comments Preprint
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI
Comments 6 pages, 3 figures
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
Comments Submitted to the 10th International Conference on Informatics, Electronics & Vision (ICIEV), 2021
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
Comments 52 pages
Journal ref Neural Networks, Volume 139, 2021, Pages 118-139
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
Comments Presented at DATE Friday Workshop on System-level Design Methods for Deep Learning on Heterogeneous Architectures (SLOHA 2021) (arXiv:2102.00818)
专题命中 代码评测 :code generation(abstract);分类 cs.CL、cs.LG
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.LG
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.LG
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.LG
Comments 9 pages, 4 figures, 3 tables. Accepted as a conference paper to be presented at AAAI 2020
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
Comments Videos and code: https://sites.google.com/view/rlbench
专题命中 代码评测 :repository(abstract);分类 cs.CL、cs.AI
Comments In Proceedings ICLP 2019, arXiv:1909.07646
Journal ref EPTCS 306, 2019, pp. 280-294
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
Comments NeurIPS 2018 Critiquing and Correcting Trends Workshop
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.LG
Comments Accepted to IEEE Transactions on Software Engineering
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.LG
Comments ReQuEST tournament website: http://cKnowledge.org/request
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.LG
Comments to appear in IEEE International Conference on Data Mining (ICDM), Shen Zhen, China, December 2014
专题命中 代码评测 :repository(abstract);分类 cs.AI、cs.LG
Comments Accepted to be published in Artificial Intelligence Journal
专题命中 代码评测 :repository(abstract);分类 cs.SE、cs.LG
Comments Presented at the 18th International Workshop on Compilers for Parallel Computing (CPC'15), London, UK
一种评估智能体人工智能自主模型发现的实验设计方法
机构 * Department of Statistics, Virginia Tech(统计学系,弗吉尼亚理工学院) ; Department of Statistical Science, Baylor University(统计科学系,贝勒大学) ; Advanced Research Computing, Virginia Tech(高级研究计算,弗吉尼亚理工学院)
专题命中 代码评测 :coding agent(abstract);分类 cs.AI;repository(comments)
AI总结 研究大型语言模型编码智能体自主模型发现行为,提出实验设计与分析框架,将智能体视为随机模型发现算子,在多种受控因素下研究Codex和Claude Code两个算子,进行回归模型和推理,开发规范分解,通过网络造词游戏测试平台得出相关深刻发现。
Comments 39 pages, 11 figures, 6 tables. Data and code available at the GitHub repository listed in the paper
专题命中 代码评测 :repository(abstract,comments);分类 cs.CL
Comments 86 pages, 7 figures, added link to repository in abstract, minor formatting changes and typo corrections
CoSA:基于大语言模型的上下文分析实现上下文感知的漏洞严重程度评估
专题命中 代码评测 :repository(abstract);分类 cs.SE
AI总结 CoSA是一种基于大语言模型的上下文感知漏洞严重程度评估方法,通过两阶段仓库剪枝策略与Transformer预测器,在6816个CVSS标注实例上较最优基线提升了14.4%准确率与15.3% Macro-F1
大型语言模型在生成形式化程序规约方面的能力有多强?
专题命中 代码评测 :code generation(abstract);分类 cs.SE
AI总结 该研究引入基于Rocq的Coins评估框架,在HumanEval数据集上开展大规模研究,发现LLMs生成形式化程序规约仍具挑战,准确的规约评估是理解其能力的核心。
CoMedBench:合成医疗数据保真度与下游效用的多源基准
专题命中 代码评测 :repository(abstract);分类 cs.LG
AI总结 CoMedBench是涵盖多源合成医疗数据的可复现基准,评估多生成器在静态表格和时序ICU任务上的保真度与下游效用,实验显示合成数据可保留多数下游信号,不同生成器表现存在差异。
Diagram-MMU:面向科学图表的多模态基准测试
机构 * Nanjing University of Science and Technology(南京理工大学) ; Baidu Inc(百度公司) ; AIML, Adelaide University(阿德莱德大学AIML) ; SUTD(新加坡科技设计大学) ; Southeast University(东南大学) ; East China Normal University(华东师范大学) ; University of Oxford(牛津大学)
专题命中 代码评测 :code generation(abstract);分类 cs.AI
AI总结 本文构建了多模态基准测试Diagram-MMU,评估MLLMs的科学图表解析与理解能力,发现图表转代码任务更具挑战,Claude-4.6 Opus在智能体场景下表现最优。
AutoWorldModel-Bench:面向自动化世界模型研究的以状态为中心的基准
机构 * Electronic Arts(美国艺电公司) ; Simon Fraser University(西蒙菲莎大学)
专题命中 代码评测 :coding agent(abstract);分类 cs.AI
AI总结 AutoWorldModel-Bench是面向AI编码智能体的闭环基准,涵盖8个游戏环境,采用结构化状态表征,64次会话中多数智能体通过研究式修改改进了世界模型,可评估智能体的开放式研究能力。
Comments Project page: https://electronicarts.github.io/AutoWorldModelBench/
评估协议决定结果:在TwoRoom上独立复现LeWorldModel
专题命中 代码评测 :repository(abstract);分类 cs.LG
AI总结 本研究独立复现LeWorldModel在TwoRoom上的结果,发现评估协议(如目标偏移量)会显著影响性能,还揭示单步预测准确率无法预测长程规划成功、批量归一化层会夸大验证损失等关键结论。
Comments Independent reproduction of arXiv:2603.19312 - https://github.com/joyjeet-singh/tinylab
二元集成分类器的优化序贯测试
专题命中 代码评测 :repository(abstract);分类 cs.LG
AI总结 提出一种序贯测试方法,通过提前停止基模型评估来降低二元集成分类器的计算成本,同时控制与完整集成的不一致率,并利用线性规划求解最优停止策略。
Comments 33 pages, 5 figures