Confident and Wrong: Silent Semantic Failures in Coding Agents
自信且错误:编码智能体中的无声语义失败
机构 * Snowflake AI Research(Snowflake AI研究院)
AI总结 本文揭示编码智能体在任务完成率与可信度之间存在系统性偏差,提出“无声语义失败”这一危险模式,并建议通过多次运行的测试验证正确率来评估智能体。
Comments 8 pages, 8 figures. Request: please change the primary category to cs.AI