SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
SWE-EVO:在长周期软件演化场景中基准测试编码智能体
Tue Le, Minh V. T. Thai, Dung Nguyen Manh, Huy Phan Nhat, Nghi D. Q. Bui
机构
*
FPT Software AI Center(FPT软件人工智能中心)
;
School of Computing and Information Systems(计算与信息系统学院)
;
University of Melbourne(墨尔本大学)
;
Center of AI Research(人工智能研究中心)
;
VinUniversity(文大学)
机构
*
Gradient
;
Soochow University(苏州大学)
;
Independent Researcher(独立研究者)
;
University of Southern California(南加州大学)
;
Rice University(Rice大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
University of the Chinese Academy of Sciences(中国科学院大学)
;
University of California, Los Angeles(加州大学洛杉矶分校)
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
GT-HarmBench:通过博弈论视角评估AI安全风险
Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang, Zhijing Jin
机构
*
ETH Zürich(苏黎世联邦理工学院)
;
Berea College(贝雷学院)
;
University of Toronto(多伦多大学)
;
Vector Institute(向量研究所)
;
Max Planck Institute for Intelligent Systems, Tübingen, Germany(图宾根德国智能系统马克斯·普朗克研究所)
机构
*
Hong Kong Polytechnic University(香港理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai AI Lab(上海人工智能实验室)
;
National University of Singapore(新加坡国立大学)
CommentsWe substantially updated the previous version "Diffusion Models for Tabular Data: Challenges, Current Progress, and Future Directions" by including flow matching models for tabular data
机构
*
Department of Computer Science, National University of Singapore(新加坡国立大学计算机科学系)
;
Singapore-MIT Alliance for Research and Technology Centre(新加坡-麻省理工联合研究中心)
;
CSAIL, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)
;
Worcester Polytechnic Institute(沃斯堡理工学院)
CommentsThis paper introduces SafeBarrier, a framework that enforces safety in large language models by steering their latent representations with control barrier functions during inference, reducing adversarial and unsafe outputs