机构
*
University of Notre Dame(圣母大学)
;
University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Nanjing University(南京大学)
Zheng-Xin Yong, Parv Mahajan, Andy Wang, Ida Caspary, Yernat Yestekov, Zora Che, Mosh Levy, Elle Najt, Dennis Murphy, Prashant Kulkarni, Lev McKinney, Kei Nishimura-Gasparian, Ram Potham, Aengus Lynch, Michael L. Chen
机构
*
Constellation
;
Anthropic Fellows Program(Anthropic研究员项目)
;
Brown University(布朗大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
Imperial College London(伦敦帝国学院)
;
University of Maryland, College Park(马里兰大学帕克分校)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Bar Ilan University(巴伊兰大学)
;
University of Toronto(多伦多大学)
;
University of Oxford(牛津大学)
专题命中
Agent评测
:agentic(abstract);分类 cs.AI、cs.CL
AI总结
Kimi K2.5作为开源大模型,在安全评估中显示出潜在风险,包括CBRNE滥用、网络安全漏洞及政治偏见,但其在拒绝恶意请求方面表现较弱,凸显开源模型的安全挑战。
Let's Have a Conversation: Designing and Evaluating LLM Agents for Interactive Optimization
让我们交谈:设计和评估用于交互优化的LLM代理
Joshua Drossman, Alexandre Jacquillat, Sébastien Martin
机构
*
Operations Research Center and Sloan School of Management, Massachusetts Institute of Technology(麻省理工学院运筹学研究中心和斯隆管理学院)
;
Kellogg School of Management, Northwestern University(西北大学凯洛格管理学院)