Large-Scale ChatBot Validation Through Customer Digital Twin Simulations
基于客户数字孪生模拟的大规模聊天机器人验证
Cristovao Iglesias, Devesh Batra, Alankar Atreya, Stefan Wagner, Robert Hankache, Patrick Sinclair, Giulio Pelosio, Michael McMillan, Greig A. Cowan, Raad Khraishi
GPT-Red: Automated Red Teaming via Self-Play at Scale
GPT-Red:基于大规模自博弈的自动化红队测试
Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cerón Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen
Conformal Changepoint Localization and Root Cause Analysis with Corrupted Observations
含损坏观测值的保序变点定位与根因分析
Seunghun Yu, Meiyi Zhu, Petar Popovski, Joonhyuk Kang, Osvaldo Simeone
机构
*
Korea Advanced Institute of Science and Technology(韩国科学技术院)
;
King’s College London(伦敦国王学院)
;
Aalborg University(奥尔堡大学)
;
Northeastern University London(伦敦东北大学)
CommentsPeer-reviewed and presented at the 1st Workshop on Toward Trustworthy Vision-Language Models in the Wild (TrustVLM), co-located with ACM ICMR 2026, Amsterdam. Non-archival workshop. Reviews public on OpenReview. 5 pages, 2 figures
Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance
重新思考胸部X射线机器学习中的临床相关性:评估参考如何定义性能
Panagiotis Fytas, Ian Selby, Clemens Karner, Judith Babar, Simon Baker, Jake Beckford, Timothy J. Sadler, Shahab Shahipasand, Arthikkaa Thavakumar, John Li Chen, Alex Sawer, Michael Roberts, Jonathan Weir-McCall, J. H. F. Rudd, Carola-Bibiane Schönlieb, Anna Korhonen, Anna Breger
Comments† Equal contribution. Affiliations: 1: The University of Hong Kong 2: The University of Sydney 3: University of Electronic Science and Technology of China Corresponding authors: Zihan Deng (zhdeng@hku.hk), Chuanzhi Xu (chuanzhi.xu@sydney.edu.au) Project page: https://frankdengai.github.io/SciFigQual-Bench Source code & dataset: https://github.com/FrankDengAI/SciFigQual-Bench