CommentsAfter further review of the current submission, we have identified potential legal and intellectual property concerns associated with keeping the manuscript publicly available as a preprint. In particular, there are ongoing considerations regarding institutional affiliation information, intellectual property ownership, and related compliance matters
Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI
缩小人工智能信任差距:可信人工智能独立认证的案例
Trisevgeni Papakonstantinou, Cansu Canca, Farah Nanji, Waheedullah Pardess, Jen Weedon, Jasmijn Remmers, Eliza Krigman, Matthew Ball, Yalda Daryani, Kiran Iqbal, Francielle Vargas, María Llorente Sánchez, Joe Humphreys, Fendi Tsim, Kelly Fitzpatrick, Jeff Dunn, Catherine Feldman
机构
*
University College London(伦敦大学学院)
;
AI Ethics Lab(人工智能伦理实验室)
;
University of Cambridge(剑桥大学)
;
Columbia University(哥伦比亚大学)
;
MATS
;
Tech with Intention(技术与意图)
;
University of Southern California(南加州大学)
;
Kyushu University(九州大学)
;
University of Chile(智利大学)
;
BehSci Meets AI(行为科学与人工智能)
;
Digital Trust Council(数字信任委员会)
CommentsAccepted as an oral presentation at HARMONY 2026 (Human-centered AI Research for Mental Health, an Open Networking Symposium), co-located with IEEE/ACM Conference on Connected Health: Applications, Systems, and Engineering Technologies (CHASE 2026) held in Pittsburgh, August 6, 2026
Safety-Aware Evaluation of LLM-Generated Driver Intervention Messages through Multi-Task Risk Fusion
通过多任务风险融合的LLM生成驾驶员干预信息的安全感知评估
Keito Inoshita
机构
*
Faculty of Business and Commerce, Kansai University(关西大学商学部)
;
Data Science and AI Innovation Research Promotion Center, Shiga University(滋贺大学数据科学与人工智能创新研究推进中心)
机构
*
University of Georgia(佐治亚大学)
;
University of South Florida(南佛罗里达大学)
;
Rutgers University(罗格斯大学)
;
University of Southern California(南加州大学)
;
Johns Hopkins University(约翰霍普金斯大学)
机构
*
School of Artificial Intelligence, Shanghai Jiao Tong University, Shanghai, China(上海交通大学人工智能学院)
;
Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室)
;
University of Science and Technology of China, Hefei, Anhui, China(中国科学技术大学)
;
The Chinese University of Hong Kong, Hong Kong, China(香港中文大学)
;
University of Illinois Urbana-Champaign, Urbana, IL, USA(伊利诺伊大学厄巴纳-香槟分校)
Journal refProceedings of The fourth international workshop on the role of resources in the age of large language models RESOURCEFUL-2026 at LREC 2026, Palma de Mallorca, Spain, 2026
Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization
对齐但脆弱:通过零阶优化增强LLM安全鲁棒性
Zhihao Liu, Yifan Wu, Jian Lou, Di Wang, Yuxi Zhou, Yuke Hu
机构
*
The State Key Laboratory of Blockchain and Data Security(区块链与数据安全国家重点实验室)
;
Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高科技园区(滨江)区块链与数据安全研究院)
;
Sun Yat-sen University(中山大学)
;
KAUST(卡塔尔大学)
AI总结
本文通过 Hateful Memes Challenge 数据集系统分析 GPT-4o mini 在多模态仇恨言论检测中的安全架构,发现并实验验证了“单模态瓶颈”缺陷,即上下文无关的安全过滤器会优先阻断多模态推理,导致误报。
CommentsThis paper reports preliminary findings from a small-scale study whose sample size is insufficient to support the stated conclusions. The authors are withdrawing it to conduct a more comprehensive evaluation
CommentsThe Algorithm 1 is not entirely correct and they may affect the results as well. We are restarting the experimentations and will upload the new version as soon as possible