Repurposing Adversarial Perturbations for Continual Learning: From Defense to Active Alignment
重新利用对抗扰动进行持续学习:从防御到主动对齐
Ran Liu, Min Yu, Mingqi Liu, Jianguo Jiang, Gang Li, Rongsheng Li, Ning Li, Zhen Xu, Weiqing Huang, Ming Liu
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
Deakin University(德肯大学)
;
Harbin Engineering University(哈尔滨工程大学)
dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment
dashi: 一个用于数据集偏移表征以支持可信AI开发和部署的Python库
David Fernández-Narro, Pablo Ferri, Ángel Sánchez-García, Juan M. García-Gómez, Carlos Sáez
机构
*
Biomedical Data Science Lab, Instituto Universitario de Tecnologías de la Información y Comunicaciones, Universitat Politècnica de Valéncia(生物医学数据科学实验室,信息与通信技术大学,巴塞罗那理工大学)
Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities
通过反应语气建模社区态度:评估LLM与在线社区语言行为对齐的人机协作框架
Nuan Wen, Xuezhe Ma
机构
*
Information Sciences Institute University of Southern California(南加州大学信息科学研究所)
机构
*
Shanghai Academy of Artificial Intelligence for Science, Shanghai, China.(上海人工智能科学研究院)
;
School of Biomedical Engineering, Shanghai Jiao Tong University, Shanghai, China.(上海交通大学生物医学工程学院)
;
Incubation Institute, Fudan University, Shanghai, China.(复旦大学孵化院)
No More, No Less: Task Alignment in Terminal Agents
无需更多,无需更少:终端代理的任务对齐
Sina Mavali, David Pape, Jonathan Evertz, Samira Abedini, Devansh Srivastav, Thorsten Eisenhofer, Sahar Abdelnabi, Lea Schönherr
机构
*
CISPA Helmholtz Center for Information Security(CISPA 涅槃中心信息安全研究所)
;
ELLIS Institute Tübingen and MPI for Intelligent Systems(图宾根 ELLIS 研究所和智能系统 MPI)
;
Tübingen AI Center(图宾根人工智能中心)
A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?
对计算机使用代理的安全性和安全威胁的综述:贾维斯或乌tron?
Ada Chen, Yongjiang Wu, Junyuan Zhang, Jingyu Xiao, Shu Yang, Jen-tse Huang, Kun Wang, Wenxuan Wang, Shuai Wang
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
KAUST(卡塔尔科技大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
Nanyang Technological University(南洋理工大学)
;
Renmin University of China(中国人民大学)
;
The Hong Kong University of Science and Technology(香港科学大学)
Comments23 pages, 9 figures; editorial and LaTeX revisions for clarity; improved presentation of methodology and results; updated figures, tables, and float placement; clarified temperature sensitivity and deployment-risk analysis; expanded reporting from the same experiments; results unchanged in substance
Green Shielding: A User-Centric Approach Towards Trustworthy AI
绿色防护:面向可信AI的用户导向方法
Aaron J. Li, Nicolas Sanchez, Hao Huang, Ruijiang Dong, Jaskaran Bains, Katrin Jaradeh, Zhen Xiang, Bo Li, Feng Liu, Aaron Kornblith, Bin Yu
机构
*
University of California, Berkeley(加州大学伯克利分校)
;
University of Melbourne(墨尔本大学)
;
University of California, San Francisco(加州大学旧金山分校)
;
University of Georgia(佐治亚大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构
*
Nanjing University(南京大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
Can "AI" Be a Doctor? A Study of Empathy, Readability, and Alignment in Clinical LLMs
AI能成为医生吗?对临床LLM中共情、可读性和对齐性的研究
Mariano Barone, Francesco Di Serio, Roberto Moio, Marco Postiglione, Giuseppe Riccio, Antonio Romano, Vincenzo Moscato
机构
*
Department of Electrical Engineering and Information Technology(电子工程与信息科技系)
;
University of Naples Federico II(那不勒斯费德里科二世大学)
;
Department of Translational Medical Sciences(转化医学科学系)
;
University of Campania ”Luigi Vanvitelli”(坎帕尼亚“路易吉·范维蒂利”大学)
;
Department of Computer Science, McCormick School of Engineering and Applied Science(计算机科学系,麦科马克工程与应用科学学院)
;
Northwestern University(西北大学)
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
MHSafeEval:面向大语言模型中心理健康安全的交互层面角色感知评估
Suhyun Lee, Palakorn Achananuparp, Neemesh Yadav, Ee-Peng Lim, Yang Deng
机构
*
Department of Artificial Intelligence, Hanyang University(翰艺大学人工智能系)
;
School of Computing and Information Systems, Singapore Management University(新加坡国立管理学院信息系)
Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models
评估多模态大语言模型在住院诊断中的表现:在十个前沿模型上的真实世界性能、安全性和成本
Bruce A. Bassett, Amy Rouillard, Sitwala Mundia, Michael Cameron Gramanie, Linda Camara, Ziyaad Dangor, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Ismail Kalla, Haroon Saloojee
机构
*
School of Computer Science and Applied Mathematics, University of the Witwatersrand(沃斯兰大学计算机科学与应用数学学院)
;
Wits MIND Institute, University of the Witwatersrand(沃斯兰大学Wits MIND研究所)
;
Grai Labs, Cape Town, South Africa(南非开普敦Grai实验室)
;
Faculty of Health Sciences, University of the Witwatersrand(沃斯兰大学健康科学学院)
;
South African Medical Research Council Vaccines and Infectious Diseases Analytics Research Unit, Faculty of Health Sciences, University of the Witwatersrand(南非医学研究委员会疫苗与传染病分析研究单位,沃斯兰大学健康科学学院)
A Multi-Stage Validation Framework for Trustworthy Large-scale Clinical Information Extraction using Large Language Models
一种用于可信大规模临床信息提取的多阶段验证框架
Maria Mahbub, Gregory M. Dams, Josh Arnold, Caitlin Rizy, Sudarshan Srinivasan, Elliot M. Fielstein, Minu A. Aghevli, Kamonica L. Craig, Elizabeth M. Oliva, Joseph Erdos, Jodie Trafton, Ioana Danciu
机构
*
Oak Ridge National Laboratory(橡树岭国家实验室)
;
Program Evaluation and Resource Center, Office of Mental Health and Office of Suicide Prevention, Department of Veterans Affairs(退伍军人事务部心理健康办公室和自杀预防办公室项目评估与资源中心)
;
Vanderbilt University Medical Center(范德比尔特大学医学中心)
;
VA Maryland Health Care System(退伍军人事务部马里兰医疗保健系统)
;
VA Desert Pacific Healthcare Network(退伍军人事务部沙漠太平洋医疗网络)
;
VA Connecticut Health Care System(退伍军人事务部康涅狄格医疗保健系统)
;
Yale School of Medicine(耶鲁医学院)