The Forgotten Shield: Safety Grafting in Parameter-Space for Medical MLLMs
被遗忘的盾牌:参数空间中的医疗大语言模型安全性移植
Jiale Zhao, Xing Mou, Jinlin Wu, Hongyuan Yu, Mingrui Sun, Yang Shi, Xuanwu Yin, Zhen Chen, Zhen Lei, Yaohua Wang
机构
*
National University of Defense Technology(国防科技大学)
;
Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统(MAIS))
;
Multimedia Department, Xiaomi Inc(小米公司多媒体部门)
;
Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences, Hong Kong(香港科学院人工智能与机器人中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences, UCAS(中国科学院大学人工智能学院)
CommentsEmail by the arXiv Support Team: "Dear Gjergji, Thank you for your patience. Your appeal was accepted. You are welcome to resubmit this work at your convenience to cs.CY (cs.AI, cs.HC) Regards, arXiv Support"
Journal refAI and Ethics (2026) 6:295, Springer Nature
RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations
RiskNet:一个来自新闻的大规模AI风险事件数据集,包含对齐和多维标注
Leihan Zhang, Wecheng Ye, Xianlong Ma, Haochuan Liu, Yang Li, Qianyu Zhang, Jinliang Chen, Qiang Yan
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Beijing Key Laboratory of Multimodal Data Intelligent Perception and Governance(多模态数据智能感知与治理北京市重点实验室)
Alignment Risks from Capability-Seeking RL Training
从能力寻求强化学习训练中产生的对齐风险
Yujun Zhou, Yue Huang, Han Bao, Kehan Guo, Zhenwen Liang, Pin-Yu Chen, Tian Gao, Werner Geyer, Nuno Moniz, Nitesh V Chawla, Xiangliang Zhang
机构
*
University of California, Berkeley(加州大学伯克利分校)
;
Stanford University(斯坦福大学)
;
University of Washington(华盛顿大学)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
University of Toronto(多伦多大学)
;
University of Cambridge(剑桥大学)
机构
*
Kyoto University(京都大学)
;
Hohai University(河海大学)
;
The University of Tokyo(东京大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Hong Kong Polytechnic University(香港理工大学)
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
MoralityGym:用于评估序列决策代理中分层道德对齐的基准
Simon Rosen, Siddarth Singh, Ebenezer Gelo, Helen Sarah Robertson, Ibrahim Suder, Victoria Williams, Benjamin Rosman, Geraud Nangue Tasse, Steven James
Journal refProc of the 25th International Conference on Autonomous Agents and Multiagent Systems AAMAS 2026, Paphos, Cyprus, May 25 to 29, 2026, IFAAMAS
The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play
镜中的攻击者:通过锚定双策略自我博弈打破安全性中的自我一致性
Gabriele La Malfa, Emanuele La Malfa, Saar Cohen, Jie M. Zhang, Michael Luck, Michael Wooldridge, Elizabeth Black
机构
*
Department of Informatics, King’s College London(伦敦国王学院信息学院)
;
Department of Computer Science, University of Oxford(牛津大学计算机科学系)
;
University of Sussex(Sussex大学)
;
Institute for Decentralized AI(去中心化人工智能研究所)