机构
*
Tongji University(同济大学)
;
The City University of New York(纽约城市大学)
;
University of Technology Sydney(悉尼大学)
;
Huazhong University of Science and Technology(华中科技大学)
;
Shenzhen University of Advanced Technology(深圳先进技术大学)
SSP: Safety-guaranteed Surgical Policy via Joint Optimization of Behavioral and Spatial Constraints
SSP:通过行为和空间约束的联合优化实现安全的手术策略
Jianshu Hu, ZhiYuan Guan, Lei Song, Kantaphat Leelakunwet, Hesheng Wang, Wei Xiao, Qi Dou, Yutong Ban
机构
*
Global College, Shanghai Jiao Tong University(上海交通大学全球学院)
;
Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系)
;
Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)
Journal refIn Proceedings of the Second Conference of the International Association for Safe and Ethical Artificial Intelligence (IASEAI'26), Paris, France, 2026
"Dark Triad" Model Organisms of Misalignment: Narrow Fine-Tuning Mirrors Human Antisocial Behavior
黑暗三联征模型生物:对齐偏差:狭窄微调映射人类反社会行为
Roshni Lulla, Fiona Collins, Sanaya Parekh, Thilo Hagendorff, Jonas Kaplan
机构
*
Brain & Creativity Institute, University of Southern California(大脑与创造力研究所,南加州大学)
;
Department of Psychology, University of Southern California(心理学系,南加州大学)
;
Interchange Forum for Reflecting on Intelligent Systems, University of Stuttgart(智能系统反思交流论坛,斯图加特大学)
Journal refProceedings of the 2025 IEEE 7th International Conference on Trust, Privacy and Security in Intelligent Systems and Applications (TPS-ISA), Pittsburgh, PA, USA, November 12-14, 2025. IEEE
Stealth Fine-Tuning: Efficiently Breaking Alignment in RVLMs Using Self-Generated CoT
隐形微调:通过自动生成的CoT打破RVLMs的对齐
Le Yu, Zhengyue Zhao, Yawen Zheng, Yunhao Liu
机构
*
Machine Intelligence Laboratory, Sichuan University(四川大学人工智能实验室)
;
University of Wisconsin--Madison(威斯康星大学麦迪逊分校)
;
Department of Automation, Tsinghua University(清华大学自动化系)
;
Global Innovation Exchange, Tsinghua University(清华大学全球创新交流中心)