机构
*
Qwen Large Model Application Team, Alibaba(阿里巴巴大模型应用团队)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Shenzhen Research Institute of Big Data(深圳大数据研究院)
From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails
从拒绝到恢复:一种生成AI防护机制的控制论方法
Ravi Pandya, Madison Bland, Duy P. Nguyen, Changliu Liu, Jaime Fernández Fisac, Andrea Bajcsy
机构
*
Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)
;
Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系)
;
Equal Advising(平等指导)
Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
在偏离时回溯:缓解大语言模型推理蒸馏中的双重暴露偏差
Bing Wang, Shaotian Yan, Chen Shen, kaiyuan liu, Sinan Fan, Ximing Li, Rui Miao, Xiaosong Yuan, Zhanming Shen, Jieping Ye
机构
*
College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院)
;
Key Laboratory of Symbolic Computation and Knowledge Engineering, MoE, Jilin University(吉林大学符号计算与知识工程重点实验室)
;
Tongyi Lab, Alibaba Group(阿里集团通义实验室)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
School of Artificial Intelligence, Jilin University(吉林大学人工智能学院)
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
通过场景分割策略对文本到视频模型进行劫持
Wonjun Lee, Haon Park, Doehyeon Lee, Bumsub Ham, Suhyun Kim
机构
*
Yonsei University(延世大学)
;
Korea Institute of Science and Technology(韩国科学技术院)
;
AIM Intelligence(AIM智能)
;
Seoul National University(首尔国立大学)
;
Kyung Hee University(庆熙大学)
Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection
财富低语:通过提示注入对谷歌的Agent支付协议进行红队测试
Tanusree Debi, Wentian Zhu, Pranjol Sen Gupta
机构
*
School of Computing University of Georgia Athens, USA(佐治亚大学计算机学院 美国亚特兰大)
;
Department of Information Technology Kennesaw State University Kennesaw, USA(凯斯韦尔州立大学信息科技学院 美国凯斯韦尔)