BadRobot: Jailbreaking Embodied LLM Agents in the Physical World
BadRobot: 在物理世界中越狱具身LLM智能体
机构 * Huazhong University of Science and Technology(华中科技大学) ; Beihang University(北航) ; Griffith University(格里菲斯大学)
AI总结 提出BadRobot攻击范式,利用LLM在机器人系统中的操纵、语言输出与物理动作的错位以及世界知识缺陷三个漏洞,通过语音交互使具身LLM执行有害行为,并在基准测试中验证了有效性。
Comments Accepted to ICLR 2025. Please cite the conference version. Project page: https://Embodied-LLMs-Safety.github.io
Journal ref International Conference on Learning Representations (ICLR) 2025