SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents
SafeRelBench:用于VLM驱动的具身智能体过程级安全的空间关系感知基准测试
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,字节跳动公司)
专题命中 评测与基准 :language model(abstract);分类 cs.AI
AI总结 研究VLM驱动具身智能体的过程级安全问题,引入SAFERELBENCH基准测试,含507个样本。评估七个智能体发现任务成功与安全合规有差距,该基准明确测试行动前安全条件,凸显空间关系在安全评估中的核心地位。
Comments Preprint. 10 pages, 6 figures, 4 tables