Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
对冲与非肯定:量化大语言模型在人权问题上的对齐
Rafiya Javed, Cassandra Parent, Jackie Kay, David Yanni, Abdullah Zaini, Anushe Sheikh, Maribeth Rauh, Walter Gerych, Ramona Comanescu, Iason Gabriel, Marzyeh Ghassemi, Laura Weidinger
机构
*
Google Deepmind(谷歌DeepMind)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Independent Researcher(独立研究员)
;
Google(谷歌)
;
AI Accountability Lab, Trinity College Dublin(都柏林圣三一学院人工智能问责实验室)
专题命中
其他LLM
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
MARL-GPT: Foundation Model for Multi-Agent Reinforcement Learning
MARL-GPT:多智能体强化学习的基础模型
Maria Nesterova, Mikhail Kolosov, Anton Andreychuk, Egor Cherepanov, Oleg Bulichev, Alexey Kovalev, Konstantin Yakovlev, Aleksandr Panov, Alexey Skrynnik
机构
*
MIRAI \& Innopolis University Moscow Russia
;
MIRAI \& Innopolis University
Foundations for Agentic AI Investigations from the Forensic Analysis of OpenClaw
从OpenClaw的取证分析为基础的代理AI研究基础
Jan Gruber, Jan-Niclas Hilgert
机构
*
KASTEL Security Research Labs, Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院 KASTEL 安全研究实验室)
;
Fraunhofer Institute for Communication, Information Processing and Ergonomics FKIE(弗劳恩霍夫通信、信息处理与人体工程学研究所 FKIE)
专题命中
其他LLM
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI
CommentsThis is the accepted version of a paper that will appear in the proceedings of the 21st International Conference on Evaluation of Novel Approaches of Software Engineering (ENASE 2026). The final published version will be available from Science and Technology Publications (SCITEPRESS). 15 pages, 3 figures, 7 tables