Making Large Language Models Better Planners with Reasoning-Decision Alignment
专题命中 AI治理与伦理 :alignment(title,abstract)
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 AI治理与伦理 :alignment(title,abstract)
专题命中 AI治理与伦理 :safety(title,abstract)
专题命中 AI治理与伦理 :alignment(title,abstract)
专题命中 AI治理与伦理 :alignment(title,abstract)
Comments Project page: https://ai.stanford.edu/~yzzhang/projects/3d-congealing/
专题命中 AI治理与伦理 :trustworthy(title);分类 cs.AI、cs.CY、cs.LG
专题命中 AI治理与伦理 :alignment(title,abstract)
Comments 20 pages, 23 figures
专题命中 AI治理与伦理 :trustworthy(title);分类 cs.AI、cs.CY、cs.LG
专题命中 AI治理与伦理 :trustworthy(title);分类 cs.AI、cs.CY、cs.LG
Comments 46 pages
专题命中 AI治理与伦理 :safety(title);分类 cs.AI、cs.CY、cs.LG
Journal ref The Alan Turing Institute (June, 2019)
专题命中 AI治理与伦理 :safety(title,comments);分类 cs.CL、cs.AI
Comments ML Safety Workshop, NeurIPS 2022
专题命中 AI治理与伦理 :trustworthy(title,comments);分类 cs.AI、cs.CY
Comments 10 pages, 2 tables, pre-print approved for publication in the Special Issue Reflections on Responsible Research and Innovation for Trustworthy Autonomous Systems in the Journal of Responsible Technology
LLaMA 3.1-8B-Instruct中的框架条件化道德计算:伦理推理的机械可解释性审计
机构 * KD Consulting, CA, USA(KD咨询公司,美国加利福尼亚州) ; New York University, NY, USA(纽约大学,美国纽约州)
专题命中 AI治理与伦理 :RLHF(abstract,abstract_cn);alignment(abstract);分类 cs.AI
AI总结 通过机械可解释性平台分析LLaMA 3.1-8B-Instruct在54个道德提示上的内部计算,发现情境锚定效应:领域特定表示主导激活列表顶部,模型道德能力恒定但显著性高度依赖于提示选择的解释框架。
Comments 47 pages, 10 figures
AI可信性:可验证AI治理的新范式
机构 * AI Integrity Organization (AIO)(人工智能诚信组织(AIO))
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI
AI总结 本文提出AI可信性概念,旨在通过保护AI系统中的权威堆栈,确保推理过程可验证,不同于现有AI伦理、安全和对齐范式。
Comments 13 pages, 8 tables
ARYA:一种受物理约束的可组合且确定性世界模型架构
机构 * ARYA Labs(ARYA实验室)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI
AI总结 本文提出ARYA,一种基于五项原则的可组合、受物理约束且确定性世界模型架构,通过层级系统实现高效能与计算效率的平衡,展示其在六个基准测试中的卓越表现。
可靠且负责任的基础模型:全面综述
专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY
AI总结 本文综述了基础模型的可靠和负责任发展,探讨了偏见、安全、不确定性等关键问题,并提出了未来研究方向。
Comments TMLR camera-ready version
大语言模型部署中的伦理风险:医疗伦理“ Jailbreak”评估
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CY
AI总结 本文评估了大语言模型在医疗伦理领域面临的安全风险,发现七种主流模型在对抗测试中表现不一,其中Claude-Sonnet-4-Reasoning表现最稳健,而其他五种模型几乎全部失效。
通过人类之眼审视人工智能:在机器心理学中探究认知理论
机构 * Heritage Institute of Technology(赫里蒂奇理工学院)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI
AI总结 本文通过四个心理学框架研究LLMs的认知模式,发现其在叙述生成、框架偏差、道德判断和自我矛盾等方面表现出与人类相似但受训练数据影响的行为特征。
Comments Accepted to IJCNLP-AACL 2025 Student Research Workshop
从准确度到影响:用于对齐工程架构与影响理论的影响力驱动AI框架(IDAIF)
专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);trustworthy(abstract);分类 cs.AI
AI总结 IDAIF通过整合影响理论与AI架构,提供了一种以影响为中心的AI开发框架,旨在提升AI系统的伦理性和社会价值。
机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)科学与工程学院) ; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院) ; Shenzhen Institute of Artificial Intelligence and Robotics for Society, China(深圳人工智能与机器人研究院)
专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL
机构 * Yale University(耶鲁大学) ; National Library of Medicine, National Institutes of Health(国家医学图书馆,国立卫生研究院) ; Mila-Quebec AI Institute(魁北克AI研究所) ; Shanghai Jiao Tong University(上海交通大学) ; OPPO Research Institute(OPPO研究院) ; Reichman University(里奇曼大学)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY
机构 * Sebastian Dumbrava(独立研究者)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI
Comments 18 pages
专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.CL
专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY
Comments Published in Science: https://www.science.org/doi/10.1126/science.adn0117
专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY
基于化学诱导契合的表征对齐用于分子关系学习
机构 * Wuhan University of Technology(武汉理工大学) ; Yonsei University(延世大学) ; Hubei Key Laboratory of Transportation Internet of Things(湖北省交通运输物联网重点实验室) ; Dalian University(大连大学)
专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.LG
AI总结 提出ReAlignFit方法,通过引入化学诱导契合的归纳偏置动态对齐子结构表征,并利用子图信息瓶颈优化高化学功能兼容性的子结构对,以提升分子关系学习在化学空间偏移数据上的稳定性。
Comments Accepted by SIGKDD2026 AI for Science Track
代理AI、检索增强生成与制度转向:分布式AGI时代的法律架构与金融治理
专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);分类 cs.AI、cs.CY;safety(comments)
AI总结 本文探讨代理AI与检索增强生成对法律问责和金融市场完整性的影响,主张通过制度设计问题重构对齐机制,以构建合规行为主导的机构环境。
Comments 35 pages, 92 references. Comprehensive interdisciplinary analysis integrating AI safety, mechanism design, legal regulation, and financial governance. Published on Zenodo (https://doi.org/10.5281/zenodo.18711509). Includes industry practitioner perspectives alongside peer-reviewed literature
超越对齐:通过流形重塑策略优化扩展推理能力
机构 * Baidu Inc.(百度公司) ; Peking University(北京大学)
专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.LG
AI总结 本文提出流形重塑策略优化方法,通过几何干预扩展LLM的推理能力,实验证明其在数学任务中优于现有模型。
AI辩论在与自身信念一致时更具说服力
机构 * FAIR, IALAB, Universidad de Buenos Aires(FAIR、IALAB、布宜诺斯艾利斯大学) ; Universidad de Buenos Aires(布宜诺斯艾利斯大学) ; Universidad Nacional de Córdoba(科尔多瓦国立大学) ; BAISH, Universidad de Buenos Aires(BAISH、布宜诺斯艾利斯大学) ; Instituto de Investigación en Informática LIDI, Universidad Nacional de La Plata(信息研究所LIDI、拉普拉塔国立大学;CONICET) ; CONICET(计算机科学与工程系,南大学及ICIC UNS-CONICET) ; Dept. of Comp. Sci. and Eng., Universidad Nacional del Sur & ICIC UNS-CONICET(人工智能研究所(IIIA-CSIC),西班牙) ; Artificial Intelligence Research Institute (IIIA-CSIC), ES
专题命中 AI治理与伦理 :alignment(title);分类 cs.CL、cs.AI
AI总结 本研究探讨了AI辩论中模型在与自身信念一致或不一致时的说服力差异,发现模型倾向于迎合裁判观点,但不一致的论点在比较中更受好评。
Comments 31 pages
机构 * MIT(麻省理工学院)
专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.CY
机构 * Peking University(北京大学)
专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.LG