arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2502.14888 2026-04-28 cs.CV cs.AI 79%

Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models

超越跨模态对齐:测量和利用模态差距在视觉-语言模型中

Hanqi Yan, Xiangxiang Cui, Lu Yin, Jindong Gu, Paul Pu Liang, Yulan He, Yifei Wang

机构 * King’s College London(伦敦国王学院) University of Surrey(萨里大学) University of Oxford(牛津大学) MIT CSAIL(麻省理工学院CSAIL实验室) The Alan Turing Institute(阿兰·图灵研究院)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出通过测量和利用模态差距改进视觉-语言模型的下游任务,引入模态主导分数和自动可解释性度量,实现轻量级编辑和系统分析。

Comments accepted by ACL26-findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14133 2026-04-16 cs.AI cs.CR cs.MA 79%

Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems

形式化代理AI系统的安全、安全性和功能性属性

Edoardo Allegrini, Ananth Shreekumar, Z. Berkay Celik

机构 * Sapienza University of Rome(罗马萨皮恩扎大学) Purdue University(普渡大学)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI

AI总结 本文提出一个统一的语义框架,用于分析、设计和部署正确、可靠且稳健的代理AI系统,通过形式化验证确保系统行为的安全性和功能性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12851 2026-04-15 cs.CY 79%

Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment

基于人物提示的LLM能否模拟子群体价值观?对文化契合性一般性和公平性的实证分析

Bryan Chen Zhengyu Tan, Zhengyuan Liu, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Roy Ka-Wei Lee

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CY

AI总结 本文探讨LLM在细粒度价值观对齐中的挑战,通过新加坡案例研究发现,即使使用GPT-4.1等先进模型,预测子群体偏好准确率仅57.4%,但通过结构化数值偏好微调可提升17.4%,但存在预存的公平性偏见。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04265 2026-04-07 cs.CR cs.AI cs.MA 79%

Governance-Constrained Agentic AI: Blockchain-Enforced Human Oversight for Safety-Critical Wildfire Monitoring

受治理约束的代理AI:区块链强制的人类监督用于安全关键的野火监测

Ali Akarma, Toqeer Ali Syed, Salman Jan, Hammad Muneer, Abdul Khadar Jilani

机构 * AI Center, Faculty of Computer and Information Systems, Islamic University of Madinah(伊斯兰大学麦地那分校计算机与信息系统学院人工智能中心) Faculty of Computer Studies, Arab Open University-Bahrain(阿拉伯开放大学巴林分校计算机研究学院) Department of Computer Science, The Islamia University of Bahawalpur(巴哈瓦尔布尔伊斯兰大学计算机科学系) College of Computer Studies, University of Technology Bahrain(巴林科技大学计算机研究学院)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI

AI总结 本文提出基于区块链的治理意识代理AI架构,用于安全关键的野火预警,通过约束部分可观测马尔可夫决策过程模型,实现动态风险适应的无人机资源分配,并通过智能合约强制人类授权,确保警报完整性、人类控制和非否认性。

Comments This paper was presented at ICETAS 2026 Bahrain

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09232 2026-04-01 cs.CL cs.SD 79%

POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation

POTSA:一种用于语音到文本翻译的跨语言语音对齐框架

Xuanchen Li, Chenrui Cui, Tianrui Wang, Meng Ge, Zikang Huang, Yizhou Peng, Jin Li, Yuheng Lu, Yu Jiang, Nyima Tashi, Longbiao Wang, Jianwu Dang

机构 * Tianjin Key Laboratory of Cognitive Computing and Application, Tianjin University(天津大学认知计算与应用天津市重点实验室) Nanyang Technological University(南洋理工大学) Huiyan Technology Company, Ltd.(慧眼科技有限公司) Chinese Academy of Sciences(中国科学院) Tibet University(西藏大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL

AI总结 本文提出POTSA框架,通过最优传输和跨语言平行语音对缓解多语言语音翻译中的语义偏见问题,实验显示在FLEURS数据集上取得SOTA性能,提升BLEU分数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21435 2026-03-24 cs.AI econ.GN q-fin.EC 79%

Behavioural feasible set: Value alignment constraints on AI decision support

行为可行集:人工智能决策支持中的价值对齐约束

Taejin Park

机构 * Bank for International Settlements(国际清算银行)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

AI总结 研究探讨了人工智能决策支持系统中价值对齐约束对行为可行集的影响,通过实验显示对齐显著压缩了推荐范围,并揭示了组织在选择供应商时面临的价值导向问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20604 2026-01-29 cs.AI 79%

Dialogical Reasoning Across AI Architectures: A Multi-Model Framework for Testing AI Alignment Strategies

跨AI架构的对话推理:一种多模型框架用于测试AI对齐策略

Gray Cox

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出一种多模型框架,通过对话推理测试AI对齐策略,展示不同AI模型在处理复杂对齐框架时的表现及新兴见解。

Comments 23 pages, 5 tables, 5 appendices. Code and data: https://github.com/jgraycox-coa/vcw-multi-ai-dialogue

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16444 2026-01-27 cs.CL 79%

Exploring the Effects of Alignment on Numerical Bias in Large Language Models

探索对大型语言模型中数值偏差的影响

Ayako Sato, Hwichan Kim, Zhousi Chen, Masato Mita, Mamoru Komachi

机构 * Tokyo Metropolitan University(东京 Metropolitan 大学) Hitotsubashi University(立命馆大学) CyberAgent Inc.(CyberAgent 公司)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL

AI总结 本文研究了对齐对大型语言模型数值偏差的影响,发现对齐会增加偏差,并提出分数范围调整作为缓解策略。

Comments Accepted at AIBSD 2026 (Workshop at AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09212 2026-01-21 cs.CL 79%

Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment

针对对齐偏差:一种基于冲突意识的奖励模型驱动LLM对齐框架

Zixuan Liu, Siavash H. Khajavi, Guangkai Jiang, Xinru Liu

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL

AI总结 本文提出了一种基于冲突意识的框架,通过识别和缓解代理模型与策略间的冲突来提升大语言模型的对齐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01495 2026-01-05 cs.CL 79%

C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models

C-VARC:一个大规模的中文价值观规则语料库用于大语言模型的价值对齐

Ping Wu, Guobin Shen, Dongcheng Zhao, Yuwei Wang, Yiting Dong, Yu Shi, Enmeng Lu, Feifei Zhao, Yi Zeng

机构 * BrainCog Lab, Institute of Automation, Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所脑认知实验室,北京,中国) Beijing Key Laboratory of Safe AI and Superalignment, Beijing, China(北京安全人工智能与超对齐关键实验室,北京,中国) Beijing Institute of AI Safety and Governance, Beijing, China(北京人工智能安全与治理研究所,北京,中国) The School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学人工智能学院,北京,中国) Long-term AI, Beijing, China(长期人工智能,北京,中国)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL

AI总结 C-VARC通过构建大规模中文价值观规则语料库,提供了一种基于中国核心价值观的价值对齐方法,有效提升了大语言模型在不同文化背景下的伦理评估能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11581 2025-12-17 cs.AI econ.GN q-fin.EC 79%

Trustworthy, responsible, ethical AI in manufacturing and supply chains: synthesis and emerging research questions

制造业与供应链中的可信、负责任、道德AI:综合与新兴研究问题

Alexandra Brintrup, George Baryannis, Ashutosh Tiwari, Svetan Ratchev, Giovanna Martinez-Arellano, Jatinder Singh

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.AI

AI总结 本文探讨制造业中负责任、道德和可信AI的应用,提出研究问题以指导未来研究,确保AI应用的安全性和责任性。

Comments Pre-print under peer-review

Journal ref Data-Centric Engineering 6 (2025) e53

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06301 2025-12-10 cs.AI 79%

Large Language Models and Their Applications in Roadway Safety and Mobility Enhancement: A Comprehensive Review

大语言模型及其在道路安全与出行提升中的应用:全面综述

Muhammad Monjurul Karim, Yan Shi, Shucheng Zhang, Bingzhang Wang, Mehrdad Nasri, Yinhai Wang

机构 * Department of Civil and Environmental Engineering, University of Washington(土木与环境工程系,华盛顿大学)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI

AI总结 本文综述了大语言模型在道路安全与出行提升中的应用,探讨其在交通领域的适应策略及面临的挑战。

Journal ref Artificial Intelligence for Transportation, 1, 100004, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05383 2025-12-08 cs.SE cs.AI 79%

Fuzzing the brain: Automated stress testing for the safety of ML-driven neurostimulation

对大脑进行模糊测试:用于ML驱动的神经刺激安全性的自动化压力测试

Mara Downing, Matthew Peng, Jacob Granley, Michael Beyeler, Tevfik Bultan

机构 * Department of Computer Science, University of California, Santa Barbara, CA, USA(计算机科学系,加州大学圣巴巴拉分校)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI

AI总结 本文提出了一种基于模糊测试的自动化方法,用于检测和表征ML驱动神经刺激系统中的不安全刺激模式,以提升神经接口的安全性。

Comments 20 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08068 2025-12-02 cs.LG 79%

Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions

分位数奖励策略优化:与点估计回归和精确分区函数对齐

Simon Matrenok, Skander Moalla, Caglar Gulcehre

机构 * CLAIRE, EPFL(CLAIRE,瑞士联邦理工学院)

专题命中 AI治理与伦理 :alignment(title);DPO(abstract);分类 cs.LG

AI总结 QRPO通过分位数奖励实现点绝对奖励学习,保留DPO的简单性和离线适用性,同时提升在聊天和编码任务中的性能。

Comments 58 pages, NeurIPS2025 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19749 2025-11-26 cs.AI 79%

Scaling Item-to-Standard Alignment with Large Language Models: Accuracy, Limits, and Solutions

基于大语言模型的项目-标准对齐扩展:准确性、限制与解决方案

Farzan Karimi-Malekabadi, Pooya Razavi, Sonya Powers

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

AI总结 本研究探讨了大语言模型在提升项目-标准对齐效率方面的应用,通过实验发现LLMs在识别不一致项目和筛选候选技能方面表现优异,结合过滤策略可显著减少人工工作量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21456 2025-11-21 cs.CL 79%

Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes

在道德一致性中的性能权衡诊断:关于性别刻板印象的案例研究

Guangliang Liu, Bocheng Chen, Han Zi, Xitong Zhang, Kristen Marie Johnson

机构 * Michigan State University(密歇根州立大学) University of Mississippi(密苏里大学) Northeastern University(东北大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL

AI总结 本文通过分析性别刻板印象缓解中的性能权衡,发现当前公平性目标存在局限,整体遗忘与下游任务性能密切相关,选择性遗忘无法有效降低整体遗忘。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13931 2025-10-17 cs.CL 79%

Robust or Suggestible? Exploring Non-Clinical Induction in LLM Drug-Safety Decisions

Siying Liu, Shisheng Zhang, Indu Bala

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CL

Comments Preprint of a paper accepted as a poster at the NeurIPS 2025 Workshop on Generative AI for Health (GenAI4Health). The final camera-ready workshop version may differ. Licensed under CC BY 4.0

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12844 2025-10-16 cs.CY 79%

AI Alignment vs. AI Ethical Treatment: 10 Challenges

Adam Bradley, Bradford Saad

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CY

Comments author order is arbitrary

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04528 2025-10-07 cs.CR cs.AI 79%

Unified Threat Detection and Mitigation Framework (UTDMF): Combating Prompt Injection, Deception, and Bias in Enterprise-Scale Transformers

Santhosh KumarRavindran

机构 * Microsoft Corporation(微软公司)

专题命中 AI治理与伦理 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04073 2025-10-07 cs.AI 79%

Moral Anchor System: A Predictive Framework for AI Value Alignment and Drift Prevention

Santhosh Kumar Ravindran

机构 * Microsoft Corporation(微软公司)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

Comments 11 pages Includes simulations with over 4 million steps

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12754 2025-08-19 cs.AI 79%

Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants

Alessio Galatolo, Luca Alberto Rappuoli, Katie Winkle, Meriem Beloucif

机构 * Uppsala University(乌普萨拉大学) University of St. Andrews(圣安德鲁大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

Comments Full version of the paper published in ECAI 2025 proceedings (IOS Press, CC BY-NC 4.0)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09656 2025-08-08 cs.AI 79%

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives

Wei Zeng, Hengshu Zhu, Chuan Qin, Han Wu, Yihang Cheng, Sirui Zhang, Xiaowei Jin, Yinuo Shen, Zhenxing Wang, Feimin Zhong, Hui Xiong

机构 * Business School, Hunan University(湖南大学商学院) Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) University of Chinese Academy of Sciences(中国科学院大学) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) School of Business, Hunan University(湖南大学商学院) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能研究所) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR(香港科技大学(香港)计算机科学与工程学院)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22267 2025-07-31 cs.HC cs.AI 79%

Promoting Online Safety by Simulating Unsafe Conversations with LLMs

Owen Hoffman, Kangze Peng, Zehua You, Sajid Kamal, Sukrit Venkatagiri

机构 * Department of Computer Science, Swarthmore College(计算机科学系,斯沃斯里学院)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI

Journal ref ACM 2025 Conference on Conversational User Interfaces Workshop on Personas Evolved: Designing Ethical LLM-Based Conversational Agent Personalities

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05400 2025-07-10 cs.CY econ.GN q-fin.EC 79%

Strategic Alignment Patterns in National AI Policies

Mohammad Hossein Azin, Hessam Zandhessami

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23815 2025-07-02 cs.HC cs.AI 79%

The Impact of AI on Educational Assessment: A Framework for Constructive Alignment

Patrick Stokkink

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14502 2025-06-18 cs.AI 79%

Toward Safety-First Human-Like Decision Making for Autonomous Vehicles in Time-Varying Traffic Flow

Xiao Wang, Junru Yu, Jun Huang, Qiong Wu, Ljubo Vacic, Changyin Sun

机构 * University Scientific Research Program of Anhui Province(安徽省大学科学研究计划) National Natural Science Foundation of China(中国国家自然科学基金)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04253 2025-06-06 cs.AI cs.HC 79%

HADA: Human-AI Agent Decision Alignment Architecture

Tapio Pitkäranta, Leena Pitkäranta

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Aalto University(阿尔托大学) Department of Industrial Engineering and Management(工业工程与管理系)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

Comments 18 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18779 2025-05-27 cs.CY 79%

Evaluating Intra-firm LLM Alignment Strategies in Business Contexts

Noah Broestl, Benjamin Lange, Cristina Voinea, Geoff Keeling, Rachael Lam

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CY

Comments 9 pages, 0 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17112 2025-05-26 cs.CL 79%

Cultural Value Alignment in Large Language Models: A Prompt-based Analysis of Schwartz Values in Gemini, ChatGPT, and DeepSeek

Robin Segerer

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL

Comments 15 pages, 1 table, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07205 2025-05-13 cs.CL 79%

Benchmarking Ethical and Safety Risks of Healthcare LLMs in China-Toward Systemic Governance under Healthy China 2030

Mouxiao Bian, Rongzhao Zhang, Chao Ding, Xinwei Peng, Jie Xu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏