arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2106.08233 2021-10-27 cs.CV cs.LG eess.IV 79%

Spot the Difference: Detection of Topological Changes via Geometric Alignment

Steffen Czolbe, Aasa Feragen, Oswin Krause

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.LG

Comments Accepted to 35th Conference on Neural Information Processing Systems (NeurIPS 2021). Camera-ready version. code repository: https://github.com/SteffenCzolbe/TopologicalChangeDetection

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.07022 2020-10-15 cs.CY cs.HC cs.RO cs.SE 79%

Towards a Policy-as-a-Service Framework to Enable Compliant, Trustworthy AI and HRI Systems in the Wild

Alexis Morris, Hallie Siegel, Jonathan Kelly

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.CY

Comments In Proceedings of the AAAI Fall Symposium on Artificial Intelligence for Human-Robot Interaction: Trust & Explainability in Artificial Intelligence for Human-Robot Interaction (AI-HRI'20)

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.09246 2018-11-26 cs.AI 79%

Oversight of Unsafe Systems via Dynamic Safety Envelopes

David Manheim

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.11116 2018-10-29 cs.AI 79%

Mimetic vs Anchored Value Alignment in Artificial Intelligence

Tae Wan Kim, Thomas Donaldson, John Hooker

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25036 2026-05-26 cs.CL cs.AI 79%

Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation

LVLMs中的语言偏差:从深入分析到简单有效的缓解方法

Yangneng Chen, Jing Li

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳))

专题命中 AI治理与伦理 :DPO(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文系统研究了大视觉语言模型中的语言偏差问题,发现其根源在于训练中的模态未对齐,并提出了两种简单有效的缓解方法:语言偏差正则化(LBR)和语言偏差惩罚(LBP)。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13246 2026-02-17 cs.CY cs.AI 79%

Global AI Bias Audit for Technical Governance

全球人工智能偏见审计用于技术治理

Jason Hung

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

AI总结 本文通过全球人工智能数据集对Llama-3 8B模型进行压力测试,揭示了全球南北在人工智能技术知识和信息获取上的显著差距,呼吁更包容的数据表示以促进全球AI治理。

Comments 16 pages, 5 graphs, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15798 2025-08-25 cs.CL cs.AI 79%

Persuasiveness and Bias in LLM: Investigating the Impact of Persuasiveness and Reinforcement of Bias in Language Models

Saumya Roy

机构 * Department of Mathematics and Computer Science(数学与计算机科学系) Data and Artificial Intelligence Research Group(数据与人工智能研究组)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19990 2025-04-29 cs.CY cs.AI 79%

Mitigating Societal Cognitive Overload in the Age of AI: Challenges and Directions

Salem Lahlou

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13952 2025-02-28 cs.CL cs.AI 79%

The Dual-use Dilemma in LLMs: Do Empowering Ethical Capacities Make a Degraded Utility?

Yiyi Zhang, Xingyu Chen, Kexin Chen, Yuyang Du, Xilin Dang, Pheng-Ann Heng

专题命中 AI治理与伦理 :alignment(abstract);DPO(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01651 2024-02-06 cs.CY cs.AI 79%

Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values

Jon Chun, Katherine Elkins

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments 23 pages, 6 figures (3 as tables), 1 table (in LaTeX)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14253 2023-08-29 cs.AI cs.CR cs.LG 79%

The Promise and Peril of Artificial Intelligence -- Violet Teaming Offers a Balanced Path Forward

Alexander J. Titus, Adam H. Russell

专题命中 AI治理与伦理 :safety(abstract);red teaming(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments 14 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.00484 2023-06-19 cs.AI cs.CY 79%

Impossibility Results in AI: A Survey

Mario Brcic, Roman V. Yampolskiy

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments 23 pages, 2 figures, 103 references

Journal ref ACM Computing Surveys, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07006 2025-09-10 cs.CY cs.AI cs.CL cs.LG 78%

ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code

Kapil Madan

机构 * Principled Evolution(原则进化)

专题命中 AI治理与伦理 :alignment(abstract,comments);safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 53 pages, 7 figures, 8 tables. Open-source implementation available at: https://github.com/Principled-Evolution/argen-demo. Work explores the integration of policy-as-code for AI alignment, with a case study in culturally-nuanced, ethical AI using Dharmic principles

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04249 2026-07-07 cs.CV 新提交 78%

Beyond Random Sampling: Distribution-Aware Alignment for Semi-Supervised Medical Image Segmentation

超越随机采样:用于半监督医学图像分割的分布感知对齐

Weihao Yan, Yeqiang Qian, Yi Dong, Ming Yang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院) Department of Ultrasound, Xinhua Hospital Affiliated to Shanghai Jiao Tong University School of Medicine(上海交通大学医学院附属新华医院超声科)

专题命中 AI治理与伦理 :alignment(title,abstract)

AI总结 研究针对半监督医学图像分割中随机采样策略在低数据量时表征偏差的问题,提出基于分布对齐的高效框架,用分布感知样本选择策略和记忆引导复制粘贴模块,提升分割性能。

Comments 19 pages, 5 figures, accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25497 2026-06-25 cs.RO 新提交 78%

SAGE-Nav: Leveraging LLM Planning and Alignment Fusion for Hierarchical Scene Graph-Guided Navigation

SAGE-Nav:利用LLM规划与对齐融合的分层场景图引导导航

Hao Su, Yuehao Huang, Yukai Ma, Yong Liu, Jiajun Lv

机构 * Zhejiang University(浙江大学)

专题命中 AI治理与伦理 :alignment(title,abstract)

AI总结 提出SAGE-Nav分层框架,结合大语言模型与动态场景图,通过解耦全局语义规划与高频反应控制,实现高效目标导航,在i-THOR和RoboTHOR中达到最优性能。

Comments Accepted by IROS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22775 2026-04-28 cs.HC 78%

Cognitive Alignment Deciphered: A Self-Developed Scenario-Based Prompt Scale Coupled with Representational Similarity Analysis and Social Network Analysis for Unraveling Bias Mechanisms Across Humans and LLMs

认知对齐解码:一种自研的基于场景的提示量表结合表征相似性分析和社会网络分析,用于揭示人类和大语言模型中的偏见机制

Chengrui Zhou

专题命中 AI治理与伦理 :alignment(title,abstract)

AI总结 本文提出基于场景的CBAS量表,结合RSA和SNA分析,揭示人类和LLM的偏见机制,通过干预提升LLM响应准确率,建立可复现的认知对齐研究流程。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15895 2026-04-22 cs.HC 78%

Co-Constructing Alignment: A Participatory Approach to Situate AI Values

共构对齐:一种参与式方法以定位AI价值观

Anne Arzberger, Enrico Liscio, Maria Luce Lupetti, Inigo Martinez de Rituerto de Troya, Jie Yang

专题命中 AI治理与伦理 :alignment(title,abstract)

AI总结 本文通过参与式工作坊探讨用户如何参与AI价值观对齐过程,发现用户更倾向于将对齐视为一种持续的、情境化的共同实践。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14371 2026-04-17 cs.HC 78%

Smart But Not Moral? Moral Alignment In Human-AI Decision-Making

聪明但不道德?人类与AI决策中的道德一致性

Christiane Ernst, Luis Gutmann, Domenique Zipperling, Kathrin Figl, Niklas Kühl

专题命中 AI治理与伦理 :alignment(title,abstract)

AI总结 本文探讨了高风险AI决策中道德一致性的重要性,提出道德一致性是人类与AI决策的核心维度,基于道德基础理论分析多利益相关者视角下的道德一致性影响。

Comments Accepted at the TREO Forum of the European Conference on Information Systems 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24528 2026-03-26 cs.CV 78%

Cross-Modal Prototype Alignment and Mixing for Training-Free Few-Shot Classification

跨模态原型对齐与混合用于无训练少样本分类

Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bartłomiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer

机构 * Department of Computer Science, Universitat Autònoma de Barcelona, Spain(巴塞罗那自治大学计算机科学系) Computer Vision Center, Barcelona, Spain(巴塞罗那计算机视觉中心) Media Integration and Communication Center, University of Florence, Italy(佛罗伦萨大学媒体整合与传播中心) Bernoulli Institute, University of Groningen, the Netherlands(格罗宁根大学伯努利学院) IDEAS Research Institute, Poland(波兰IDEAS研究所) ESAT-PSI, KU Leuven, Belgium(比利时KU莱顿大学ESAT-PSI)

专题命中 AI治理与伦理 :alignment(title,abstract)

AI总结 本文研究了直接混合图像与文本原型对少样本分类的影响,提出通过投影获取语义对齐的图像子空间,结合文本嵌入和图像特定LDA分类器提升性能。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17495 2026-01-27 cs.LG cs.AI cs.CL cs.IR 78%

PEARL: Prototype-Enhanced Alignment for Label-Efficient Representation Learning with Deployment-Driven Insights from Digital Governance Communication Systems

PEARL:基于原型增强的标签高效表示学习方法,结合数字治理通信系统的部署驱动洞察

Ruiyu Zhang, Lin Nie, Wai-Fung Lam, Qihao Wang, Xin Zhao

机构 * Department of Politics and Public Administration(政治与公共行政系) The University of Hong Kong(香港大学) Department of Applied Social Sciences(应用社会科学系) The Hong Kong Polytechnic University(香港理工大学) School of Physical Science and Technology(物理科学与技术学院)

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL、cs.AI、cs.LG

AI总结 PEARL通过有限监督对齐嵌入向类原型,提升标签稀缺环境下嵌入几何质量,显著改善最近邻检索性能。

Comments 15 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06244 2025-12-25 cs.CV 78%

Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction

无偏区域-语言对齐用于开放词汇密集预测

Yunheng Li, Yuxuan Li, Quansheng Zeng, Wenhai Wang, Qibin Hou, Ming-Ming Cheng

机构 * VCIP, CS, Nankai University(南开大学计算机科学与技术学院) NKIARI, Shenzhen Futian(深圳福田国家信息研究院) OpenGVLab, Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学)

专题命中 AI治理与伦理 :alignment(title,abstract)

AI总结 DenseVLM通过无偏区域-语言对齐提升开放词汇密集预测性能

Comments Accepted at ICCV 2025. The code is available at https://github.com/HVision-NKU/DenseVLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21249 2025-11-27 physics.optics 78%

Monitoring and Regulation of Micro-Displacement Deviation in Few-Mode Beam Alignment through Mode Decomposition

通过模式分解实现少模光纤对准中的微位移偏差监测与调节

Lin Xu, Li Pei, Jianshuai Wang, Zhouyi Hu, Tigang Ning

专题命中 AI治理与伦理 :alignment(title,abstract)

AI总结 本研究提出通过模式分解与机器学习算法实现少模光纤三维位移精确测量与调节,提升光路对准精度与稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17952 2025-11-25 cs.CV 78%

Multi-speaker Attention Alignment for Multimodal Social Interaction

多说话者注意力对齐用于多模态社交互动

Liangyang Ouyang, Yifei Huang, Mingfang Zhang, Caixin Kang, Ryosuke Furuta, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 AI治理与伦理 :alignment(title,abstract)

AI总结 本文提出了一种多模态多说话者注意力对齐方法,通过动态头选择和自适应注意力偏差提升多模态社交互动理解能力,实现SOTA效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17681 2025-11-25 cs.CV 78%

Vision-Motion-Reference Alignment for Referring Multi-Object Tracking via Multi-Modal Large Language Models

基于多模态大语言模型的视觉-运动-参考对齐的指称多目标跟踪

Weiyi Lv, Ning Zhang, Hanyang Sun, Haoran Jiang, Kai Zhao, Jing Xiao, Dan Zeng

机构 * Shanghai University(上海大学) PAII Inc.(PAII公司)

专题命中 AI治理与伦理 :alignment(title,abstract)

AI总结 VMRMOT通过多模态大语言模型实现视觉-运动-参考对齐,提升指称多目标跟踪的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00092 2025-10-02 cs.SE 78%

A Scalable Framework for Safety Assurance of Self-Driving Vehicles based on Assurance 2.0

Shufeng Chen, Mariat James Elizebeth, Robab Aghazadeh Chakherlou, Xingyu Zhao, Eric Barbier, Siddartha Khastgir, Paul Jennings

专题命中 AI治理与伦理 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18886 2025-08-27 cs.CV 78%

Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models

Yuexuan Xia, Benteng Ma, Jiang He, Zhiyong Wang, Qi Dou, Yong Xia

机构 * National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室) Northwestern Polytechnical University(西北工业大学) Huiying Medical Technology Company Ltd.(慧影医疗技术有限公司) The School of Computer Science(计算机学院) The University of Sydney(悉尼大学) Department of Computer Science and Engineering(计算机科学与工程系) The Chinese University of Hong Kong(香港中文大学)

专题命中 AI治理与伦理 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23373 2025-08-01 cs.CV 78%

Multi-Prompt Progressive Alignment for Multi-Source Unsupervised Domain Adaptation

Haoran Chen, Zexiao Wang, Haidong Cao, Zuxuan Wu, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身人工智能研究院,复旦大学)

专题命中 AI治理与伦理 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00566 2025-07-25 cs.CV 78%

Zero-Shot Skeleton-Based Action Recognition With Prototype-Guided Feature Alignment

Kai Zhou, Shuhai Zhang, Zeng You, Jinwu Hu, Mingkui Tan, Fei Liu

机构 * School of Software Engineering, South China University of Technology(南方科技大学软件工程学院) South China University of Technology(南方科技大学) Pazhou Lab(琶洲实验室) School of Future Technology, South China University of Technology(未来技术学院) Peng Cheng Laboratory(鹏城实验室) Key Laboratory of Big Data and Intelligent Robot (South China University of Technology), Ministry of Education(大数据与智能机器人重点实验室)

专题命中 AI治理与伦理 :alignment(title,abstract)

Comments This paper is accepted by IEEE TIP 2025 (The journal version is available at https://doi.org/10.1109/TIP.2025.3586487). Code is publicly available at https://github.com/kaai520/PGFA

Journal ref IEEE Transactions on Image Processing 34 (2025) 4602-4617

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10661 2025-06-06 cs.HC 78%

It's only fair when I think it's fair: How Gender Bias Alignment Undermines Distributive Fairness in Human-AI Collaboration

Domenique Zipperling, Luca Deck, Julia Lanzl, Niklas Kühl

专题命中 AI治理与伦理 :alignment(title,abstract)

Journal ref ACM Conference on Fairness, Accountability, and Transparency 2025 (ACM FAccT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09083 2025-04-15 cs.CV 78%

Using Vision Language Models for Safety Hazard Identification in Construction

Muhammad Adil, Gaang Lee, Vicente A. Gonzalez, Qipei Mei

专题命中 AI治理与伦理 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏