arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1824 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1824 篇

2507.21091 2025-07-30 cs.CY cs.AI 81%

The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment

Lenart Motnikar, Katharina Baum, Alexander Kagan, Sarah Spiekermann-Hoff

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments Thirty-Third European Conference on Information Systems (ECIS 2025), Amman, Jordan

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15181 2025-07-24 cs.CY cs.AI 81%

Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures

Lily Stelling, Mick Yang, Rokas Gipiškis, Leon Staufer, Ze Shen Chin, Siméon Campos, Ariel Gil, Michael Chen

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

Comments 166 pages, the Oxford Martin AI Governance Initiative

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02242 2025-06-19 cs.LG cs.CY 81%

From Street Views to Urban Science: Discovering Road Safety Factors with Multimodal Large Language Models

Yihong Tang, Ao Qu, Xujing Yu, Weipeng Deng, Jun Ma, Jinhua Zhao, Lijun Sun

机构 * McGill University(麦吉尔大学) Massachusetts Institute of Technology(麻省理工学院) The University of Hong Kong(香港大学)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00751 2025-06-03 cs.AI cs.LG 81%

Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?

Zhuojun Gu, Quan Wang, Shuchu Han

机构 * Department of Information Systems and Business Analytics(信息系统与商业分析系) University at Albany - State University of New York(阿尔巴尼大学 - 新 york 状态大学) Santa Clara, CA(圣克拉拉,加州) Princeton Junction, NJ(普林斯顿 junction,新泽西)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00136 2025-05-29 cs.CL cs.AI 81%

A Checks-and-Balances Framework for Context-Aware Ethical AI Alignment

Edward Y. Chang

机构 * stan(斯坦福大学) Computer Science, Stanford University(斯坦福大学计算机科学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments 20 pages, 7 tables, 6 figures. arXiv admin note: substantial text overlap with arXiv:2405.07076

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17841 2025-05-26 cs.CY cs.AI eess.AS 81%

TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation

Wiebke Hutiri, Mircea Cimpoi, Morgan Scheuerman, Victoria Matthews, Alice Xiang

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10231 2025-05-16 cs.CV cs.AI cs.CL 81%

On the Interplay of Human-AI Alignment,Fairness, and Performance Trade-offs in Medical Imaging

Haozhe Luo, Ziyu Zhou, Zixin Shu, Aurélie Pahud de Mortanges, Robert Berke, Mauricio Reyes

机构 * University of Bern(伯恩大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01031 2025-05-09 cs.CL cs.AI cs.SI 81%

ValuesRAG: Enhancing Cultural Alignment Through Retrieval-Augmented Contextual Learning

Wonduk Seo, Zonghao Yuan, Yi Bu

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00137 2025-04-30 cs.CL cs.AI 81%

Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment

Sangwon Yu, Jongyoon Song, Bongkyu Hwang, Hoyoung Kang, Sooah Cho, Junhwa Choi, Seongho Joe, Taehee Lee, Youngjune L. Gwon, Sungroh Yoon

机构 * Department of Electrical and Computer Engineering, Seoul National University(电子与计算机工程系,首尔国立大学) Samsung SDS, Korea(三星SDS,韩国) AIIS, ASRI, INMC, ISRC, and IPAI, Seoul National University(人工智能研究所、人工智能研究机构、智能网络中心、智能系统研究室和人工智能政策研究所,首尔国立大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments NAACL 2025 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10915 2025-04-23 cs.MA cs.AI cs.CY 81%

LOKA Protocol: A Decentralized Framework for Trustworthy and Ethical AI Agent Ecosystems

Rajesh Ranjan, Shailja Gupta, Surya Narayan Singh

专题命中 AI治理与伦理 :trustworthy(title);alignment(abstract);分类 cs.AI、cs.CY

Comments 4 Figures, 1 Table

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12911 2025-04-22 cs.CL cs.AI 81%

Benchmarking Multi-National Value Alignment for Large Language Models

Weijie Shi, Chengyi Ju, Chengzhong Liu, Jiaming Ji, Jipeng Zhang, Ruiyuan Zhang, Jia Zhu, Jiajie Xu, Yaodong Yang, Sirui Han, Yike Guo

机构 * The Hong Kong University of Science and Technology(香港科技大学) Peking University(北京大学) Zhejiang Normal University(浙江师范大学) Soochow University(苏州大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08972 2025-04-08 cs.CL cs.AI 81%

Tuning-Free Personalized Alignment via Trial-Error-Explain In-Context Learning

Hyundong Cho, Karishma Sharma, Nicolaas Jedema, Leonardo F. R. Ribeiro, Alessandro Moschitti, Ravi Krishnan, Jonathan May

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments NAACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01947 2025-03-06 cs.CL cs.CY 81%

Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts

Akito Nakanishi, Yukie Sano, Geng Liu, Francesco Pierri

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CL、cs.CY

Comments This paper has been submitted to IEEE Transactions on Artificial Intelligence for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09402 2025-01-07 cs.CY cs.AI cs.GT 81%

CERN for AI: A Theoretical Framework for Autonomous Simulation-Based Artificial Intelligence Testing and Alignment

Ljubisa Bojic, Matteo Cinelli, Dubravko Culibrk, Boris Delibasic

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments 32 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17686 2024-12-25 cs.AI cs.CL 81%

Large Language Model Safety: A Holistic Survey

Dan Shi, Tianhao Shen, Yufei Huang, Zhigen Li, Yongqi Leng, Renren Jin, Chuang Liu, Xinwei Wu, Zishan Guo, Linhao Yu, Ling Shi, Bojian Jiang, Deyi Xiong

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CL、cs.AI

Comments 158 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00929 2024-12-10 cs.CL cs.AI 81%

A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias

Yuemei Xu, Ling Hu, Jiayi Zhao, Zihan Qiu, Kexin XU, Yuqi Ye, Hanwen Gu

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments The article has been accepted by Frontiers of Computer Science (FCS), with the DOI: {10.1007/s11704-024-40579-4}

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19749 2024-10-29 cs.CY cs.AI 81%

Using AI Alignment Theory to understand the potential pitfalls of regulatory frameworks

Alejandro Tlaie

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14096 2024-09-20 cs.CL cs.AI 81%

Cultural Bias and Cultural Alignment of Large Language Models

Yan Tao, Olga Viberg, Ryan S. Baker, Rene F. Kizilcec

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

Journal ref PNAS Nexus, Volume 3, Issue 9, September 2024, pgae346

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01431 2024-08-15 cs.CY cs.AI 81%

Building an Ethical and Trustworthy Biomedical AI Ecosystem for the Translational and Clinical Integration of Foundational Models

Simha Sankar Baradwaj, Destiny Gilliland, Jack Rincon, Henning Hermjakob, Yu Yan, Irsyad Adam, Gwyneth Lemaster, Dean Wang, Karol Watson, Alex Bui, Wei Wang, Peipei Ping

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10567 2024-06-18 cs.CL cs.AI 81%

InSaAF: Incorporating Safety through Accuracy and Fairness | Are LLMs ready for the Indian Legal Domain?

Yogesh Tripathi, Raghav Donakanti, Sahil Girhepuje, Ishan Kavathekar, Bhaskara Hanuma Vedula, Gokul S Krishnan, Shreya Goyal, Anmol Goel, Balaraman Ravindran, Ponnurangam Kumaraguru

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15479 2024-03-27 cs.CY cs.AI 81%

Antisocial Analagous Behavior, Alignment and Human Impact of Google AI Systems: Evaluating through the lens of modified Antisocial Behavior Criteria by Human Interaction, Independent LLM Analysis, and AI Self-Reflection

Alan D. Ogilvie

专题命中 AI治理与伦理 :alignment(title);safety(abstract);分类 cs.AI、cs.CY

Comments 48 pages including addendum of transcripts

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10934 2023-11-28 cs.AI cs.CY cs.HC 81%

Case Repositories: Towards Case-Based Reasoning for AI Alignment

K. J. Kevin Feng, Quan Ze Chen, Inyoung Cheong, King Xia, Amy X. Zhang

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments MP2 workshop @ NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08943 2023-11-16 cs.CY cs.AI cs.SY eess.SY 81%

Safety, Trust, and Ethics Considerations for Human-AI Teaming in Aerospace Control

Kerianne L. Hobbs, Bernard Li

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08487 2023-11-16 cs.CL cs.AI 81%

Alignment is not sufficient to prevent large language models from generating harmful information: A psychoanalytic perspective

Zi Yin, Wei Ding, Jia Liu

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03718 2023-11-09 cs.CY cs.AI 81%

Frontier AI Regulation: Managing Emerging Risks to Public Safety

Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O'Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager, Kevin Wolf

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

Comments Update July 11th: - Added missing footnote back in. - Adjusted author order (mistakenly non-alphabetical among the first 6 authors) and adjusted affiliations (Jess Whittlestone's affiliation was mistagged and Gillian Hadfield had SRI added to her affiliations) Updated September 4th: Various typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.09250 2023-10-16 cs.LG cs.AI stat.ML 81%

It's an Alignment, Not a Trade-off: Revisiting Bias and Variance in Deep Models

Lin Chen, Michal Lukasik, Wittawat Jitkrittum, Chong You, Sanjiv Kumar

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.11208 2023-03-15 cs.CY cs.AI 81%

Artificial Intelligence Ethics and Safety: practical tools for creating "good" models

Nicholas Kluge Corrêa

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

Comments I would like to withdraw this preprint because it has incorrect and outdated information

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.10566 2022-10-04 cs.LG cs.AI 81%

Scalable Safety-Critical Policy Evaluation with Accelerated Rare Event Sampling

Mengdi Xu, Peide Huang, Fengpei Li, Jiacheng Zhu, Xuewei Qi, Kentaro Oguchi, Zhiyuan Huang, Henry Lam, Ding Zhao

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.13294 2022-09-08 cs.LG cs.AI cs.RO cs.SE cs.SY eess.SY 81%

The missing link: Developing a safety case for perception components in automated driving

Rick Salay, Krzysztof Czarnecki, Hiroshi Kuwajima, Hirotoshi Yasuoka, Toshihiro Nakae, Vahdat Abdelzad, Chengjie Huang, Maximilian Kahn, Van Duong Nguyen

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.00002 2021-05-04 cs.CY cs.AI 81%

Ethics-Based Auditing to Develop Trustworthy AI

Jakob Mokander, Luciano Floridi

专题命中 AI治理与伦理 :trustworthy(title);alignment(abstract);分类 cs.AI、cs.CY

Comments Minds & Machines (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏