arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9331 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9331 篇

2306.08161 2023-06-19 cs.CL cs.AI cs.HC cs.IR cs.LG 67%

h2oGPT: Democratizing Large Language Models

Arno Candel, Jon McKinney, Philipp Singer, Pascal Pfeiffer, Maximilian Jeblick, Prithvi Prabhu, Jeff Gambera, Mark Landry, Shivam Bansal, Ryan Chesler, Chun Ming Lee, Marcos V. Conde, Pasha Stetsenko, Olivier Grellier, SriSatish Ambati

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Work in progress by H2O.ai, Inc

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02622 2023-06-06 cs.LG cs.AI cs.CL 67%

What Makes Entities Similar? A Similarity Flooding Perspective for Multi-sourced Knowledge Graph Embeddings

Zequn Sun, Jiacheng Huang, Xiaozhou Xu, Qijin Chen, Weijun Ren, Wei Hu

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted in the 40th International Conference on Machine Learning (ICML 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08726 2023-06-01 cs.CL cs.AI cs.IR cs.LG 67%

RARR: Researching and Revising What Language Models Say, Using Language Models

Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y. Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, Kelvin Guu

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16504 2023-05-29 cs.CL cs.AI cs.LG 67%

On the Tool Manipulation Capability of Open-source Large Language Models

Qiantong Xu, Fenglu Hong, Bo Li, Changran Hu, Zhengyu Chen, Jian Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14095 2023-04-28 cs.NI 67%

Securing Autonomous Air Traffic Management: Blockchain Networks Driven by Explainable AI

Louise Axon, Dimitrios Panagiotakopoulos, Samuel Ayo, Carolina Sanchez-Hernandez, Yan Zong, Simon Brown, Lei Zhang, Michael Goldsmith, Sadie Creese, Weisi Guo

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

Comments under review in IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.03052 2023-01-10 cs.LG cs.AI cs.CY 67%

AI Maintenance: A Robustness Perspective

Pin-Yu Chen, Payel Das

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted to IEEE Computer Magazine. To be published in 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00355 2023-01-06 cs.CL cs.AI cs.CY 67%

Second Thoughts are Best: Learning to Re-Align With Human Values from Text Edits

Ruibo Liu, Chenyan Jia, Ge Zhang, Ziyu Zhuang, Tony X Liu, Soroush Vosoughi

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments In proceedings of the 36th Conference on Neural Information Processing Systems (NeurIPS 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00544 2022-12-02 cs.RO 67%

Towards Explainability in Modular Autonomous Vehicle Software

Hongrui Zheng, Zirui Zang, Shuo Yang, Rahul Mangharam

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12757 2022-11-24 cs.LG cs.AI cs.CY 67%

FAIRification of MLC data

Ana Kostovska, Jasmin Bogatinovski, Andrej Treven, Sašo Džeroski, Dragi Kocev, Panče Panov

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments This paper was accepted ECML PKDD 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10693 2022-10-20 cs.CL cs.AI cs.LG 67%

Robustness of Demonstration-based Learning Under Limited Data Scenario

Hongxin Zhang, Yanzhe Zhang, Ruiyi Zhang, Diyi Yang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 14 pages, EMNLP 2022 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04995 2022-10-12 cs.LG cs.AI cs.CY 67%

FEAMOE: Fair, Explainable and Adaptive Mixture of Experts

Shubham Sharma, Jette Henderson, Joydeep Ghosh

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.05862 2022-09-21 cs.CY cs.AI cs.LG 67%

X-Risk Analysis for AI Research

Dan Hendrycks, Mantas Mazeika

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.09079 2022-08-22 cs.LG cs.AI cs.CV cs.CY 67%

A Multi-Modal Wildfire Prediction and Personalized Early-Warning System Based on a Novel Machine Learning Framework

Rohan Tan Bhowmik

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08080 2022-08-18 cs.AI cs.CL cs.CV cs.LG cs.MM 67%

Multimodal Lecture Presentations Dataset: Understanding Multimodality in Educational Slides

Dong Won Lee, Chaitanya Ahuja, Paul Pu Liang, Sanika Natu, Louis-Philippe Morency

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.05811 2022-07-14 cs.LG cs.AI cs.CY 67%

Revealing Unfair Models by Mining Interpretable Evidence

Mohit Bajaj, Lingyang Chu, Vittorio Romaniello, Gursimran Singh, Jian Pei, Zirui Zhou, Lanjun Wang, Yong Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.12725 2022-06-28 cs.CV 67%

Empirical Evaluation of Physical Adversarial Patch Attacks Against Overhead Object Detection Models

Gavin S. Hartnett, Li Ang Zhang, Caolionn O'Connell, Andrew J. Lohn, Jair Aguirre

专题命中 安全评测 :safety(abstract);AI safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06490 2022-03-22 cs.CL cs.AI cs.LG 67%

Dict-BERT: Enhancing Language Model Pre-training with Dictionary

Wenhao Yu, Chenguang Zhu, Yuwei Fang, Donghan Yu, Shuohang Wang, Yichong Xu, Michael Zeng, Meng Jiang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ACL 2022 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.02776 2022-02-08 cs.AI cs.CY cs.HC cs.LG 67%

Human rights, democracy, and the rule of law assurance framework for AI systems: A proposal

David Leslie, Christopher Burr, Mhairi Aitken, Michael Katell, Morgan Briggs, Cami Rincon

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 341 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.07754 2021-11-09 cs.AI cs.CY cs.LG stat.ML 67%

Counterfactual Explanations as Interventions in Latent Space

Riccardo Crupi, Alessandro Castelnovo, Daniele Regoli, Beatriz San Miguel Gonzalez

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 34 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.11036 2021-06-22 cs.CY cs.AI cs.LG 67%

Know Your Model (KYM): Increasing Trust in AI and Machine Learning

Mary Roszel, Robert Norvill, Jean Hilger, Radu State

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.09539 2020-12-18 cs.LO 67%

Online Shielding for Stochastic Systems

Bettina Könighofer, Julian Rudolf, Alexander Palmisano, Martin Tappler, Roderick Bloem

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

Comments 18 Pages, 6 Figures, under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.01318 2019-04-03 cs.CV 67%

Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents

Christian Rupprecht, Cyril Ibrahim, Christopher J. Pal

专题命中 安全评测 :safety(abstract);AI safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.08757 2018-10-09 cs.CV 67%

Siamese Generative Adversarial Privatizer for Biometric Data

Witold Oleszkiewicz, Peter Kairouz, Karol Piczak, Ram Rajagopal, Tomasz Trzcinski

专题命中 安全评测 :safety(abstract);AI safety(abstract)

Comments Paper accepted to ACCV 2018 (Asian Conference on Computer Vision)

详情

展开后加载摘要…

URL PDF HTML 收藏
1506.04089 2015-12-18 cs.CL cs.AI cs.LG cs.NE cs.RO 67%

Listen, Attend, and Walk: Neural Mapping of Navigational Instructions to Action Sequences

Hongyuan Mei, Mohit Bansal, Matthew R. Walter

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments To appear at AAAI 2016 (and an extended version of a NIPS 2015 Multimodal Machine Learning workshop paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03910 2026-08-05 cs.AI cs.LG cs.MA 新提交 66%

Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory

基于社会基础的智能体人工智能:通过社会理论协调多元视角

Matt Ratto, Abhishek Moturu, Daniel Silver

专题命中 安全评测 :alignment(abstract,comments);分类 cs.AI、cs.LG

AI总结 本文提出将社会理论用于人工智能系统设计,以解决多元对齐问题,将多元对齐重新定位为基于社会基础的协调问题,勾勒了相关系统的设计空间并指明未来研究方向。

Comments Pluralistic Alignment Workshop @ ICML 2026, Seoul, South Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24259 2026-06-24 cs.CL cs.AI 新提交 66%

SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization

SURGELLM: 通过任务感知特征门控与类别平衡归一化重新思考多任务评估

Noor Islam S. Mohammad, Ulug Bayazit

机构 * Istanbul Technical University(伊斯坦布尔理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI;trustworthy(comments)

AI总结 针对异构NLP任务中归纳偏置不匹配、类别不平衡和缺乏外部词汇知识的问题,提出统一Transformer框架SURGELLM,包含手术特征门控、任务条件前缀令牌和实例加权归一化,在四个任务上平均F1提升0.036。

Comments Proceedings of the 6th Workshop on Trustworthy NLP (TrustNLP 2026), ACL 2026, San Diego, California, USA. Available at https://openreview.net/forum?id=WJCalficPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20668 2026-06-23 cs.CR cs.AI cs.LG 新提交 66%

BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems

BELLS-O:评估LLM监督系统的运营权衡

Leonhard Waibl, Felix Michalak, Hadrien Mariaccia

机构 * University of Graz, Graz, Austria(格拉茨大学) Supervised Program for Alignment Research (SPAR)(对齐研究监督计划 (SPAR)) Centre pour la Sécurité de l'IA (CeSIA), Paris, France(人工智能安全研究中心 (CeSIA),巴黎,法国)

专题命中 安全评测 :jailbreak(abstract);分类 cs.AI、cs.LG;trustworthy(comments)

AI总结 提出首个独立运营基准BELLS-O,评估28个LLM监督系统在检测率、误报率、延迟和成本上的权衡,发现专用护栏在内容审核中占优,而前沿通用模型在越狱检测中表现更好但成本更高。

Comments Accepted at the ICML 2026 Workshop on Trustworthy AI for Good (AI4GOOD). 2 figures; main text plus appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13776 2026-06-10 cs.CY cs.CL cs.CR cs.CV 版本更新 66%

Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking

谁被标记?AI内容水印中的多元评估差距

Alexander Nemecek, Osama Zafar, Yuqiao Xu, Wenbiao Li, Erman Ayday

机构 * Case Western Reserve University(凯斯西储大学)

专题命中 安全评测 :alignment(abstract,comments);分类 cs.CL、cs.CY

AI总结 本文揭示AI内容水印在不同语言、文化和群体间存在系统性偏差,提出跨语言检测一致性、文化多样性覆盖和检测指标人口统计分解三个评估维度,主张水印部署前必须进行公平性审计。

Comments 7 pages. Accepted at the Multimodal Alignment for a Pluralistic Society (MAPS) Workshop, CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20508 2026-06-03 cs.MA cs.AI cs.CL 66%

Measuring Weak-to-Strong Legibility of Reasoning Models

衡量推理模型的弱到强可读性

Dani Roytburg, Shreya Sridhar, Daphne Ippolito

机构 * University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI;trustworthy(comments)

AI总结 针对推理语言模型在多智能体场景中生成的中间思维链,提出“弱到强可读性”概念,并设计衡量指标以评估强模型输出对弱模型的易理解性。

Comments Accepted to Trustworthy AI4GOOD Workshop @ ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04069 2026-03-05 cs.CL cs.AI 66%

Monitoring Emergent Reward Hacking During Generation via Internal Activations

通过内部激活监控生成过程中的涌现奖励黑客行为

Patrick Wilhelm, Thorsten Wittkopp, Odej Kao

机构 * Technical University of Berlin(柏林技术大学) BIFOLD - Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究所)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI;trustworthy(journal_ref)

AI总结 通过分析内部激活来监控生成过程中的奖励黑客行为,提升微调语言模型的安全性。

Journal ref ICLR2026 Workshop: Principled Design for Trustworthy AI

详情

展开后加载摘要…

URL PDF HTML 收藏