arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7971 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7971 篇

2501.18438 2025-02-03 cs.SE cs.AI 70%

o3-mini vs DeepSeek-R1: Which One is Safer?

Aitor Arrieta, Miriam Ugarte, Pablo Valle, José Antonio Parejo, Sergio Segura

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2501.17749

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16513 2025-01-31 cs.CL 70%

Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models

Sudarshan Kamath Barkur, Sigurd Schacht, Johannes Scholl

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments Corrected Version - Solved Some Issues with reference compilation by latex

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06167 2025-01-08 cs.AI 70%

Predictable Artificial Intelligence

Lexin Zhou, Pablo A. Moreno-Casares, Fernando Martínez-Plumed, John Burden, Ryan Burnell, Lucy Cheke, Cèsar Ferri, Alexandru Marcoci, Behzad Mehrbakhsh, Yael Moros-Daval, Seán Ó hÉigeartaigh, Danaja Rutar, Wout Schellaert, Konstantinos Voudouris, José Hernández-Orallo

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

Comments Paper Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05285 2024-12-03 cs.AI cs.SE 70%

AgentOps: Enabling Observability of LLM Agents

Liming Dong, Qinghua Lu, Liming Zhu

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13944 2024-10-21 cs.CL 70%

Boosting LLM Translation Skills without General Ability Loss via Rationale Distillation

Junhong Wu, Yang Zhao, Yangyifan Xu, Bing Liu, Chengqing Zong

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17287 2024-10-01 cs.CL 70%

When to Trust LLMs: Aligning Confidence with Response Quality

Shuchang Tao, Liuyi Yao, Hanxing Ding, Yuexiang Xie, Qi Cao, Fei Sun, Jinyang Gao, Huawei Shen, Bolin Ding

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments Accepted by ACL 2024. Code: https://github.com/TaoShuchang/CONQORD

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04792 2024-07-12 cs.GT cs.AI 70%

Playing Large Games with Oracles and AI Debate

Xinyi Chen, Angelica Chen, Dean Foster, Elad Hazan

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03820 2024-06-24 cs.CL 70%

CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues

Makesh Narsimhan Sreedhar, Traian Rebedea, Shaona Ghosh, Jiaqi Zeng, Christopher Parisien

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13669 2024-05-29 cs.CL 70%

Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning

Zhaorui Yang, Tianyu Pang, Haozhe Feng, Han Wang, Wei Chen, Minfeng Zhu, Qian Liu

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16967 2024-04-29 cs.LG cs.CR 70%

ML2SC: Deploying Machine Learning Models as Smart Contracts on the Blockchain

Zhikai Li, Steve Vott, Bhaskar Krishnamachar

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10636 2024-04-18 cs.CY cs.AI cs.CL cs.HC cs.LG 70%

What are human values, and how do we align AI to them?

Oliver Klingefjord, Ryan Lowe, Joe Edelman

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13812 2024-03-22 cs.DL cs.AI cs.CL cs.CY cs.LG stat.OT 70%

Quantitative Analysis of AI-Generated Texts in Academic Research: A Study of AI Presence in Arxiv Submissions using AI Detection Tool

Arslan Akram

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 8 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12893 2023-12-29 cs.RO cs.AI cs.CV 70%

A Safer Vision-based Autonomous Planning System for Quadrotor UAVs with Dynamic Obstacle Trajectory Prediction and Its Application with LLMs

Jiageng Zhong, Ming Li, Yinliang Chen, Zihang Wei, Fan Yang, Haoran Shen

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18762 2023-10-31 cs.LG cs.CR 70%

Purify++: Improving Diffusion-Purification with Advanced Diffusion Models and Control of Randomness

Boya Zhang, Weijian Luo, Zhihua Zhang

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01508 2023-10-10 cs.LG cs.CR cs.CV 70%

Circumventing Concept Erasure Methods For Text-to-Image Generative Models

Minh Pham, Kelly O. Marshall, Niv Cohen, Govind Mittal, Chinmay Hegde

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12833 2023-08-25 cs.CL cs.CR 70%

Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities

Maximilian Mozes, Xuanli He, Bennett Kleinberg, Lewis D. Griffin

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10970 2023-08-03 cs.LG 70%

Can GPT-4 Perform Neural Architecture Search?

Mingkai Zheng, Xiu Su, Shan You, Fei Wang, Chen Qian, Chang Xu, Samuel Albanie

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14784 2023-05-25 cs.AI cs.CL cs.CY cs.LG 70%

Anthropomorphization of AI: Opportunities and Risks

Ameet Deshpande, Tanmay Rajpurohit, Karthik Narasimhan, Ashwin Kalyan

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01299 2023-05-03 cs.LG cs.SY eess.SY 70%

An Improved Yaw Control Algorithm for Wind Turbines via Reinforcement Learning

Alban Puech, Jesse Read

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

Journal ref Amini, MR., Canu, S., Fischer, A., Guns, T., Kralj Novak, P., Tsoumakas, G. (eds) Machine Learning and Knowledge Discovery in Databases. ECML PKDD 2022. Lecture Notes in Computer Science(), vol 13717. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10513 2022-11-23 cs.AI 70%

Computable Artificial General Intelligence

Michael Timothy Bennett

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

Comments Experiment code available on TechRxiv: https://www.techrxiv.org/articles/preprint/Computable_Artificial_General_Intelligence/19740190

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.12016 2021-08-30 cs.LG 70%

DeepFlow: Abnormal Traffic Flow Detection Using Siamese Networks

Sepehr Sabour, Sanjeev Rao, Majid Ghaderi

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

Comments 7 pages, 12 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.11872 2021-06-23 cs.LG cs.NE 70%

Randomness In Neural Network Training: Characterizing The Impact of Tooling

Donglin Zhuang, Xingyao Zhang, Shuaiwen Leon Song, Sara Hooker

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

Comments 21 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.10247 2021-01-26 cs.LG 70%

Incorporating Expert Guidance in Epidemic Forecasting

Alexander Rodríguez, Bijaya Adhikari, Naren Ramakrishnan, B. Aditya Prakash

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

Comments Appears in SIGKDD 2020 epiDAMIK

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.05651 2020-07-16 cs.LG stat.ML 70%

Bayesian Variational Autoencoders for Unsupervised Out-of-Distribution Detection

Erik Daxberger, José Miguel Hernández-Lobato

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

Comments 21 pages, extended version with supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.04877 2018-08-29 cs.LG cs.CV cs.HC 70%

Learning via social awareness: Improving a deep generative sketching model with facial feedback

Natasha Jaques, Jennifer McCleary, Jesse Engel, David Ha, Fred Bertsch, Rosalind Picard, Douglas Eck

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.08476 2017-07-27 cs.AI cs.CR 70%

Guidelines for Artificial Intelligence Containment

James Babcock, Janos Kramar, Roman V. Yampolskiy

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06039 2026-07-08 cs.SE 新提交 69%

Automating Quality Assessment with NLP of LLM-Generated Defeaters

利用自然语言处理对大语言模型生成的反驳进行质量评估自动化

Tihomir Rohlinger, Daniel Ratiu, Stefan Wagner

专题命中 其他安全 :safety(abstract,journal_ref);alignment(abstract)

AI总结 研究针对大语言模型生成的反驳质量评估依赖人工且主观的问题,提出结合保证案例图结构特征、语义嵌入和元分类器的自动化评估方法,经案例研究验证,该方法能减少主观差异,为保证案例审查提供决策支持。

Comments 10 pages, 2 figures. Author preprint version of a paper published at ICSRS 2025

Journal ref 2025 9th International Conference on System Reliability and Safety (ICSRS), Turin, Italy, 2025, pp. 101 110

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15794 2026-08-19 cs.LG cs.AI cs.CL 版本更新 67%

Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting

自蒸馏作为大语言模型的性能恢复机制:对抗压缩与灾难性遗忘

Chi Liu, Xin Chen, Xu Zhou, Fangbo Tu, Srinivasan Manoharan

机构 * PayPal AI

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出基于自蒸馏微调的性能恢复框架,通过理论分析和实验验证,证明自蒸馏能有效恢复模型能力,揭示了高维流形对齐与性能恢复之间的强相关性。

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06416 2026-08-19 cs.CL cs.AI cs.LG 版本更新 67%

Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries

注意力流动:通过故事摘要追踪LLM概念性参与

Rebecca M. M. Hicke, Sil Hamilton, David Mimno, Ross Deans Kristensen-McLachlan

机构 * Cornell University(康奈尔大学) Aarhus University(奥胡斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过比较人类与LLM生成的故事摘要,分析模型在文本中的概念性参与模式,发现模型更关注文本结尾,揭示了摘要生成任务的复杂性。

Comments Error found in data creation pipeline

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13925 2026-08-17 cs.LG cs.AI cs.CL 新提交 67%

CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing

CForce:通过一致性强制提升扩散大语言模型(dLLMs)的并行解码性能

Yuji Ren, Chenkai Xu, Zhuocheng Gong, Jianguo Li, Zhijie Deng

机构 * Shanghai Jiao Tong University(上海交通大学) Ant Group(蚂蚁集团)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出CForce方法,通过一致性强制提升dLLMs的并行解码性能,在LLaDA模型上实验证实其在高并行解码预算下可优化速度-质量权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏