arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 684 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 隐私与版权 684 篇

2504.13959 2025-07-14 cs.CY cs.AI cs.CL econ.GN q-fin.EC 89%

AI Safety Should Prioritize the Future of Work

Sanchaita Hazra, Bodhisattwa Prasad Majumder, Tuhin Chakrabarty

机构 * University of Utah Allen Institute for AI Stony Brook University \& Salesforce AI Research

专题命中 隐私与版权 :safety(title,abstract);AI safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10450 2025-06-12 cs.CR cs.AI cs.CL cs.LG 89%

Trustworthy AI: Safety, Bias, and Privacy -- A Survey

Xingli Fang, Jianwei Li, Varun Mulchandani, Jung-Eun Kim

专题命中 隐私与版权 :safety(title,abstract);trustworthy(title);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06778 2025-11-12 cs.CL 88%

SAFENLIDB: A Privacy-Preserving Safety Alignment Framework for LLM-based Natural Language Database Interfaces

Ruiheng Liu, XiaoBing Chen, Jinyu Zhang, Qiongwen Zhang, Yu Zhang, Bailong Yang

专题命中 隐私与版权 :alignment(title,abstract);safety(title);DPO(abstract);分类 cs.CL

Comments AAAI 2026 Extended Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.09006 2023-05-09 cs.SD cs.LG eess.AS 88%

A Review of Speech-centric Trustworthy Machine Learning: Privacy, Safety, and Fairness

Tiantian Feng, Rajat Hebbar, Nicholas Mehlman, Xuan Shi, Aditya Kommineni, and Shrikanth Narayanan

专题命中 隐私与版权 :safety(title,abstract);trustworthy(title,abstract);分类 cs.LG

Journal ref APSIPA Transactions on Signal and Information Processing, vol. 12, no. 3, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26494 2026-08-11 cs.HC cs.AI cs.CY cs.ET 版本更新 87%

Culturally Situated AI Safety for Youth: Saudi Arabian Perspectives of Youth, Parents and Teachers

面向青少年的文化感知生成式AI风险:来自非西方背景的青少年、父母和教师视角

Aljawharah Alzahrani, Tanusree Sharma

机构 * Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 隐私与版权 :safety(title,abstract);AI safety(title);分类 cs.AI、cs.CY

AI总结 研究从非西方视角探讨青少年使用生成式AI工具的风险,分析文化、宗教和社会因素对隐私和安全的影响,揭示家庭经济因素加剧的共享账户使用问题,并提出符合文化规范的家长控制设计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00475 2026-01-21 cs.CY cs.AI 87%

Probabilistic Analysis of Copyright Disputes and Generative AI Safety

版权纠纷的概率分析与生成式AI安全

Hiroaki Chiba-Okabe

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 隐私与版权 :safety(title,abstract);AI safety(title);分类 cs.AI、cs.CY

AI总结 本文通过概率方法分析版权纠纷,并评估生成式AI的版权安全,揭示了NAF条件的局限性。

Comments 5 pages

Journal ref Proc. 20th Int. Conf. on Artificial Intelligence and Law (ICAIL '25), ACM, pp. 470-474 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20957 2026-03-31 cs.CL cs.AI cs.CY 87%

Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models

对齐的打砖块:微调激活大型语言模型对版权书籍的原文回忆

Xinyue Liu, Niloofar Mireshghallah, Jane C. Ginsburg, Tuhin Chakrabarty

机构 * Stony Brook University(石溪大学) Carnegie Mellon University(卡内基梅隆大学) Columbia Law School(哥伦比亚法学院)

专题命中 隐私与版权 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究显示微调可绕过安全对齐策略,使GPT-4o等模型能回忆85-90%的版权书籍,揭示模型权重存储版权作品的漏洞。

Comments Preprint Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05739 2026-01-12 cs.AI cs.CL cs.CR cs.CV 86%

PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility

PII-VisBench: 在可见性连续体上评估视觉语言模型中个人可识别信息安全性

G M Shahariar, Zabir Al Nazi, Md Olid Hasan Bhuiyan, Zhouxing Shi

机构 * University of California, Riverside(加州大学河滨分校)

专题命中 隐私与版权 :safety(title,abstract);alignment(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

AI总结 PII-VisBench 评估视觉语言模型在可见性连续体上对个人可识别信息安全性的保护,通过4000个探测器分析不同可见性水平下的隐私泄露情况。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16606 2026-04-21 cs.CR cs.LG 85%

SafeLM: Unified Privacy-Aware Optimization for Trustworthy Federated Large Language Models

SafeLM:面向可信联邦大语言模型的统一隐私意识优化

Noor Islam S. Mohammad, Uluğ Bayazıt

机构 * Istanbul Technical University(伊斯坦布尔技术大学)

专题命中 隐私与版权 :trustworthy(title,abstract);alignment(abstract);safety(abstract);分类 cs.LG

AI总结 SafeLM通过整合隐私、安全、虚假信息和对抗鲁棒性四大支柱,提升联邦大语言模型的可信度,实现98%有害内容检测准确率,降低通信开销并提升隐私保护效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14312 2025-10-17 cs.AI cs.CL cs.CR 84%

Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies

Mason Nakamura, Abhinav Kumar, Saaduddin Mahmud, Sahar Abdelnabi, Shlomo Zilberstein, Eugene Bagdasarian

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) ELLIS Institute(ELLIS研究所) MPI for Intelligent Systems(智能系统研究所) Tübingen AI Center(图宾根人工智能中心)

专题命中 隐私与版权 :safety(title,abstract);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01311 2026-08-04 cs.CL 新提交 83%

RH-RAG: Trustworthy Long-Form Generation for Privacy-Constrained Settings

RH-RAG:适用于隐私受限场景的可信长文本生成

Raj Shekhar Singh

机构 * Indian Institute of Technology, Roorkee(鲁尔基印度理工学院)

专题命中 隐私与版权 :trustworthy(title,abstract);alignment(abstract);分类 cs.CL

AI总结 该研究针对隐私受限场景下的长文本生成难题,提出基于本地语言模型的多智能体框架RH-RAG,通过三阶段协同生成与双层检索索引提升生成质量,且兼顾数据隐私。

Comments accepted in KDD 2026 SeT-LLM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13933 2026-04-17 cs.CL 83%

OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset

OmniCompliance-100K:一个多领域、基于规则、现实世界安全合规数据集

Wenbin Hu, Huihao Jing, Haochen Shi, Changxuan Fan, Haoran Li, Yangqiu Song

机构 * Hong Kong University of Science and Technology(香港科技大学)

专题命中 隐私与版权 :safety(title,abstract);alignment(abstract);分类 cs.CL

AI总结 本文构建了一个涵盖74项法规的多领域安全合规数据集,包含12,985条规则和106,009个现实案例,通过基准测试评估了不同规模LLM的安全合规能力。

Comments Accepted to ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12308 2026-04-15 cs.CL 83%

ContextLens: Modeling Imperfect Privacy and Safety Context for Legal Compliance

ContextLens:建模不完美隐私和安全上下文以满足法律合规

Haoran Li, Yulin Chen, Huihao Jing, Wenbin Hu, Tsz Ho Li, Chanhou Lou, Hong Ting Tsang, Sirui Han, Yangqiu Song

机构 * Beihang University(北京航空航天大学) HKUST(香港科技大学) National University of Singapore(新加坡国立大学) Faculty of Law, University of Macau(澳门大学法学院)

专题命中 隐私与版权 :safety(title,abstract);AI safety(abstract);分类 cs.CL

AI总结 本文提出ContextLens框架,利用LLM建模法律领域上下文,识别已知和未知因素以提升合规评估,实验表明其能有效提升LLM合规性评估性能。

Comments Accepted by ACL 26

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.10209 2024-03-06 cs.LG cs.CR 83%

On the Alignment of Group Fairness with Attribute Privacy

Jan Aalmoes, Vasisht Duddu, Antoine Boutet

专题命中 隐私与版权 :alignment(title,abstract);trustworthy(abstract);分类 cs.LG

Comments arXiv admin note: text overlap with arXiv:2202.02242

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03546 2026-06-30 cs.CL cs.AI cs.HC cs.LG 82%

Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict

大语言模型在隐私与亲社会冲突下的价值-行动对齐

Guanyu Chen, Chenxiao Yu, Xiyang Hu

机构 * Arizona State University(亚利桑那州立大学) University of Southern California(南加州大学)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究探讨大语言模型在隐私与亲社会冲突中的价值-行动对齐,通过多组结构方程模型分析隐私关注与亲社会性对数据共享的影响,提出价值-行动对齐率(VAAR)作为评估指标。

Comments Findings of the Association for Computational Linguistics: ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24819 2026-06-24 cs.CR 新提交 82%

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice

HelpBench:评估LLM提供隐私、安全和安全建议的能力

Sarah Meiklejohn, Sunny Consolvo, Patrick Gage Kelley, Tara Matthews, Sai Teja Peddinti, Renee Shelby, Lenin Simicich, Kurt Thomas

专题命中 隐私与版权 :safety(title,abstract);trustworthy(abstract)

AI总结 提出HelpBench基准,通过450个真实用户问题评估18个LLM的隐私、安全建议准确性,发现平均得分82%但10%回答低于65%存在有害建议。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16128 2026-04-20 cs.CR 82%

PolicyGapper: Automated Detection of Inconsistencies Between Google Play Data Safety Sections and Privacy Policies Using LLMs

PolicyGapper:利用LLMs自动检测Google Play数据安全部分与隐私政策之间的不一致之处

Luca Ferrari, Billel Habbati, Meriem Guerar, Mariano Ceccato, Luca Verderame

专题命中 隐私与版权 :safety(title,abstract);alignment(abstract)

AI总结 本文提出PolicyGapper,一种基于LLM的方法,用于自动检测Google Play数据安全部分与隐私政策之间的不一致之处,通过四个阶段处理330款应用,发现2689处遗漏披露,验证精度达75%。

Comments Submitted for consideration to the Journal of Information Security and Applications (JISA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18444 2026-03-18 cs.CV 82%

SineProject: Machine Unlearning for Stable Vision Language Alignment

SineProject:用于稳定视觉语言对齐的机器反遗忘

Arpit Garg, Hemanth Saratchandran, Simon Lucey

机构 * Australian Institute for Machine Learning(澳大利亚机器学习研究所)

专题命中 隐私与版权 :alignment(title,abstract);safety(abstract)

AI总结 SineProject通过在冻结的投影器中加入正弦调制的可训练参数,提升Jacobian的谱条件数,稳定反遗忘过程中的视觉语言对齐,减少无害查询拒绝并实现目标信息遗忘。

Comments Accepted at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16667 2024-08-30 cs.LG cs.AI cs.CL cs.MA 82%

Iterative Graph Alignment

Fangyuan Yu, Hardeep Singh Arora, Matt Johnson

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24411 2026-07-28 cs.AI cs.CL cs.CV cs.HC 版本更新 81%

OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows

OS-Sentinel: 通过现实工作流中的混合验证实现安全增强的移动GUI代理

Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong

机构 * The University of Hong Kong(香港大学) Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) Nanyang Technological University(南洋理工大学) Nanjing University(南京大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 隐私与版权 :safety(title,abstract);分类 cs.CL、cs.AI

AI总结 OS-Sentinel通过结合形式验证器和VLM上下文判断者,提升移动GUI代理的安全性,实验显示在多个指标上优于现有方法。

Comments ACL 2026 (Oral) & Best Paper at AIWILD @ ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21710 2026-06-23 cs.CL cs.AI cs.IR 新提交 81%

PrivacyAlign: Contextual Privacy Alignment for LLM Agents

PrivacyAlign: LLM代理的上下文隐私对齐

Manveer Singh Tamber, Abhay Puri, Marc-Etienne Brunet, Perouz Taslakian, Jimmy Lin, Spandana Gella

机构 * University of Waterloo(滑铁卢大学) ServiceNow AI Research(ServiceNow AI 研究) McGill University(麦吉尔大学) Mila -- Quebec AI Institute(米拉——魁北克人工智能研究所)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 提出PrivacyAlign数据集,通过人类标注对齐LLM代理的隐私决策,并引入基于标注的奖励建模以提升隐私对齐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22373 2026-05-25 cs.LG cs.CL 81%

Boundary-targeted Membership Inference Attacks on Safety Classifiers

针对安全分类器的边界目标成员推断攻击

Anthony Hughes, Alexander Goldberg, Prince Jha, Adam Perer, Nikolaos Aletras, Niloofar Mireshghallah

机构 * University of Sheffield(谢菲尔德大学) Carnegie Mellon University(卡内基梅隆大学) MBZUAI

专题命中 隐私与版权 :safety(title,abstract);分类 cs.CL、cs.LG

AI总结 提出一种边界目标选择策略,通过识别低置信度样本来放大成员推断信号,在安全分类器上以5%假阳性率恢复19%的敏感对话,比现有方法高3.5倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19219 2026-05-19 cs.CR cs.AI cs.DC cs.LG 81%

Sherpa.ai Privacy-Preserving Multi-Party Entity Alignment without Intersection Disclosure for Noisy Identifiers

Sherpa.ai 保护隐私的多方实体对齐无需披露交集

Daniel M. Jimenez-Gutierrez, Dario Pighin, Enrique Zuazua, Georgios Kellaris, Joaquin Del Rio, Oleksii Sliusarenko, Xabi Uribe-Etxebarria

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出Sherpa.ai多方PSU协议,用于垂直联邦学习中的隐私保护实体对齐,实现精确和噪声匹配,同时隐藏交集成员信息,适用于多机构医疗疾病检测等场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01699 2026-05-08 cs.LG cs.AI cs.CR cs.NE 81%

Probe-Geometry Alignment: Erasing the Cross-Sequence Memorization Signature Below Chance

探测器几何对齐:在偶然机会以下消除跨序列记忆签名

Anamika Paul Rupa, Anietie Andy

机构 * Howard University(霍华德大学)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了大语言模型中记忆痕迹的消除方法,提出通过探测器几何对齐技术在不损害能力的前提下,有效消除跨序列记忆签名,且在多种模型上验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11253 2026-03-16 cs.SI cs.CL cs.CY 81%

LLMs Can Infer Political Alignment from Online Conversations

大语言模型能从在线对话中推断政治倾向

Byunghwee Lee, Sangyeon Kim, Filippo Menczer, Yong-Yeol Ahn, Haewoon Kwak, Jisun An

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.CL、cs.CY

AI总结 研究展示大语言模型能可靠推断隐藏的政治倾向,优于传统机器学习模型,通过聚合文本推断和使用相关领域提升预测精度,揭示LLM在利用社会文化关联方面的潜力与风险。

Comments 56 pages; 4 figures in the main text and 18 supplementary figures, 11 supplementary tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06049 2026-01-13 cs.CY cs.AI 81%

The Violation State: Safety State Persistence in a Multimodal Language Model Interface

违规状态:多模态语言模型接口中的安全状态持续

Bentley DeVilling

专题命中 隐私与版权 :safety(title,abstract);分类 cs.AI、cs.CY

AI总结 研究发现多模态语言模型在面对版权拒绝后,会持续拒绝无关图像生成请求,揭示了会话级安全状态的持续性问题。

Comments 19 pages, 1 figure, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14698 2025-10-17 cs.LG cs.AI 81%

FedPPA: Progressive Parameter Alignment for Personalized Federated Learning

Maulidi Adi Prasetia, Muhamad Risqi U. Saputra, Guntur Dharma Putra

机构 * Universitas Gadjah Mada, Indonesia(加雅玛大学) Monash University, Indonesia(墨尔本大学)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 8 pages, TrustCom 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18532 2025-03-21 cs.CL cs.LG 81%

Differentially Private Steering for Large Language Model Alignment

Anmol Goel, Yaxi Hu, Iryna Gurevych, Amartya Sanyal

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments ICLR 2025 Camera Ready; Code: https://github.com/UKPLab/iclr2025-psa

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19113 2024-12-02 cs.CL cs.IR 80%

Integration of Contextual Descriptors in Ontology Alignment for Enrichment of Semantic Correspondence

Eduard Manziuk, Oleksander Barmak, Pavlo Radiuk, Vladislav Kuznetsov, Iurii Krak, Sergiy Yakovlev

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.CL

Comments Ontology alignment, contextual descriptors, semantic matching, knowledge representation, essential descriptors, ontology integration, hierarchical structure, semantic heterogeneity, ethical AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10804 2024-07-16 cs.CL 80%

Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format Alignment

Jinhao Jiang, Junyi Li, Wayne Xin Zhao, Yang Song, Tao Zhang, Ji-Rong Wen

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.CL

Comments LLM, CPT, knowledge learning, format alignment; work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏