arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 684 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 隐私与版权 684 篇

2601.21369 2026-08-04 cs.LG 版本更新 79%

Rethinking Federated Graph Foundation Models: A Graph-Language Alignment-based Approach

重新思考联邦图基础模型:基于图-语言对齐的方法

Yinlin Zhu, Di Wu, Xianzhi Zhang, Yuming Ai, Xunkai Li, Miao Hu, Guocong Quan

机构 * Sun Yat-sen University, Guangzhou, China(中山大学) Beijing Institute of Technology, Beijing, China(北京理工大学)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.LG

AI总结 FedGALA通过图-语言对齐解决联邦图基础模型中的语义-结构正交性问题,提升多任务适应性与效率

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15687 2026-07-20 cs.LG 新提交 79%

Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework

迈向联邦多模态图基础模型:一种拓扑感知多模态对齐框架

Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.LG

AI总结 研究针对多模态属性图分散在不同隐私受限孤岛的问题,提出FedGAMMA框架,将联邦多模态图基础学习视为两阶段语义结构对齐问题,经实验验证该框架在下游任务中表现出色,超越诸多基线,在少样本学习场景下也有优势。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25533 2026-06-25 cs.CR cs.CL 新提交 79%

Security and Privacy in Retrieval-Augmented Generation: Architectures, Threats, Defenses, and Future Directions for Building Trustworthy Systems

检索增强生成中的安全与隐私:构建可信系统的架构、威胁、防御与未来方向

Balamurugan Palanisamy, G S S Chalapathi, Vikas Hassija, Rajkumar Buyya

机构 * Department of Electrical and Electronics Engineering(电子与电气工程系) Birla Institute of Technology and Science, Pilani, Pilani Campus(比拉理工学院和科学学院,比拉校区) Department of Computer Engineering, KIIT University(计算机工程系,KIIT大学) Quantum Cloud Computing and Distributed Systems (qCLOUDS) Laboratory(量子云计算与分布式系统(qCLOUDS)实验室) Department of Computing and Information Systems(计算与信息系统系) The University of Melbourne(墨尔本大学)

专题命中 隐私与版权 :trustworthy(title,abstract);分类 cs.CL

AI总结 本文综述了检索增强生成系统在集中式、设备端、联邦和混合范式下的隐私与安全挑战,提出了涵盖检索、上下文构建和生成阶段的统一威胁分类,并分析了攻击类别及防御策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30289 2026-05-29 cs.LG stat.AP stat.ML 79%

Statistical Embeddings for Similarity, Retrieval, and Interpretable Alignment of Numeric Tabular Datasets

用于数值表格数据集的相似性、检索和可解释对齐的统计嵌入

M. Ross Kunz, John Merickel, Keith Wilson

机构 * Idaho National Laboratory(爱达荷国家实验室)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.LG

AI总结 提出一种通过结构化探索性数据分析描述符、句子变换器嵌入和典型相关分析(CCA)来表征和比较数值表格数据集的方法,实现跨数据集的相似性检索和可解释变量级对齐,并支持差分隐私。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28975 2026-05-29 cs.LG 79%

A Training-Time Diagnostic for Generalization via the Log-Alignment Ratio

基于对数对齐比率的训练时泛化诊断

Ali Shehper, Ashish Vaswani

机构 * Essential AI

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.LG

AI总结 提出对数对齐比率(LAR)作为参数-激活对齐的度量,通过捕捉训练中权重谱与激活谱的扩散来跟踪记忆与泛化的转换,并在grokking和语言模型预训练中预测泛化差距。

Comments 32 pages, 25 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18908 2026-05-26 cs.AI 79%

Characterizing Linear Alignment Across Language Models

表征语言模型间的线性对齐

Matt Gorbett, Suman Jana

机构 * Independent Researcher(独立研究者) Department of Computer Science(计算机科学系) Columbia University(哥伦比亚大学)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.AI

AI总结 研究独立训练的大语言模型间是否存在线性对齐,并探索其在文本生成、嵌入分类、分布外检测及隐私保护跨孤岛推理中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17341 2026-05-19 cs.CV cs.AI 79%

Single-Sample Black-Box Membership Inference Attack against Vision-Language Models via Cross-modal Semantic Alignment

通过跨模态语义对齐实现面向视觉-语言模型的单样本黑盒成员推断攻击

Jiaqing Li, Yajuan Lu, Xiaochuan Shi, Gang Wu, ZhongYuan Wang, Chao Liang

机构 * Wuhan University(武汉大学) Tarim University(塔里木大学)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出了一种基于跨模态语义对齐的新型成员推断攻击框架,针对视觉-语言模型在单样本和黑盒场景下的数据安全风险进行评估,通过量化联合嵌入空间中的对齐程度,显著提升了攻击性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14966 2026-04-09 cs.CY cs.CR 79%

Towards trustworthy management of AIGC copyright: blockchain-enabled full lifecycle recording and multi-party auditing approach

面向可信的AIGC版权管理:基于区块链的全生命周期记录与多方审计方法

Jiajia Jiang, Moting Su, Fengshu Li, Xiangli Xiao, Yushu Zhang

专题命中 隐私与版权 :trustworthy(title,abstract);分类 cs.CY

AI总结 本文提出AIGC-Chain系统,通过区块链实现AIGC全生命周期数据记录与多方审计,解决现有解决方案在中间数据管理上的不足,提升版权确认与多方收益分配的可信度。

Journal ref Cybersecurity 9, 151 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19894 2026-04-08 cs.SE cs.AI 79%

TransAgent: Enhancing LLM-Based Code Translation via Fine-Grained Execution Alignment

TransAgent:通过细粒度执行对齐增强基于LLM的代码翻译

Zhiqiang Yuan, Weitong Chen, Hanlin Wang, Xin Peng, Zhenpeng Chen, Yiling Lou

机构 * Fudan University(复旦大学) Tsinghua University(清华大学)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.AI

AI总结 TransAgent通过细粒度执行对齐本地化错误代码块,提升LLM代码翻译的准确性与修复性能,优于现有方法33.3%和56.7%。

Comments Accepted by FSE'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29777 2026-04-01 cs.CV cs.AI 79%

From Skeletons to Semantics: Design and Deployment of a Hybrid Edge-Based Action Detection System for Public Safety

从骨架到语义:一种混合边缘基于的动作检测系统在公共安全中的设计与部署

Ganen Sethupathy, Lalit Dumka, Jan Schagen

机构 * Sopra Steria Germany(Sopra Steria德国) Sopra Steria India(Sopra Steria印度)

专题命中 隐私与版权 :safety(title,abstract);分类 cs.AI

AI总结 本文提出一种混合边缘动作检测系统,结合骨架运动分析与视觉语言模型,以提升公共安全中的实时视频分析能力,通过系统级对比展示两种方法在边缘计算中的互补性与局限性。

Comments Preprint version of a manuscript currently under review at IEEE Access

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22690 2026-03-25 cs.CV cs.AI 79%

WiFi2Cap: Semantic Action Captioning from Wi-Fi CSI via Limb-Level Semantic Alignment

WiFi2Cap:从Wi-Fi CSI进行语义动作描述的肢体级语义对齐

Tzu-Ti Wei, Chu-Yu Huang, Yu-Chee Tseng, Jen-Jee Chen

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出WiFi2Cap框架,通过Wi-Fi CSI生成动作描述,引入镜像一致性损失解决方向敏感问题,并在基准数据集上验证了其在隐私保护语义感知中的有效性。

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16918 2026-03-19 cs.HC cs.AI 79%

Privacy and Safety Experiences and Concerns of U.S. Women Using Generative AI for Seeking Sexual and Reproductive Health Information

美国女性使用生成式AI寻求性与生殖健康信息的隐私和安全经验与担忧

Ina Kaleva, Xiao Zhan, Ruba Abu-Salma, Jose Such

机构 * King's College London London United Kingdom VRAIN, Universitat Politècnica de València \& University of Cambridge Valencia Spain Cambridge United Kingdom King's College London VRAIN, Universitat Politècnica de València \& University of Cambridge

专题命中 隐私与版权 :safety(title,abstract);分类 cs.AI

AI总结 研究探讨了美国女性使用生成式AI获取性与生殖健康信息时的隐私和安全问题,发现用户面临数据收集、政府监控等风险,提出针对性设计和政策建议。

Comments 21 pages, 2 tables, CHI conference on Human Factors in Computing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05954 2026-03-17 cs.LG 79%

A Survey on Deep Learning Approaches for Tabular Data Generation: Utility, Alignment, Fidelity, Privacy, Diversity, and Beyond

关于表格数据生成的深度学习方法的综述:效用、对齐、保真度、隐私、多样性及其他

Mihaela Cătălina Stoian, Eleonora Giunchiglia, Thomas Lukasiewicz

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.LG

AI总结 本文综述了表格数据生成的深度学习方法,从效用、对齐、保真度、隐私、多样性等方面探讨不同需求下的生成方法,并讨论评估方法和未来发展方向。

Comments Accepted to Transactions on Machine Learning Research (02/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09214 2026-03-11 cs.AI 79%

PrivPRISM: Automatically Detecting Discrepancies Between Google Play Data Safety Declarations and Developer Privacy Policies

PrivPRISM:自动检测Google Play数据安全声明与开发者隐私政策之间的差异

Bhanuka Silva, Dishanika Denipitiyage, Anirban Mahanti, Aruna Seneviratne, Suranga Seneviratne

机构 * University of Sydney(悉尼大学) University of New South Wales(新南威尔士大学)

专题命中 隐私与版权 :safety(title,abstract);分类 cs.AI

AI总结 PrivPRISM通过对比隐私政策与数据安全声明,自动检测应用数据实践中的不一致,揭示系统性问题并强调自动化监管的必要性。

Comments 21 pages, 18 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00061 2026-03-03 cs.CR cs.LG 79%

The Hidden Costs of Domain Fine-Tuning: Pii-Bearing Data Degrades Safety and Increases Leakage

领域微调的隐性成本:包含个人身份信息的数据会削弱安全性和增加泄露

Jayesh Choudhari, Piyush Kumar Singh

专题命中 隐私与版权 :safety(title,abstract);分类 cs.LG

AI总结 研究揭示领域微调导致安全性和隐私泄露风险增加,包含PII数据加剧了有害合规和信息泄露问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08862 2025-12-10 cs.CR cs.LG 79%

Secure and Privacy-Preserving Federated Learning for Next-Generation Underground Mine Safety

面向下一代地下矿山安全的安全且隐私保护的联邦学习

Mohamed Elmahallawy, Sanjay Madria, Samuel Frimpong

机构 * 2 School of Engineering Applied Science, Washington State University, Richland, WA 99354, USA 3 Computer Science Department, Missouri University of Science 4 Explosive \& Mining Engineering Department, Missouri University of Science

专题命中 隐私与版权 :safety(title,abstract);分类 cs.LG

AI总结 FedMining通过去中心化功能加密和平衡聚合机制,解决地下采矿中隐私保护和模型收敛问题,实现安全高效的联邦学习应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14067 2025-09-22 cs.CV cs.AI 79%

VLA-Mark: A cross modal watermark for large vision-language alignment model

Shuliang Liu, Qi Zheng, Jesse Jiaxi Xu, Yibo Yan, Junyan Zhang, He Geng, Aiwei Liu, Peijie Jiang, Jia Liu, Yik-Cheung Tam, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) University of Toronto(多伦多大学) Ant Group, Alibaba(蚂蚁集团,阿里巴巴) New York University Shanghai(纽约大学上海分校)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.AI

Comments Accepted by the main conference, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14585 2025-09-05 cs.CL 79%

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning

Wenbin Hu, Haoran Li, Huihao Jing, Qi Hu, Ziqian Zeng, Sirui Han, Heli Xu, Tianshu Chu, Peizhao Hu, Yangqiu Song

机构 * HKUST(香港科技大学) South China University of Technology(华南理工大学) Huawei Technologies(华为技术有限公司)

专题命中 隐私与版权 :safety(title,abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12727 2025-08-19 cs.LG 79%

FedSODA: Federated Fine-tuning of LLMs via Similarity Group Pruning and Orchestrated Distillation Alignment

Manning Zhu, Songtao Guo, Pengzhan Zhou, Yansong Ning, Chang Han, Dewen Qiao

机构 * Chongqing University(重庆大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Third Military Medical University(第三军医大学)

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00972 2025-07-16 cs.AI cs.RO 79%

Seeking to Collide: Online Safety-Critical Scenario Generation for Autonomous Driving with Retrieval Augmented Large Language Models

Yuewen Mei, Tong Nie, Jian Sun, Ye Tian

机构 * Dept. of Traf. Eng. Tongji University Shanghai, China Dept. of Civ. \& Envir. Eng. The Hong Kong Polytechnic University Hong Kong SAR, China

专题命中 隐私与版权 :safety(title,abstract);分类 cs.AI

Comments Accepted at IEEE ITSC 2025

Journal ref IEEE International Conference on Intelligent Transportation Systems, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20000 2025-06-26 cs.CR cs.DC cs.LG 79%

Can One Safety Loop Guard Them All? Agentic Guard Rails for Federated Computing

Narasimha Raghavan Veeraragavan, Jan Franz Nygård

机构 * Division of Cancer Registry of Norway, Norwegian Institute of Public Health, Oslo, Norway(挪威癌症登记处,挪威公共卫生研究所,奥斯陆) Department of Physics and Technology, The Arctic University of Norway, Tromsø, Norway(挪威极地大学物理与技术系)

专题命中 隐私与版权 :safety(title,abstract);分类 cs.LG

Comments Accepted at ICML 2025 Workshop on Collaborative and Federated Agentic Workflows (CFAgentic@ICML'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06443 2025-04-29 cs.AI 79%

Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment

Qizhang Feng, Siva Rajesh Kasa, Santhosh Kumar Kasa, Hyokun Yun, Choon Hui Teo, Sravan Babu Bodapati

机构 * Amazon Inc.(亚马逊公司)

专题命中 隐私与版权 :alignment(title);DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07501 2025-03-11 cs.LG 79%

Trustworthy Machine Learning via Memorization and the Granular Long-Tail: A Survey on Interactions, Tradeoffs, and Beyond

Qiongxiu Li, Xiaoyu Luo, Yiyi Chen, Johannes Bjerva

专题命中 隐私与版权 :trustworthy(title,abstract);分类 cs.LG

Comments 28 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06652 2025-02-11 cs.CL 79%

Transparent NLP: Using RAG and LLM Alignment for Privacy Q&A

Anna Leschanowsky, Zahra Kolagar, Erion Çano, Ivan Habernal, Dara Hallinan, Emanuël A. P. Habets, Birgit Popp

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.CL

Comments Submitted to ARR

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07764 2024-02-20 cs.AI cs.NI 79%

When Large Language Model Agents Meet 6G Networks: Perception, Grounding, and Alignment

Minrui Xu, Dusit Niyato, Jiawen Kang, Zehui Xiong, Shiwen Mao, Zhu Han, Dong In Kim, Khaled B. Letaief

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08629 2023-12-15 cs.AI 79%

ChatSOS: LLM-based knowledge Q&A system for safety engineering

Haiyang Tang, Zhenyi Liu, Dongping Chen, Qingzhao Chu

专题命中 隐私与版权 :safety(title,abstract);分类 cs.AI

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06707 2023-12-13 cs.CY 79%

Exploring Public's Perception of Safety and Video Surveillance Technology: A Survey Approach

Babak Rahimi Ardabili, Armin Danesh Pazho, Ghazal Alinezhad Noghre, Vinit Katariya, Gordon Hull, Shannon Reid, Hamed Tabkhi

专题命中 隐私与版权 :safety(title,abstract);分类 cs.CY

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.16735 2023-09-01 cs.CV cs.AI 79%

Post-Deployment Adaptation with Access to Source Data via Federated Learning and Source-Target Remote Gradient Alignment

Felix Wagner, Zeju Li, Pramit Saha, Konstantinos Kamnitsas

专题命中 隐私与版权 :alignment(title,abstract);分类 cs.AI

Comments This version was accepted for the Machine Learning in Medical Imaging (MLMI 2023) workshop at MICCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.09369 2021-08-24 cs.CR cs.CV cs.CY cs.HC 79%

OSRM-CCTV: Open-source CCTV-aware routing and navigation system for privacy, anonymity and safety (Preprint)

Lauri Sintonen, Hannu Turtiainen, Andrei Costin, Timo Hamalainen, Tuomo Lahtinen

专题命中 隐私与版权 :safety(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14442 2026-07-17 cs.CR 新提交 78%

Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale

披露差异:大规模衡量隐私政策与数据安全的不一致性

Mst Eshita Khatun, Lamine Noureddine, Sideeq Bello, Aisha Ali-Gombe

专题命中 隐私与版权 :safety(title,abstract)

AI总结 针对移动应用隐私政策与数据安全标签表述不一致问题,通过对6051款安卓应用大规模实证研究,利用基于大语言模型的框架和统一模式,测量一致性并引入风险评分,发现敏感类别受影响大,凸显披露机制差距,强调加强验证与提高透明度。

Journal ref Proceedings on Privacy Enhancing Technologies (PoPETs) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏