arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-04-24 至 2026-04-24 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 10 篇

2601.06033 2026-04-24 cs.HC cs.CR cs.CY 79%

How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape

生成式人工智能如何在信任与安全领域赋能攻击者和防御者

Patrick Gage Kelley, Steven Rousso-Schindler, Renee Shelby, Kurt Thomas, Allison Woodruff

专题命中 其他安全 :safety(title,abstract);分类 cs.CY

AI总结 本文通过43名专家的质性研究,探讨生成式AI在信任与安全领域对攻击者和防御者的影响,揭示其在提升攻击规模与速度的同时,也为防御方提供了检测有害内容、调查取证等新手段。

Comments 21 pages, 4 tables, 1 figure

Journal ref In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26). Association for Computing Machinery, New York, NY, USA, Article 1316, 1-21

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20938 2026-04-24 cs.LG cs.AI 62%

HARBOR: Automated Harness Optimization

HARBOR:自动化Harness优化

Biswa Sengupta, Jinhua Wang

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出HARBOR框架,通过约束噪声贝叶斯优化解决Harness配置问题,强调自动化配置优于手动堆叠,并在生产编码代理中实例化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21361 2026-04-24 cs.AI 57%

Time, Causality, and Observability Failures in Distributed AI Inference Systems

时间、因果性和可观察性故障在分布式AI推理系统中

Ankur Sharma, Deep Shah, David Lariviere, Hesham ElBakoury

机构 * Open Compute Project(开放计算项目)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 研究揭示了分布式AI系统中时间同步对可观察性和因果性的影响,发现微小时钟偏差会导致因果错误,但系统功能和性能不受影响,强调时间对齐的重要性。

Comments 17 pages, 6 figures. Produced as part of the Unified Intelligent Infrastructure workstream at the Open Compute Project (OCP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21352 2026-04-24 cs.CL 57%

CARE: Counselor-Aligned Response Engine for Online Mental-Health Support

CARE:面向在线心理健康支持的咨询师对齐响应引擎

Hagai Astrin, Ayal Swaid, Avi Segal, Kobi Gal

机构 * Ben-Gurion University(本· Gurion 大学) School of Informatics, University of Edinburgh(爱丁堡大学信息学院) School of Informatics, University of Edinburgh Edinburgh UK(爱丁堡大学信息学院爱丁堡 英国)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 CARE通过针对希伯来语和阿拉伯语的微调,提升心理健康领域低资源语言下的响应质量,增强咨询师与求助者对话的动态结构和情感语境。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21255 2026-04-24 cs.CL 57%

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors

当代理看起来相同:量化蒸馏诱导的工具使用行为相似性

Chenghao Yang, Yuning Zhang, Zhoufutu Wen, Tao Gong, Jiaheng Liu, Qi Chu, Nenghai Yu

机构 * School of Cyber Science and Technology, USTC(USTC计算机科学与技术学院) Anhui Province Key Laboratory of Digital Security(安徽省数字安全重点实验室) M-A-P

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文提出两种新指标RPS和AGS,用于区分任务成功必需行为与模型自主偏好,通过评估18个模型发现AGS能区分教师特定收敛与通用改进。

Comments Accepted by ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12994 2026-04-24 cs.CR cs.AI 57%

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software

LogicEval: 一个系统框架用于评估自动修复技术在现实软件中逻辑漏洞的修复

Syed Md Mukit Rashid, Abdullah Al Ishtiaq, Kai Tu, Yilu Dong, Tianwei Wu, Ali Ranjbar, Tianchang Yang, Najrin Sultana, Shagufta Mehnaz, Syed Rafiul Hussain

机构 * The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 本文提出LogicEval框架,通过LogicDS数据集评估传统和基于LLM的修复方法,发现编译和测试失败主要由提示敏感性、代码上下文丢失和补丁定位困难导致。

Comments To appear in ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11569 2026-04-24 cs.AI 57%

SemaPop: Semantic-Persona Conditioned and Controllable Population Synthesis

SemaPop:基于语义-人设的可控人口合成

Zhenlin Qin, Yancheng Ling, Leizhen Wang, Francisco Câmara Pereira, Zhenliang Ma

机构 * Department of Civil and Architectural Engineering, KTH Royal Institute of Technology(土木与建筑系,皇家理工学院) Digital Future, KTH Royal Institute of Technology(数字未来,皇家理工学院) Department of Data Science and Artificial Intelligence, Monash University(数据科学与人工智能系,墨尔本大学) Department of Technology, Management and Economics, Technical University of Denmark(技术、管理与经济学系,丹麦技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 SemaPop通过引入人设表示作为生成条件,实现可控的人口合成,提升生成性能并保持统计一致性与样本多样性。

Comments Submitted to Transportation Research Part C: Emerging Technologies

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04395 2026-04-24 cs.CV cs.MM 50%

BiTDiff: Fine-Grained 3D Conducting Motion Generation via BiMamba-Transformer Diffusion

BiTDiff:通过BiMamba-Transformer扩散实现细粒度3D导体运动生成

Tianzhi Jia, Kaixing Yang, Xiaole Yang, Xulong Tang, Ke Qiu, Shikui Wei, Yao Zhao

机构 * Institute of Information Science, Beijing Jiaotong University(信息科学学院,北京交通大学) Renmin University of China(中国人民大学)

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出BiTDiff框架,结合BiMamba-Transformer模型和扩散策略,解决3D导体运动生成中的数据和方法限制,构建了首个大规模CM-Data数据集,并实现高质量运动合成与训练自由的联合级运动编辑。

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12845 2026-04-24 cs.CV 50%

Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation

多模态蛋白质语言模型用于酶动力学参数:从底物识别到构象适应

Fei Wang, Xinye Zheng, Kun Li, Yanyan Wei, Yuxin Liu, Ganpeng Hu, Tong Bao, Jingwen Yang

机构 * School of Computer Science and Information Engineering(计算机科学与信息工程学院) Institute of Artificial Intelligence(人工智能研究院) CVLab, College of Information Technology(CV实验室,信息学院) Intelligent Interconnected Systems Laboratory of Anhui Province(安徽省智能互联系统实验室) School of Food Biological Engineering(食品生物工程学院)

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出多阶段多模态条件建模方法,通过ERBA模块在蛋白质语言模型中注入跨模态信息,提升酶动力学参数预测的准确性与生物合理性。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17955 2026-04-24 cs.SE 50%

Mining Type Constructs Using Patterns in AI-Generated Code

在AI生成代码中利用模式挖掘类型构造

Imgyeong Lee, Tayyib Ul Hassan, Abram Hindle

专题命中 其他安全 :safety(abstract)

AI总结 研究AI在类型构造任务中的表现,发现AI更易误用any关键字及高级类型构造,但其PR接受率高于人类。

详情

展开后加载摘要…

URL PDF HTML 收藏