arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7971 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7971 篇

2603.00062 2026-03-03 cs.CY 70%

How much technical talent is there? A systematic estimate of the ML research pool among 3 million consultants

有多少技术人才?对300万顾问中机器学习研究池的系统估计

Maximilian Schons, Red Bermejo, Florian Aldehoff-Zeidler, Niccolò Zanichelli, Oliver Evans, Gavin Leech, Samuel Härgestam

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CY

AI总结 研究通过系统方法估计300万顾问中机器学习研究人才数量,发现技术人才数量超过MATS培训计划校友,且通过工作测试的AI模型尚未出现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04398 2026-03-02 cs.CL 70%

Interpreting Transformers Through Attention Head Intervention

通过注意力头干预解读变换器

Mason Kadem, Rong Zheng

机构 * Computing and Software, Faculty of Engineering, McMaster University(工程学院计算机与软件系,麦斯特大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL

AI总结 通过注意力头干预方法,研究变换器的机理可解释性,实现对模型行为的定向控制,提升AI安全性和实用性。

Comments minor citation fix

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20727 2026-02-25 cs.CL 70%

ID-LoRA: Efficient Low-Rank Adaptation Inspired by Matrix Interpolative Decomposition

ID-LoRA: 基于矩阵插值分解的高效低秩适应

Xindian Ma, Rundong Kong, Peng Zhang, Ruoxiang Huang, Yongyu Jiang

机构 * Tianjin University(天津大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

AI总结 ID-LoRA通过基于矩阵插值分解的低秩适应方法,在减少可训练参数的同时保持模型容量,优于现有PEFT技术。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20696 2026-02-25 cs.AI 70%

PromptCD: Test-Time Behavior Enhancement via Polarity-Prompt Contrastive Decoding

PromptCD:通过极性提示对比解码提升测试时行为

Baolong Bi, Yuyao Ge, Shenghua Liu, Yuchen He, Siqian Tong, Lizhe Chen, Lingrui Mei, Zehao Li, Yiwei Wang, Yujun Cai, Ming-Hsuan Yang, Xueqi Cheng

专题命中 其他安全 :alignment(abstract);harmlessness(abstract);分类 cs.AI

AI总结 PromptCD通过极性提示对比解码在测试时提升LLM和VLM的行为一致性,实现无需额外训练的广泛增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18705 2026-02-24 cs.MA cs.AI 70%

EDU-MATRIX: A Society-Centric Generative Cognitive Digital Twin Architecture for Secondary Education

EDU-MATRIX: 一种面向次级教育的以社会为中心的生成认知数字孪生架构

Wenjing Zhai, Jianbin Zhang, Tao Liu

机构 * The High School Affiliated to Beijing Normal University(北京师范大学附属高中) Department of Electronic and Communication Engineering, North China Electric Power University(华北电力大学电子与通信工程系)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 EDU-MATRIX通过社会引力与认知流体的交互,构建了面向次级教育的生成认知数字孪生架构,实现了教育价值观与复杂社交动态的对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00973 2026-02-17 cs.CR cs.AI eess.SP 70%

Keys in the Weights: Transformer Authentication Using Model-Bound Latent Representations

权重中的钥匙:使用模型绑定的潜在表示进行变换器认证

Ayşe S. Okatan, Mustafa İlhan Akbaş, Laxima Niure Kandel, Berker Peköz

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本文提出MoBLE,一种基于模型绑定的潜在表示变换器认证方法,通过零次解码不可转移性实现身份验证,适用于安全关键领域。

Comments Cite as A. S. Okatan, M. I. Akbas, L. N. Kandel, and B. Pekoz, "Keys in the weights: Transformer authentication using model-bound latent representations," in Proc. 2025 Cyber Awareness and Research Symp. (IEEE CARS 2025), Grand Forks, ND, Oct. 2025, pp. 6

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05228 2026-02-12 cs.AI 70%

Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention Sink

手术:通过注意力sink缓解大型语言模型的有害微调

Guozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo, Qi Mu, Xiumin Wang, Li Shen

机构 * South China University of Technology(南方科技大学) Pengcheng Laboratory(鹏城实验室) Sun Yat-sen University(中山大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 Surgery通过注意力sink机制在微调阶段抑制有害模式学习,提升模型安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08630 2026-02-10 cs.AI cs.CC 70%

Debate is efficient with your time

与时间高效地进行辩论

Jonah Brown-Cohen, Geoffrey Irving, Simon C. Marshall, Ilan Newman, Georgios Piliouras, Mario Szegedy

机构 * Google DeepMind(谷歌DeepMind) UK AI Security Institute(英国人工智能安全研究所) University of Haifa(海法大学) Rutgers University(罗格斯大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 通过辩论实现人工智能安全,证明辩论查询复杂度与电路复杂性密切相关。

Comments 11 Pages, 0 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01618 2026-02-03 cs.CL 70%

SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia

SEA-Guard:面向东南亚的文化根基多语言安全防护

Panuthep Tasawong, Jian Gang Ngui, Alham Fikri Aji, Trevor Cohn, Peerat Limkonchotiwat

机构 * VISTEC Google(谷歌) AI Singapore(AI新加坡)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

AI总结 SEA-Guard是首个基于东南亚文化背景的多语言安全模型,通过代理数据生成框架构建区域特定安全数据集,有效检测地区敏感内容并保持良好的通用安全性能。

Comments Under reivew

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12812 2026-01-21 cs.CL 70%

Do Clinical Question Answering Systems Really Need Specialised Medical Fine Tuning?

临床问答系统真的需要专门的医学微调吗?

Sushant Kumar Ray, Gautam Siddharth Kashyap, Sahil Tripathi, Nipun Joshi, Vijay Govindarajan, Rafiq Ali, Jiechao Gao, Usman Naseem

机构 * University of Delhi(德里大学) Jamia Hamdard(贾迈亚哈马尔德大学) Cornell University(康奈尔大学) Expedia Group(Expedia集团) DSEU-Okhla Center for SDGC(SDGC中心) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

AI总结 MEDASSESS-X通过推理时对齐技术,无需领域微调即可提升临床问答系统性能,解决专门化谬误问题。

Comments Accepted at EACL 2026 (Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04170 2026-01-08 cs.AI 70%

Agent Drift: Quantifying Behavioral Degradation in Multi-Agent LLM Systems Over Extended Interactions

智能体漂移:多智能体大语言模型系统在长时间交互中的行为退化量化

Abhishek Rath

机构 * Independent Researcher(独立研究者)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本研究提出智能体漂移概念,通过理论框架和量化指标,探讨多智能体系统在长时间交互中的行为退化问题,并提出缓解策略以提升系统稳定性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03047 2026-01-07 cs.LG 70%

When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability

当咖啡特征在棺材上激活:对特征提取和转向用于机制可解释性的分析

Raphael Ronge, Markus Maier, Frederick Eberhardt

机构 * Department of Philosophy of Nature and Technology(自然哲学与技术系) Munich School of Philosophy(慕尼黑哲学学院) Division of the Humanities and Social Sciences(人文与社会科学系) California Institute of Technology(加州理工学院)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

AI总结 本文分析了通过稀疏自编码器提取特征和控制模型输出的方法,指出其在机制可解释性中的局限性和可靠性问题,强调需转向更可靠的预测与控制。

Comments 33 pages (65 with appendix), 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14500 2025-11-04 cs.AI cs.MA cs.NE 70%

The Digital Ecosystem of Beliefs: does evolution favour AI over humans?

David M. Bossens, Shanshan Feng, Yew-Soon Ong

机构 * Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR) Centre for Frontier AI Research (CFAR), Agency for Science, Technology and Research (A*STAR)(高性能计算研究所(IHPC)、科技研究局(A*STAR)前沿人工智能研究中心(CFAR)、科技研究局(A*STAR)) School of Computer Science Wuhan University(计算机科学学院 武汉大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11040 2025-10-14 cs.CL 70%

Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks

Wenya Xie, Qingying Xiao, Yu Zheng, Xidong Wang, Junying Chen, Ke Ji, Anningzhe Gao, Prayag Tiwari, Xiang Wan, Feng Jiang, Benyou Wang

机构 * Shenzhen Research Institute of Big Data(大数据研究 institute) National Health Data Institute(国家健康数据研究所) Halmstad University(哈马碧大学) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01363 2025-10-03 cs.AI 70%

Retrieval-Augmented Framework for LLM-Based Clinical Decision Support

Leon Garza, Anantaa Kotal, Michael A. Grasso, Emre Umucu

机构 * Dept. of Computer Science, The University of Texas at El Paso, USA(计算机科学系,德克萨斯大学埃尔帕索分校) Dept. of Emergency Medicine, University of Maryland School of Medicine, USA(急诊医学系,马里兰大学医学院) Dept. of Public Health Sciences, The University of Texas at El Paso, USA(公共卫生科学系,德克萨斯大学埃尔帕索分校)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00300 2025-10-02 cs.AI 70%

ICL Optimized Fragility

Serena Gomez Wannaz

机构 * Serena Gomez Wannaz(独立研究者)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16660 2025-09-23 cs.CL 70%

Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation

Zuhair Hasan Shaik, Abdullah Mazhar, Aseem Srivastava, Md Shad Akhtar

机构 * IIIT Dharwad, India(印度IIIT达尔瓦德大学) IIIT Delhi, India(印度IIIT德里大学) FLaME-NLP Lab, IIIT Delhi(IIIT德里大学FLaME-NLP实验室)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL

Comments Accepted to the NeurIPS 2025 Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19512 2025-08-29 cs.CL 70%

Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging

Hua Farn, Hsuan Su, Shachi H Kumar, Saurav Sahay, Shang-Tse Chen, Hung-yi Lee

机构 * National Taiwan University(国立台湾大学) Intel Lab(英特尔实验室)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15796 2025-08-25 cs.CL cs.AI cs.CY cs.LG 70%

Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases

Nouar AlDahoul, Yasir Zaki

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01059 2025-08-05 cs.CR cs.AI 70%

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report

Sajana Weerawardhena, Paul Kassianik, Blaine Nelson, Baturay Saglam, Anu Vellore, Aman Priyanshu, Supriti Vijay, Massimo Aufiero, Arthur Goldblatt, Fraser Burch, Ed Li, Jianliang He, Dhruv Kedia, Kojin Oshiba, Zhouran Yang, Yaron Singer, Amin Karbasi

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

Comments 34 pages - Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20028 2025-07-29 cs.CV cs.AI 70%

TAPS : Frustratingly Simple Test Time Active Learning for VLMs

Dhruv Sarkar, Aprameyo Chakrabartty, Bibhudatta Bhanja

机构 * IIT Kharagpur(印度理工学院Kharagpur分校)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21892 2025-06-30 cs.CV cs.AI 70%

SODA: Out-of-Distribution Detection in Domain-Shifted Point Clouds via Neighborhood Propagation

Adam Goodge, Xun Xu, Bryan Hooi, Wee Siong Ng, Jingyi Liao, Yongyi Su, Xulei Yang

机构 * Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR), Singapore(信息通信研究所,科技研究局(A*STAR),新加坡) School of Computing, National University of Singapore, Singapore(计算学院,新加坡国立大学,新加坡)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17824 2025-06-24 quant-ph cs.CR cs.LG 70%

Quantum-Hybrid Support Vector Machines for Anomaly Detection in Industrial Control Systems

Tyler Cultice, Md. Saif Hassan Onim, Annarita Giani, Himanshu Thapliyal

机构 * Department of Electrical Engineering and Compute Science, University of Tennessee, Knoxville, TN, 37996 USA(电气工程与计算科学系,田纳西大学,诺克斯维尔,TN,37996 USA) GE Vernova Research Center(通用电气研究中心)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

Comments 12 pages, 6 tables, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13419 2025-05-21 cs.CL cs.AI cs.CY cs.LG cs.SC 70%

From Words to Worlds: Compositionality for Cognitive Architectures

Ruchira Dhar, Anders Søgaard

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to ICML 2024 Workshop on LLMs & Cognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05501 2025-05-12 cs.CV cs.AI eess.IV 70%

Preliminary Explorations with GPT-4o(mni) Native Image Generation

Pu Cao, Feng Zhou, Junyi Ji, Qingye Kong, Zhixiang Lv, Mingjian Zhang, Xuekun Zhao, Siqi Wu, Yinghui Lin, Qing Song, Lu Yang

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04070 2025-04-08 cs.MA cs.AI 70%

Enforcement Agents: Enhancing Accountability and Resilience in Multi-Agent AI Frameworks

Sagar Tamang, Dibya Jyoti Bora

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01850 2025-04-03 cs.SE cs.AI 70%

Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks

Ali Al-Kaswan, Sebastian Deatc, Begüm Koç, Arie van Deursen, Maliheh Izadi

专题命中 其他安全 :alignment(abstract);harmlessness(abstract);分类 cs.AI

Comments FSE'25 Technical Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09555 2025-03-04 cs.CV cs.AI 70%

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis

Tingxuan Chen, Kun Yuan, Vinkle Srivastav, Nassir Navab, Nicolas Padoy

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23496 2025-03-04 cs.CL 70%

Smaller Large Language Models Can Do Moral Self-Correction

Guangliang Liu, Zhiyu Xue, Xitong Zhang, Rongrong Wang, Kristen Marie Johnson

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07587 2025-02-12 cs.LG 70%

SEMU: Singular Value Decomposition for Efficient Machine Unlearning

Marcin Sendera, Łukasz Struski, Kamil Książek, Kryspin Musiol, Jacek Tabor, Dawid Rymarczyk

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏