arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7971 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7971 篇

2601.00516 2026-01-05 cs.LG cs.AI 73%

Trajectory Guard -- A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI

轨迹守护 -- 一种轻量级、序列感知的模型,用于代理AI中的实时异常检测

Laksh Advani

机构 * Laksh Advani

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 轨迹守护通过对比学习和重建学习,实现对代理AI中任务轨迹对齐和序列有效性的联合检测,提升实时异常检测性能。

Comments Accepted to AAAI Trustagent 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23487 2025-12-30 cs.LG cs.AI stat.ML 73%

ML Compass: Navigating Capability, Cost, and Compliance Trade-offs in AI Model Deployment

ML Compass:在AI模型部署中导航能力、成本和合规性之间的权衡

Vassilis Digalakis, Ramayya Krishnan, Gonzalo Martin Fernandez, Agni Orfanoudaki

机构 * Questrom School of Business, Boston University(波士顿大学Questrom商学院) Centre de Formació Interdisciplinària Superior and Universitat Politècnica de Catalunya(巴塞罗那理工大学跨学科教育中心)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 ML Compass通过系统性框架优化模型选择,考虑能力、成本和合规性之间的权衡,提供部署导向的推荐和排行榜。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17638 2025-11-25 cs.LG cs.AI 73%

Model-to-Model Knowledge Transmission (M2KT): A Data-Free Framework for Cross-Model Understanding Transfer

模型到模型知识传输(M2KT):一种无数据的跨模型理解迁移框架

Pratham Sorte

机构 * Department of Computer Science(计算机科学系) Engineering MIT-World Peace University, Pune, India(工程学院 MIT-世界和平大学 印度邦普尼)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 M2KT提出了一种无数据的跨模型知识传输方法,通过概念空间交换知识包,实现高效的知识迁移和模型自我改进。

Comments 8 pages including figures, prepared in IEEE conference style. Preprint. Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10509 2025-09-16 cs.LG cs.AI 73%

The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback

Sai Teja Reddy Adapala

机构 * University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments 5 pages, 3 figures, 2 tables. Code is available at: https://github.com/imsaitejareddy/ouroboros-effect-experiment

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15841 2025-08-25 cs.CL cs.LG 73%

A Review of Developmental Interpretability in Large Language Models

Ihor Kendiukhov

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00741 2025-08-04 cs.CL cs.AI 73%

Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data

Sohaib Imran, Rob Lamb, Peter M. Atkinson

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21839 2025-07-30 cs.CY cs.AI 73%

Against racing to AGI: Cooperation, deterrence, and catastrophic risks

Leonard Dung, Max Hellrigel-Holderbaum

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03708 2025-05-30 cs.CL cs.AI stat.ML 73%

Toward universal steering and monitoring of AI models

Daniel Beaglehole, Adityanarayanan Radhakrishnan, Enric Boix-Adserà, Mikhail Belkin

机构 * Computer Science and Engineering(计算机科学与工程) Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所) UC San Diego(圣地亚哥大学) Harvard SEAS(哈佛大学工程与应用科学学院) MIT Mathematics(MIT数学系) Halıcıoğlu Data Science Institute(Halıcıoğlu数据科学研究所) Harvard CMSA(哈佛大学计算机科学与应用数学系)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21664 2025-05-29 cs.CY cs.AI 73%

Expert Survey: AI Reliability & Security Research Priorities

Joe O'Brien, Jeremy Dolan, Jay Kim, Jonah Dykhuizen, Jeba Sania, Sebastian Becker, Jam Kraprayoon, Cara Labrador

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17513 2025-05-20 cs.CL cs.AI 73%

Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models

Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, Andreas Bulling

机构 * University of Stuttgart(斯图加特大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments ICML 2024 Workshop on Mechanistic Interpretability version: https://openreview.net/forum?id=yEwEVoH9Be

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11311 2025-05-19 cs.MA cs.AI cs.LG 73%

Explaining Strategic Decisions in Multi-Agent Reinforcement Learning for Aerial Combat Tactics

Ardian Selmonaj, Alessandro Antonucci, Adrian Schneider, Michael Rüegsegger, Matthias Sommer

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

Comments Published as a journal chapter in NATO Journal of Science and Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02091 2025-04-15 cs.AI cs.GT cs.LG cs.MA 73%

The Problem of Social Cost in Multi-Agent General Reinforcement Learning: Survey and Synthesis

Kee Siong Ng, Samuel Yang-Zhao, Timothy Cadogan-Cowper

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments 67 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05305 2025-03-31 cs.CL cs.AI 73%

Output Scouting: Auditing Large Language Models for Catastrophic Responses

Andrew Bell, Joao Fonseca

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments Work not ready, further experiments needed to validate the method

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10288 2025-03-03 cs.CL cs.LG 73%

Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models

Francisco Eiras, Aleksandar Petrov, Philip H. S. Torr, M. Pawan Kumar, Adel Bibi

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

Comments Accepted to ICLR'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05040 2024-11-11 cs.CL cs.AI 73%

Bottom-Up and Top-Down Analysis of Values, Agendas, and Observations in Corpora and LLMs

Scott E. Friedman, Noam Benkler, Drisana Mosaphir, Jeffrey Rye, Sonja M. Schmer-Galunder, Micah Goldwater, Matthew McLure, Ruta Wheelock, Jeremy Gottlieb, Robert P. Goldman, Christopher Miller

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13697 2024-09-24 cs.CL cs.AI 73%

Prompt Baking

Aman Bhargava, Cameron Witkowski, Alexander Detkov, Matt Thomson

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments 25 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19348 2024-08-09 cs.LG cs.AI 73%

Deep Learning for Cross-Domain Data Fusion in Urban Computing: Taxonomy, Advances, and Outlook

Xingchen Zou, Yibo Yan, Xixuan Hao, Yuehong Hu, Haomin Wen, Erdong Liu, Junbo Zhang, Yong Li, Tianrui Li, Yu Zheng, Yuxuan Liang

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

Journal ref Inform.Fusion.113(2025)102606

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19334 2024-06-11 cs.AI cs.CL cs.CV cs.MM cs.SD 73%

LLMs Meet Multimodal Generation and Editing: A Survey

Yingqing He, Zhaoyang Liu, Jingye Chen, Zeyue Tian, Hongyu Liu, Xiaowei Chi, Runtao Liu, Ruibin Yuan, Yazhou Xing, Wenhai Wang, Jifeng Dai, Yong Zhang, Wei Xue, Qifeng Liu, Yike Guo, Qifeng Chen

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments 52 Pages with 16 Figures, 12 Tables, and 545 References. GitHub Repository at: https://github.com/YingqingHe/Awesome-LLMs-meet-Multimodal-Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.12582 2024-03-13 cs.CY cs.AI 73%

Understanding and Avoiding AI Failures: A Practical Guide

Heather M. Williams, Roman V. Yampolskiy

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03096 2024-02-14 cs.LG cs.AI cs.NE 73%

What Causes Polysemanticity? An Alternative Origin Story of Mixed Selectivity from Incidental Causes

Victor Lecomte, Kushal Thaman, Rylan Schaeffer, Naomi Bashkansky, Trevor Chow, Sanmi Koyejo

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00323 2024-01-19 cs.AI cs.LG 73%

Thought Cloning: Learning to Think while Acting by Imitating Human Thinking

Shengran Hu, Jeff Clune

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to NeurIPS 2023 as a spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00667 2023-09-06 cs.CL cs.LG 73%

Taken out of context: On measuring situational awareness in LLMs

Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, Owain Evans

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.02617 2021-06-07 cs.AI cs.LG 73%

Be Considerate: Objectives, Side Effects, and Deciding How to Act

Parand Alizadeh Alamdari, Toryn Q. Klassen, Rodrigo Toro Icarte, Sheila A. McIlraith

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.07421 2021-01-05 cs.CV cs.AI cs.LG cs.MM 73%

Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models

Tong Che, Xiaofeng Liu, Site Li, Yubin Ge, Ruixiang Zhang, Caiming Xiong, Yoshua Bengio

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.04071 2020-08-11 cs.CY cs.AI 73%

On Controllability of AI

Roman V. Yampolskiy

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.06497 2019-12-16 cs.CR cs.AI cs.LG 73%

Founding The Domain of AI Forensics

Ibrahim Baggili, Vahid Behzadan

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments Accepted for presentation at SafeAI2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.05590 2018-11-15 cs.LG cs.AI stat.ML 73%

Emergence of Addictive Behaviors in Reinforcement Learning Agents

Vahid Behzadan, Roman V. Yampolskiy, Arslan Munir

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16041 2026-08-18 cs.RO 新提交 71%

ScenarioCharacterization: A Modular Toolkit for Characterizing Safety across Trajectory Datasets

ScenarioCharacterization:用于刻画轨迹数据集安全性的模块化工具包

Ingrid Navarro, Yutong Duan, Jonathan Francis, Jean Oh

机构 * Robotics Institute, School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机学院机器人研究所) Stack AV Bosch Center for Artificial Intelligence(博世人工智能中心)

专题命中 其他安全 :safety(title)

AI总结 本文提出开源模块化框架ScenarioCharacterization,可自动刻画轨迹数据集的驾驶场景安全性,适配多数据集,在Waymo等数据集上验证,支持下游应用。

Comments 9 pages, 7 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20936 2026-08-14 physics.soc-ph nlin.AO 版本更新 71%

Empathy Modeling in Active Inference Agents for Perspective-Taking and Alignment

主动推断代理中的共情建模:用于视角转换与对齐

Mahault Albarracin, Hongju Pae, Philip Wilson, Anna Mikeda, Alejandro Jimenez-Rodriguez, Sanjeev V. Namjoshi, Harshil Shah

专题命中 其他安全 :alignment(title)

AI总结 该研究提出了一种基于主动推断的共情框架,通过视角转换实现稳定合作,揭示了共情结构对社会协调的关键作用。

Comments Code and data: https://doi.org/10.5281/zenodo.21908008

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20318 2026-08-04 cs.CV cs.MM 版本更新 71%

UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval

UniCVR:从对齐到重排序的统一零样本复合视觉检索

Haokun Wen, Xuemeng Song, Haoyu Zhang, Weili Guan, Xiangyu Zhao, Liqiang Nie

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) City University of Hong Kong(香港城市大学) Southern University of Science and Technology(南方科技大学) Pengcheng Laboratory(鹏城实验室)

专题命中 其他安全 :alignment(title)

AI总结 UniCVR提出首个统一零样本复合视觉检索框架,结合多模态大语言模型与视觉语言预训练模型,通过两阶段方法实现多任务联合优化,实验表明其在五个基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏