arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7968 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7968 篇

1804.04268 2018-04-13 cs.AI 79%

Incomplete Contracting and AI Alignment

Dylan Hadfield-Menell, Gillian Hadfield

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.03449 2017-01-13 stat.ML cs.LG math.PR 79%

Manifold Alignment Determination: finding correspondences across different data views

Andreas Damianou, Neil D. Lawrence, Carl Henrik Ek

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments NIPS workshop on Multi-Modal Machine Learning, 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.04576 2016-10-17 cs.LG stat.ML 79%

Kernel Alignment Inspired Linear Discriminant Analysis

Shuai Zheng, Chris Ding

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Joint European Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD, 2014

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.03650 2016-04-28 cs.CL 79%

Smoothing parameter estimation framework for IBM word alignment models

Vuong Van Bui, Cuong Anh Le

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1109.6838 2011-10-03 cs.MA cs.AI 79%

Distributed Air Traffic Control : A Human Safety Perspective

Sarvesh Nikumbh, Joeprakash Nathaman, Rahul Vartak

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Comments Extended Abstract, 3 pages, Accepted at IBM Collaborative Academia Research Exchange (I-CARE)-2011, uses ACM-Proceeding style file

详情

展开后加载摘要…

URL PDF HTML 收藏
1106.4570 2011-06-24 cs.GT cs.AI 79%

Competitive Safety Analysis: Robust Decision-Making in Multi-Agent Systems

M. Tennenholtz

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Journal ref Journal Of Artificial Intelligence Research, Volume 17, pages 363-378, 2002

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0311045 2009-12-01 cs.AI 79%

Unsupervised Grammar Induction in a Framework of Information Compression by Multiple Alignment, Unification and Search

J Gerard Wolff

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Journal ref Proceedings of the Workshop and Tutorial on Learning Context-Free Grammars (in association with the 14th European Conference on Machine Learning and the 7th European Conference on Principles and Practice of Knowledge Discovery in Databases (ECML/PKDD 2003), September 2003, Cavtat-Dubrovnik, Croata), editors: C. de la Higuera and P. Adriaans and M. van Zaanen and J. Oncina, pp 113-124

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09702 2026-07-14 cs.GT q-fin.TR 新提交 79%

Fundamental market design as a layer of AI-agent alignment

作为人工智能-智能体对齐层次的基础市场设计

Omar Inverso, Emilio Tuosto, Dragisa Zunic

专题命中 其他安全 :alignment(title,abstract)

AI总结 研究探讨市场中人工智能-智能体对齐,提出将基础市场设计视为该对齐层次,通过对市场核心形式化建模,利用理论计算机科学严谨性构建透明盒模型,支持激励分析与机制设计,使期望行为受青睐,不良行为难维持。

Comments Accepted as at the EC'26 Workshop on Incentive-Based AI Alignment, co-located with the 27th ACM Conference on Economics and Computation, Rome, Italy, July 2026. This version is prepared for public dissemination following workshop acceptance

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15300 2026-05-18 cs.CV 79%

Deep Pre-Alignment for VLMs

视觉语言模型的深度预对齐

Tianyu Yu, Kechen Fang, Zihao Wan, Kaidong Zhang, Yicheng Zhang, Jun Song, Bo Zheng, Yuan Yao

机构 * Tsinghua University Shanghai Qi Zhi Institute Taobao \& Tmall Group of Alibaba

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文提出深度预对齐(DPA),通过替换传统ViT编码器为小型VLM作为感知器,实现视觉特征与目标大语言模型文本空间的深度对齐,提升了多模态基准性能,并降低了语言能力遗忘。

Comments Accepted by ICML 2026. Project Website: https://github.com/THUMAI-Lab/Deep-Pre-Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04295 2024-12-17 cs.CV 79%

Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment

Jiayi Guo, Junhao Zhao, Chaoqun Du, Yulin Wang, Chunjiang Ge, Zanlin Ni, Shiji Song, Humphrey Shi, Gao Huang

专题命中 其他安全 :alignment(title,abstract)

Comments GitHub: https://github.com/SHI-Labs/Diffusion-Driven-Test-Time-Adaptation-via-Synthetic-Domain-Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13266 2024-10-18 physics.soc-ph 79%

Continuous agent-based modeling of adult-child pairs based on a pseudo-energy: Relevance for public safety and egress efficiency

Chuan-Zhi Thomas Xie, Tie-Qiao Tang, Alexandre Nicolas

专题命中 其他安全 :safety(title,abstract)

Journal ref Safety Science, 2024, 177, pp.106576

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07885 2022-12-21 cs.CV 79%

Clover: Towards A Unified Video-Language Alignment and Fusion Model

Jingjia Huang, Yinan Li, Jiashi Feng, Xinglong Wu, Xiaoshuai Sun, Rongrong Ji

专题命中 其他安全 :alignment(title,abstract)

Comments Update Tri-modal Alignment task

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22826 2026-05-25 cs.CL cs.AI cs.GT cs.MA 79%

Evaluating Large Language Models in a Complex Hidden Role Game

评估大型语言模型在复杂隐藏角色游戏中的表现

Niklas Bauer

机构 * University of Göttingen(哥廷根大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过社交推理游戏《秘密希特勒》评估大型语言模型的推理、说服和欺骗能力,引入新指标并发现当前模型在复杂多轮操纵中效果不佳。

Comments Master's thesis, University of Göttingen

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06696 2026-05-11 cs.AI cs.LG cs.MA 79%

Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations

多智能体AI中的隐藏联盟:来自内部表示的谱诊断

Cameron Berg, Susan L. Schneider, Mark M. Bailey

机构 * Reciprocal Research(递归研究) Center for the Future of AI, Mind, and Society(人工智能、心智与社会未来中心) Florida Atlantic University(佛罗里达 Atlantic 大学) Biological and Computational Intelligence Center(生物与计算智能中心) National Intelligence University(国家情报大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过分析多智能体系统内部神经表示的谱分区方法,检测隐藏联盟结构,验证了该方法在强化学习和大语言模型中的有效性,揭示了代表层次结构。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16448 2026-03-31 cs.AI cs.LG 79%

Information-theoretic Distinctions Between Deception and Confusion

信息论视角下欺骗与混淆的区别

Robin Young

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 本文从信息论角度区分两种AI安全失效模式:欺骗对齐与目标漂移,揭示二者在人类-AI系统不同接口的信息分歧,提出形式化模型和思想实验,为大型语言模型对齐挑战提供新视角。

Comments Proceedings of the 14th IJCNLP and the 4th AACL (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17577 2026-01-27 cs.HC cs.AI cs.CL 79%

Status Hierarchies in Language Models

语言模型中的地位层级

Emilio Barkett

机构 * Brigham Young University–Hawaii COLUMBIA UNIVERSITY(哥伦比亚大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

AI总结 语言模型在多智能体环境中会因地位线索形成层级,高地位分配反而降低高能力模型的服从,揭示AI系统中的新兴社会行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03442 2025-11-17 cs.CL cs.AI 79%

Are language models rational? The case of coherence norms and belief revision

Thomas Hofweber, Peter Hase, Elias Stengel-Eskin, Mohit Bansal

机构 * Department of Philosophy University of North Carolina at Chapel Hill(哲学系北卡罗来纳大学教堂山分校) Department of Computer Science University of North Carolina at Chapel Hill(计算机科学系北卡罗来纳大学教堂山分校) Department of Computer Science University of Texas at Austin(计算机科学系德克萨斯大学奥斯汀分校)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments substantial expansions of sections 4 and 5, updated references, numerous smaller additions and clarifications

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05580 2025-11-11 q-bio.NC cs.CL cs.CY cs.HC 79%

Approximating the Mathematical Structure of Psychodynamics

Bryce-Allen Bagley, Navin Khoshnan

机构 * Stanford University(斯坦福大学) Mathematical Medicine Group(数学医学组) Department of Neurosurgery(神经外科系) Physician-Scientist Training Program(医师科学家培训计划)

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11613 2025-06-16 cs.LG cs.AI 79%

Model Organisms for Emergent Misalignment

Edward Turner, Anna Soligo, Mia Taylor, Senthooran Rajamanoharan, Neel Nanda

专题命中 其他安全 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15905 2026-08-18 cs.CV cs.MM 新提交 78%

CLARA: Clip-Level Multimodal Alignment with VLM-Derived Rationales for Hateful Video Detection

CLARA:基于VLM衍生理由的片段级多模态对齐用于仇恨视频检测

Yuchen Zhang, Shuang Dai, Zeyu Fu, Yunfei Long, Ravi Shekhar, Haralambos Mouratidis

机构 * Institute for Analytics and Data Science, University of Essex(埃塞克斯大学分析与数据科学研究所) University of Exeter(埃克塞特大学) Queen Mary University of London(伦敦大学玛丽皇后学院)

专题命中 其他安全 :alignment(title,abstract)

AI总结 本研究提出CLARA框架,以片段级多模态对齐结合VLM衍生理由,在三个仇恨视频数据集上实现了优于现有最优方法的仇恨视频检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15698 2026-08-18 cs.CV cs.IR 新提交 78%

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

ConceptFormer:学习自适应潜在概念以实现视觉文档检索中的查询-文档对齐

Peng Chunyi, Xu Zhipeng, Yan Yukun, Liu Zhenghao, Yu Shi, Mei Sen, Sun Yubo, Zhang Yongheng, Zhou Jie, Gu Yu, Yu Ge, Sun Maosong

机构 * Northeastern University(东北大学) Tsinghua University(清华大学) Peking University(北京大学)

专题命中 其他安全 :alignment(title,abstract)

AI总结 针对视觉文档检索中现有监督信号的局限,本文提出ConceptFormer框架,以自适应潜在概念为中间表示衔接语义鸿沟,在基准测试中较最强基线实现了16.7%、22.1%的NDCG@10相对提升,性能优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15456 2026-08-18 cs.CV 新提交 78%

AlignJEPA: Predictive Vision-Language Alignment for Remote Sensing Foundation Models

AlignJEPA:面向遥感基础模型的预测性视觉-语言对齐

Md Aminur Hossain, Omkumar Vaghasiya, Rajeev Ranjan Dwivedi, Vinod Kurmi, Biplab Banerjee

机构 * Space Applications Centre, ISRO(印度空间研究组织空间应用中心) Indian Institute of Science Education and Research (IISER) Bhopal(印度科学教育与研究学院(博帕尔)) Indian Institute of Technology Bombay(印度理工学院孟买分校)

专题命中 其他安全 :alignment(title,abstract)

AI总结 该研究针对遥感基础模型与自然语言对齐不足的问题,提出受JEPA启发的AlignJEPA框架,采用轻量级预测对齐网络,结合语义预测与双向对比检索,实现了参数高效的视觉-语言对齐。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.13536 2026-08-17 cs.OS 版本更新 78%

Don't Let AI Agents YOLO Your Files: Information and Control in Agent-Native Filesystems

别让AI代理占用你的文件:将信息和控制移交给文件系统以实现代理安全和自主性

Shawn Wanxiang Zhong, Junxuan Liao, Jing Liu, Mai Zheng, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau

专题命中 其他安全 :safety(title,abstract)

AI总结 本文研究了AI代理对文件系统的滥用问题,提出YoloFS文件系统通过三种技术提升代理的安全性和自主性,减少用户交互并提高任务完成率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09147 2026-08-11 cs.CV 新提交 78%

RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection

RefineAny3D:作为语义对齐的深度细化用于单目3D检测

Zhihao Zhang, Gengwei Zhang, Tianlong Chen, Xiaoming Liu

机构 * Michigan State University(密歇根州立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 其他安全 :alignment(title,abstract)

AI总结 RefineAny3D将单目3D检测的深度细化转化为视觉对齐问题,通过VLM实现无需数值预测的深度修正,在多类检测工具上均有性能提升且可泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08805 2026-08-11 cs.CV 新提交 78%

LASA: Language-and-Source-Anchored Alignment for Domain Generalized Semantic Segmentation

LASA:面向域泛化语义分割的语言与源锚定对齐

Jinhong Zhu, Weiqi Yan, Shengchuan Zhang, Liujuan Cao

机构 * Xiamen University(厦门大学)

专题命中 其他安全 :alignment(title,abstract)

AI总结 针对域泛化语义分割中传统方法损害特征完整性的问题,提出LASA框架,含三个协同组件,实验显示其性能优于现有最优方法。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08038 2026-08-11 cs.SE 新提交 78%

Stateful Multi-Agent LLMs for Cross-View Interface Alignment in Automotive Model-Based Systems Engineering

面向汽车基于模型的系统工程的有状态多智能体大语言模型,用于跨视图接口对齐

Aleksei Velsh, Nenad Petrovic, Alois Knoll

专题命中 其他安全 :alignment(title,abstract)

AI总结 针对LLMs在汽车MBSE中引发的架构漂移问题,提出有状态多智能体验证流水线,在ADAS场景中实现97%实体可追溯性等指标,证明对抗性审核可让LLMs可靠生成零错误MBSE架构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25129 2026-08-04 cs.CV 版本更新 78%

AirSplat: Alignment and Rating for Robust Feed-Forward 3D Gaussian Splatting

AirSplat:基于鲁棒前馈3D高斯散射的对齐与评分

Minh-Quan Viet Bui, Jaeho Moon, Munchurl Kim

机构 * KAIST(韩国科学技术院)

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文提出AirSplat框架,通过自一致性姿态对齐和基于评分的透明度匹配技术,提升无姿态视角合成的重建质量。

Comments Project page: https://kaist-viclab.github.io/airsplat-site, accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02457 2026-08-04 cs.CV 版本更新 78%

PatchAlign3D: Local Feature Alignment for Dense 3D Shape Understanding

PatchAlign3D: 本地特征对齐用于密集3D形状理解

Souhail Hadgi, Bingchen Gong, Ramana Sundararaman, Emery Pierson, Lei Li, Peter Wonka, Maks Ovsjanikov

机构 * École polytechnique(巴黎政治学院) University of Virginia(弗吉尼亚大学) KAUST(王国立阿联酋科技大学)

专题命中 其他安全 :alignment(title,abstract)

AI总结 PatchAlign3D通过点云直接生成语言对齐的片段特征,实现高效零样本3D部分分割,优于传统渲染方法。

Comments CVPR 2026. Project website: https://souhail-hadgi.github.io/patchalign3dsite/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28581 2026-07-31 cs.CV 新提交 78%

ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation

ROAD:用于3D形状生成的判别式语义的互目标对齐

Xiao Luo, Mingyang Du, Xin Zhou, Tianrui Feng, Xiwu Chen, Xiaofan Li, Jiangning Zhang, Dingkang Liang

机构 * Huazhong University of Science and Technology(华中科技大学) Megvii(旷视科技) Zhejiang University(浙江大学)

专题命中 其他安全 :alignment(title,abstract)

AI总结 ROAD框架通过迁移判别式3D基础模型的先验,采用互目标对齐策略,仅用1.5%训练数据就实现了高保真3D生成,大幅降低了计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22241 2026-07-31 cs.CV 版本更新 78%

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment

AgentHOI:通过隐式表示对齐进行人类-物体交互视频生成的多智能体推理

Ziyao Huang, Shunkai Li, Juan Cao, Chenyu Li, Youliang Zhang, Zixiang Zhou, Cong Wang, Yuan Zhou, Qinglin Lu, Fan Tang

机构 * University of Chinese Academy of Sciences(中国科学院大学) Tencent HunYuan(腾讯混元)

专题命中 其他安全 :alignment(title,abstract)

AI总结 研究针对HOI视频生成中现有方法依赖显式运动控制的问题,提出AgentHOI,通过多智能体推理和隐式文本-运动对齐策略,实现文本驱动的HOI视频生成,提升了交互自然性等,改进了复杂场景下的表现。

Comments Under review. The code is available at https://github.com/bone-11/agenthoi

详情

展开后加载摘要…

URL PDF HTML 收藏