arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7945 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7945 篇

2603.00097 2026-03-03 q-bio.BM cs.AI cs.CE cs.IR cs.LG 81%

Exploring Drug Safety Through Knowledge Graphs: Protein Kinase Inhibitors as a Case Study

通过知识图谱探索药物安全性:蛋白激酶抑制剂作为案例研究

David Jackson, Michael Gertz, Jürgen Hesser

机构 * David Jackson Institute of Informatics University of Amsterdam(大卫·杰克逊信息学研究所 阿姆斯特丹大学) Michael Gertz Institute of Computer Science Heidelberg University(迈克尔·格茨计算机科学研究所海德堡大学) Jürgen Hesser Data Analysis and Modeling in Medicine Mannheim Institute for Intelligent Systems in Medicine (MIISM) Medical Faculty Mannheim, Heidelberg University(朱尔根·赫瑟医学数据分析与建模 马尔姆研究所(MIISM)海德堡大学医学系)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于知识图谱的框架,整合多种数据源以预测蛋白激酶抑制剂的不良药物反应,通过分析疗效、靶点相似性和不良事件相关性,提升药物安全性研究。

Comments 14 pages, 5 figures. Code and data available at https://github.com/davidjackson99/PKI_KG

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23588 2026-03-02 cs.CV cs.AI cs.LG 81%

Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning

高维跨模态对齐冻结语言和图像模型以实现高效的图像描述生成

Abhishek Dalvi, Vasant Honavar

机构 * Artificial Intelligence Research Laboratory(人工智能研究实验室) The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 HDFLIM通过冻结语言和图像模型的高维计算实现高效的图像描述生成,无需参数调优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22775 2026-02-27 cs.HC cs.AI cs.CL 81%

TherapyProbe: Generating Design Knowledge for Relational Safety in Mental Health Chatbots Through Adversarial Simulation

TherapyProbe: 通过对抗性模拟生成关系安全设计知识以改进心理健康聊天机器人

Joydeep Chandra, Satyam Kumar Navneet, Yong Zhang

机构 * BNRIST, Dept. of CST, Tsinghua University(北京理工大学(Tsinghua大学)) Independent Researcher(独立研究者)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI

AI总结 TherapyProbe通过对抗性模拟生成关系安全设计知识,帮助改进心理健康聊天机器人的交互模式质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21215 2026-02-26 cs.CL cs.AI 81%

Inference-time Alignment via Sparse Junction Steering

推理时对齐 via 稀疏节点引导

Runyi Hu, Jie Zhang, Shiqian Zhao, Jiale Meng, Jiwei Li, Jason Zeng, Ming Wu, Michael Heinrich, Yonggang Wen, Tianwei Zhang

机构 * Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学) G Labs(0G实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 通过稀疏节点引导方法,实现更高效的推理过程对齐,减少计算开销并提升生成质量。

Comments 28 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15676 2026-02-18 cs.LG cs.AI 81%

Relative Geometry of Neural Forecasters: Linking Accuracy and Alignment in Learned Latent Geometry

神经预报器的相对几何:连接准确性和对齐性在学习潜在几何中的联系

Deniz Kucukahmetler, Maximilian Jean Hemmann, Julian Mosig von Aehrenfeld, Maximilian Amthor, Christian Deubel, Nico Scherf, Diaaeldin Taha

机构 * Max Planck Institute for Human Cognitive and Brain Sciences(马克斯·普朗克人类认知与脑科学研究所) School of Embedded Composite Artificial Intelligence (SECAI)(嵌入式复合人工智能学院) Leipzig University(莱比锡大学) Max Planck Institute for Mathematics in the Sciences(马克斯·普朗克数学研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本研究通过相对几何方法探讨神经预报器的对齐与准确性关系,揭示了不同模型家族在表示动态结构上的差异及预测性能的关联。

Comments Accepted to Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15368 2026-02-18 cs.CV cs.AI cs.LG eess.IV 81%

GMAIL: Generative Modality Alignment for generated Image Learning

GMAIL: 生成模态对齐用于生成图像学习

Shentong Mo, Sukmin Yun

机构 * Department of Machine Learning, CMU, USA(卡内基梅隆大学机器学习系) Department of Machine Learning, MBZUAI, UAE(马斯克大学人工智能研究所) Department of Artificial Intelligence, Hanyang University ERICA, South Korea(翰阳大学ERICA人工智能系)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 GMAIL通过多模态学习方法对齐生成图像与真实图像,提升视觉-语言任务中的生成图像学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14471 2026-02-17 cs.MA cs.AI cs.GT cs.LG 81%

Socially-Weighted Alignment: A Game-Theoretic Framework for Multi-Agent LLM Systems

社交加权对齐:多智能体大语言模型系统的博弈论框架

Furkan Mumcu, Yasin Yilmaz

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出社交加权对齐框架,通过引入社交权重解决多智能体LLM系统中的个体理性与集体稳定矛盾,通过理论分析和模拟验证了关键阈值的存在。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12968 2026-02-16 cs.IR cs.AI cs.CL 81%

RGAlign-Rec: Ranking-Guided Alignment for Latent Query Reasoning in Recommendation Systems

RGAlign-Rec: 基于排名引导的潜在查询推理中的对齐方法

Junhua Liu, Yang Jihao, Cheng Chang, Kunrong LI, Bin Fu, Kwan Hui Lim

机构 * Forth AI Singapore(Forth AI新加坡) Shopee Singapore(Shopee新加坡) Singapore Uni. of Tech. and Design(新加坡技术与设计大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 RGAlign-Rec通过闭环对齐框架结合语义推理和增强查询模型,提升电子商务推荐系统的预测准确性和服务质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02387 2026-02-11 cs.LG cs.AI 81%

BiSSL: Enhancing the Alignment Between Self-Supervised Pretraining and Downstream Fine-Tuning via Bilevel Optimization

BiSSL: 通过双层优化增强自监督预训练与下游微调之间的对齐

Gustav Wagner Zakarias, Lars Kai Hansen, Zheng-Hua Tan

机构 * Aalborg University(奥胡斯大学) Technical University of Denmark(技术大学) Pioneer Centre for AI(先锋人工智能中心)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 BiSSL通过双层优化增强自监督预训练与下游微调之间的对齐,提升模型在下游任务中的性能。

Journal ref Transactions on Machine Learning Research (TMLR), (02/2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06869 2026-02-09 cs.CL cs.LG 81%

Uncovering Cross-Objective Interference in Multi-Objective Alignment

揭示多目标对齐中的跨目标干扰

Yining Lu, Meng Jiang

机构 * University of Notre Dame(诺丁汉大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出CTWA方法,通过保持目标奖励与训练信号的正协方差来缓解多目标对齐中的跨目标干扰问题,并分析了干扰与模型几何属性的关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11085 2026-02-03 cs.LG cs.AI 81%

On the Importance of Pretraining Data Alignment for Atomic Property Prediction

关于预训练数据对齐在原子性质预测中的重要性

Yasir Ghunaim, Hasan Abed Al Kader Hammoud, Bernard Ghanem

机构 * King Abdullah University of Science and Technology (KAUST)(国王阿卜杜勒-阿齐兹大学科学与技术学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文发现,通过精心选择的任务对齐数据集预训练,可提升原子性质预测的性能,且比大规模混合数据集预训练更有效。

Comments Published in Transactions on Machine Learning Research (TMLR), 2026

Journal ref Transactions on Machine Learning Research (TMLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17652 2026-02-02 cs.LG cs.AI 81%

Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective

重新思考强化学习中用于大语言模型推理的采样标准:一种能力-难度对齐视角

Deyang Kong, Qi Guo, Xiangyu Xi, Wei Wang, Jingang Wang, Xunliang Cai, Shikun Zhang, Wei Ye

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出CDAS方法,通过能力-难度对齐提升大语言模型推理的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18255 2026-01-27 cs.LG cs.AI 81%

Beyond Retention: Orchestrating Structural Safety and Plasticity in Continual Learning for LLMs

超越保留:在连续学习中协调结构安全与可塑性

Fei Meng

机构 * Yangtze Delta Region Institute of Tsinghua University(清华大学 Yangtze Delta 区研究所)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出OSW方法,通过正交子空间唤醒协调连续学习中的结构安全与可塑性,有效保留脆弱任务能力并维持新任务学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19893 2026-01-27 cs.LG cs.AI cs.IT eess.IV math.IT 81%

Distillation-Enabled Knowledge Alignment for Generative Semantic Communications of AIGC Images

基于知识对齐的生成语义通信中的知识蒸馏

Jingzhi Hu, Geoffrey Ye Li

机构 * Department of Electrical and Electronic Engineering, Imperial College London(帝国理工学院伦敦分校电子与电气工程系)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出DeKA-g算法,通过知识蒸馏和低秩适应,提升生成语义通信中边缘与云生成图像的一致性及传输质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03747 2026-01-08 cs.LG cs.CL stat.AP 81%

Context-Alignment: Activating and Enhancing LLM Capabilities in Time Series

上下文对齐:在时间序列中激活和增强大语言模型的能力

Yuxiao Hu, Qian Li, Dongxiao Zhang, Jinyue Yan, Yuntian Chen

机构 * The Hong Kong Polytechnic University(香港理工大学) Ningbo Institute of Digital Twin(宁波数字孪生研究所) Eastern Institute of Technology(东部技术研究所) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

AI总结 本文提出上下文对齐方法,通过多模态输入和图神经网络增强LLMs在时间序列任务中的能力,提升逻辑和结构理解,提高预测性能。

Comments This paper has been accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18526 2025-12-08 cs.AI cs.LG 81%

Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models

反事实推理用于可操控的多元价值观对齐大语言模型

Hanze Guo, Jing Yao, Xiao Zhou, Xiaoyuan Yi, Xing Xie

机构 * Renmin University of China(中国人民大学) Microsoft Research Asia(微软亚洲研究院) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程研究中心,教育部)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 COUPLE通过反事实推理框架实现多元价值观对齐,解决现有方法在处理细粒度价值目标时的依赖性和优先级控制问题。

Comments NeurIPS 2025. 41 pages, 7 figures

Journal ref The Thirty-Ninth Annual Conference on Neural Information Processing Systems. (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00120 2025-12-02 cs.SD cs.AI cs.CV cs.LG cs.MM 81%

Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment

Art2Music: 通过多模态情感对齐生成艺术图像的音乐

Jiaying Hong, Ting Zhu, Thanet Markchom, Huizhi Liang

机构 * School of Computing Newcastle University(计算学院新castle大学) Department of Computer Science University of Reading(计算机科学系阅读大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 Art2Music通过多模态情感对齐生成艺术图像的音乐,利用轻量级跨模态框架提升音乐生成的感知自然性和频谱保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20931 2025-11-27 cs.CV cs.AI cs.LG 81%

Open Vocabulary Compositional Explanations for Neuron Alignment

开放词汇组合解释用于神经元对齐

Biagio La Rosa, Leilani H. Gilpin

机构 * Department of Computer Science and Engineering University of California, Santa Cruz(计算机科学与工程系加州大学圣克ruz分校)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种开放词汇组合解释框架,通过语义分割生成掩码来实现神经元对齐,提升解释的灵活性和适用性。

Comments 47 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10215 2025-11-14 cs.CL cs.AI 81%

Persona-Aware Alignment Framework for Personalized Dialogue Generation

Guanrong Li, Xinyu Liu, Zhen Wu, Xinyu Dai

机构 * National Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) Nanjing University(南京大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Pre-MIT Press publication version

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08080 2025-11-12 cs.LG cs.AI 81%

Hierarchical Structure-Property Alignment for Data-Efficient Molecular Generation and Editing

Ziyu Fan, Zhijian Huang, Yahan Li, Xiaowen Hu, Siyuan Shen, Yunliang Wang, Zeyu Zhong, Shuhong Liu, Shuning Yang, Shangqian Wu, Min Wu, Lei Deng

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Central South University(中南大学) Institute for Infocomm Research Agency for Science, Technology and Research (A* STAR)(信息通信研究所(A* STAR))

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06776 2025-11-11 cs.LG cs.AI 81%

Data Trajectory Alignment for LLM Domain Adaptation: A Two-Phase Synthesis Framework for Telecommunications Mathematics

Zhicheng Zhou, Jing Li, Suming Qiu, Junjie Huang, Linyuan Qiu, Zhijie Sun

机构 * Global Technical Service (GTS) Huawei Technologies Co., Ltd(华为技术有限公司全球技术服务部)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05968 2025-11-11 cs.CV cs.AI cs.LG 81%

DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities

Nagur Shareef Shaik, Teja Krishna Cherukuri, Adnan Masood, Dong Hye Ye

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Accepted for Oral Presentation at the 40th AAAI Conference on Artificial Intelligence (AAAI-26), Main Technical Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08872 2025-11-04 cs.AI cs.GT cs.HC cs.LG cs.MA 81%

GTAlign: Game-Theoretic Alignment of LLM Assistants for Social Welfare

Siqi Zhu, David Zhang, Pedro Cisneros-Velarde, Jiaxuan You

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) VMware Research(VMware研究)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 31 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14205 2025-10-30 cs.CL cs.AI 81%

DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans

Bingsheng Yao, Bo Sun, Yuanzhe Dong, Yuxuan Lu, Dakuo Wang

机构 * Northeastern University(东北大学) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments In Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16895 2025-10-23 cs.CV cs.AI cs.LG 81%

With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You

Fabian Gröger, Shuo Wen, Huyen Le, Maria Brbić

机构 * EPFL(瑞士联邦理工学院) University of Basel(巴塞尔大学) HSLU(苏黎世联邦理工学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12116 2025-10-15 cs.CL cs.AI 81%

Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language Models

Bajian Xiang, Shuaijiang Zhao, Tingwei Guo, Wei Zou

机构 * Beike Inc.(贝克公司) Bairong Inc.(柏睿公司)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20333 2025-10-14 cs.CL cs.AI 81%

Multi-Scale Manifold Alignment for Interpreting Large Language Models: A Unified Information-Geometric Framework

Yukun Zhang, Qi Dong

机构 * The Chinese University of Hong Kong(香港中文大学) Fudan University(复旦大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05129 2025-10-14 cs.CL cs.LG 81%

Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models

Qingshu Xu, Hong Jiao, Tianyi Zhou, Ming Li, Nan Zhang, Sydney Peters, Yanbin Fu

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Shandong Jiaotong University(山东交通大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03227 2025-10-14 cs.LG cs.AI 81%

Bridging Neural ODE and ResNet: A Formal Error Bound for Safety Verification

Abdelrahman Sayed Sayed, Pierre-Jean Meyer, Mohamed Ghazel

机构 * Univ Gustave Eiffel(法国埃菲尔大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 18 pages, 5 figures, Accepted for publication in the proceedings of the 8th International Symposium on AI Verification SAIV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20179 2025-10-13 cs.CL cs.AI cs.RO 81%

Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs

Zichao Hu, Junyi Jessy Li, Arjun Guha, Joydeep Biswas

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Conference on Language Modeling (COLM) 2025, Project site: https://amrl.cs.utexas.edu/robo-instruct/

详情

展开后加载摘要…

URL PDF HTML 收藏