arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7945 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7945 篇

2510.12847 2025-10-16 cs.LG 79%

Lifting Manifolds to Mitigate Pseudo-Alignment in LLM4TS

Liangwei Nathan Zheng, Wenhao Liang, Wei Emma Zhang, Miao Xu, Olaf Maennel, Weitong Chen

机构 * The University of Adelaide(阿德莱德大学) The University of Queensland(昆士兰大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12355 2025-10-15 cs.CL 79%

Fine-grained Analysis of Brain-LLM Alignment through Input Attribution

Michela Proietti, Roberto Capobianco, Mariya Toneva

机构 * Sapienza University of Rome(罗马萨皮恩扎大学) Sony AI(索尼人工智能) Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11738 2025-10-15 cs.SD cs.AI cs.CV cs.MM 79%

SeeingSounds: Learning Audio-to-Visual Alignment via Text

Simone Carnemolla, Matteo Pennisi, Chiara Russo, Simone Palazzo, Daniela Giordano, Concetto Spampinato

机构 * University of Catania(卡塔尼亚大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments accepted to ACM Multimedia Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10732 2025-10-14 cs.CY 79%

When Openness Fails: Lessons from System Safety for Assessing Openness in AI

Tamara Paris, Shalaleh Rismani

专题命中 其他安全 :safety(title,abstract);分类 cs.CY

Comments Accepted to Symposium on Model Accountability, Sustainability and Healthcare (SMASH) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10280 2025-10-14 cs.CL 79%

On the Entity-Level Alignment in Crosslingual Consistency

Yihong Liu, Mingyang Wang, François Yvon, Hinrich Schütze

机构 * Center for Information and Language Processing, LMU Munich(信息与语言处理中心,慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心(MCML)) Sorbonne Université, CNRS, ISIR, France(索邦大学,CNRS,ISIR,法国)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26431 2025-10-14 cs.CL 79%

Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests

Yanbin Fu, Hong Jiao, Tianyi Zhou, Nan Zhang, Ming Li, Qingshu Xu, Sydney Peters, Robert W. Lissitz

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments need updates

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05184 2025-10-08 cs.AI 79%

Representation Potentials of Foundation Models for Multimodal Alignment: A Survey

Jianglin Lu, Hailing Wang, Yi Xu, Yizhou Wang, Kuo Yang, Yun Fu

机构 * Department of Electrical and Computer Engineering, Northeastern University(东北大学电气与计算机工程系) Khoury College of Computer Science, Northeastern University(东北大学科赫里计算机科学学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Journal ref The 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08833 2025-10-08 cs.CY 79%

Position: The Pitfalls of Over-Alignment: Overly Caution Health-Related Responses From LLMs are Unethical and Dangerous

Wenqi Marshall Guo, Yiyang Du, Heidi J. S. Tworek, Shan Du

专题命中 其他安全 :alignment(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04202 2025-10-07 cs.LG 79%

Spectral Alignment as Predictor of Loss Explosion in Neural Network Training

Haiquan Qiu, You Wu, Yingjie Tan, Yaqing Wang, Quanming Yao

机构 * Tsinghua University(清华大学) Beijing Institute of Mathematical Sciences and Applications(北京数学科学研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04145 2025-10-07 cs.CV cs.CL cs.IR 79%

Automating construction safety inspections using a multi-modal vision-language RAG framework

Chenxin Wang, Elyas Asadi Shamsabadi, Zhaohui Chen, Luming Shen, Alireza Ahmadian Fard Fini, Daniel Dias-da-Costa

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

Comments 33 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01924 2025-10-03 cs.AI cs.MA 79%

To Mask or to Mirror: Human-AI Alignment in Collective Reasoning

Crystal Qian, Aaron Parisi, Clémentine Bouleau, Vivian Tsai, Maël Lebreton, Lucas Dixon

机构 * Google DeepMind(谷歌DeepMind) Paris School of Economics(巴黎经济学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22041 2025-10-02 cs.CL 79%

Taxonomy of Comprehensive Safety for Clinical Agents

Jean Seo, Hyunkyung Lee, Gibaeg Kim, Wooseok Han, Jaehyo Yoo, Seungseop Lim, Kihun Shin, Eunho Yang

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

Comments EMNLP 2025 Industry

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19607 2025-10-01 cs.HC cs.AI 79%

Enabling Rapid Shared Human-AI Mental Model Alignment via the After-Action Review

Edward Gu, Ho Chit Siu, Melanie Platt, Isabelle Hurley, Jaime Peña, Rohan Paleja

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Accepted to the Cooperative Multi-Agent Systems Decision-making and Learning:Human-Multi-Agent Cognitive Fusion Workshop at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24338 2025-09-30 cs.CL 79%

AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignment

Mengyu Bu, Shaolei Zhang, Zhongjun He, Hua Wu, Yang Feng

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS)(智能信息处理重点实验室,计算技术研究所,中国科学院) Key Laboratory of AI Safety, Chinese Academy of Sciences(人工智能安全重点实验室,中国科学院) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) Baidu Inc.(百度公司)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 Main Conference. The code will be available at https://github.com/ictnlp/AlignX

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22503 2025-09-30 cs.RO cs.AI cs.MA 79%

Communication-Efficient Desire Alignment for Embodied Agent-Human Adaptation

Yuanfei Wang, Xinju Huang, Fangwei Zhong, Yaodong Yang, Yizhou Wang, Yuanpei Chen, Hao Dong

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17030 2025-09-29 eess.IV cs.LG 79%

Distillation-Enabled Knowledge Alignment Protocol for Semantic Communication in AI Agent Networks

Jingzhi Hu, Geoffrey Ye Li

机构 * Department of Electrical and Electronic Engineering, Imperial College London(帝国理工学院伦敦分校电子与电气工程系)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Code available at https://github.com/DJ-Duke/DeKAP

Journal ref IEEE Communications Letters, early access, Aug. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21136 2025-09-26 cs.AI 79%

Embodied Representation Alignment with Mirror Neurons

Wentao Zhu, Zhining Zhang, Yuwei Ren, Yin Huang, Hao Xu, Yizhou Wang

机构 * Center on Frontiers of Computing Studies, School of Compter Science, Peking University(前沿计算研究中心,计算机科学学院,北京大学) Eastern Institute of Technology, Ningbo(宁波技术研究所) Qualcomm AI Research(高通人工智能研究) Inst. for Artificial Intelligence, Peking University(人工智能研究所,北京大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19745 2025-09-25 cs.CL cs.SD 79%

PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs

Pei Zhang, Andong Chen, Xi Chen, Baosong Yang, Derek F. Wong, Fei Huang

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团) The Chinese University of Hong Kong(香港中文大学) NLP 2 CT Lab, University of Macau(自然语言处理2CT实验室,澳门大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19329 2025-09-25 cs.CL stat.ME 79%

How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment

Julie Jung, Max Lu, Sina Chole Benker, Dogus Darici

机构 * Harvard Graduate School of Education(哈佛教育研究生院) Munster University(穆恩斯特大学) Institute of Anatomy and Neurobiology, University of Münster(解剖与神经生物学研究所,穆恩斯特大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 9 pages, 4 figures, accepted at NCME AIME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05473 2025-09-25 cs.MM cs.AI cs.SD eess.AS 79%

Embedding Alignment in Code Generation for Audio

Sam Kouteili, Hiren Madhu, George Typaldos, Mark Santolucito

机构 * Yale University(耶鲁大学) Columbia University(哥伦比亚大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Accepted to NeurIPS 2025 AI4Music Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10064 2025-09-24 cs.NE cs.LG q-bio.NC 79%

Dynamical Alignment: A Principle for Adaptive Neural Computation

Xia Chen

机构 * Georg Nemetschek Institute Munich Data Science Institute(慕尼黑数据科学研究所) Technische Universität München(慕尼黑技术大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments 16 pages, 10 figures;

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19018 2025-09-24 cs.LG 79%

OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment

Teng Xiao, Zuchao Li, Lefei Zhang

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18369 2025-09-24 cs.CV cs.AI 79%

Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning

Riad Ahmed Anonto, Sardar Md. Saffat Zabin, M. Saifur Rahman

机构 * Bangladesh University of Engineering and Technology (BUET)(孟加拉工程与技术大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17943 2025-09-23 cs.CV cs.LG 79%

Can multimodal representation learning by alignment preserve modality-specific information?

Romain Thoreau, Jessie Levillain, Dawa Derksen

机构 * institutetext(机构文本) CNES(法国国家空间研究中心) INSA-IMT(法国里尔INSA-IMT)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Accepted as a workshop paper at MACLEAN - ECML/PKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17074 2025-09-23 cs.CV cs.AI 79%

Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models

Qian Zhang, Lin Zhang, Xing Fang, Mingxin Zhang, Zhiyuan Wei, Ran Song, Wei Zhang

机构 * School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Submitted to the IEEE International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14735 2025-09-19 cs.CL 79%

Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM

Chenkun Tan, Pengyu Wang, Shaojun Zhou, Botian Jiang, Zhaowei Li, Dong Zhang, Xinghao Wang, Yaqian Zhou, Xipeng Qiu

机构 * Fudan University(复旦大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted by Findings of EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13846 2025-09-18 cs.CV cs.LG 79%

Consistent View Alignment Improves Foundation Models for 3D Medical Image Segmentation

Puru Vaish, Felix Meister, Tobias Heimann, Christoph Brune, Jelmer M. Wolterink

机构 * Department of Applied Mathematics, Technical Medical Centre, University of Twente(代尔夫特理工大学应用数学系) Digital Technology and Innovation, Siemens Healthineers, Erlangen, Germany(西门子医疗创新部,埃尔朗根,德国)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments MICCAI 2025: 1st Place in Transformer track and 2nd Place in Convolution track of SSL3D-OpenMind challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13204 2025-09-15 cs.CL 79%

Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verification

Jikai Wang, Zhenxu Tian, Juntao Li, Qingrong Xia, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Min Zhang

机构 * Soochow University(苏州大学) Key Laboratory of Data Intelligence and Advanced Computing, Soochow University(数据智能与先进计算关键实验室) Huawei Cloud(华为云)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted at EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09629 2025-09-12 cs.CL 79%

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems

Minghang Zhu, Zhengliang Shi, Zhiwei Xu, Shiguang Wu, Lingjie Wang, Pengjie Ren, Zhaochun Ren, Zhumin Chen

机构 * Shandong University(山东大学) Leiden University(莱顿大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07642 2025-09-10 cs.AI 79%

Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment

Sascha Kaltenpoth, Oliver Müller

机构 * Paderborn University, Department of Business Administration and Economics(帕德博恩大学商业管理与经济学系)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Presented at the 19th International Conference on Wirtschaftsinformatik 2024, Würzburg, Germany https://aisel.aisnet.org/wi2024/91/

详情

展开后加载摘要…

URL PDF HTML 收藏