arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7945 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7945 篇

2409.00091 2024-09-04 cs.CL cs.AI cs.LG 82%

Classification of Safety Events at Nuclear Sites using Large Language Models

Mishca de Costa, Muhammad Anwar, Daniel Lau, Issam Hammad

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Journal ref 43rd Annual CNS Conference and the 48th Annual CNS/CNA Student Conference Sheraton Cavalier Saskatoon Hotel, Saskatoon, SK, Canada, June 16-19, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02429 2024-08-27 cs.IR cs.AI cs.CL cs.LG 82%

CALRec: Contrastive Alignment of Generative LLMs for Sequential Recommendation

Yaoyiran Li, Xiang Zhai, Moustafa Alzantot, Keyi Yu, Ivan Vulić, Anna Korhonen, Mohamed Hammad

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments RecSys 2024 (Long Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04614 2024-08-15 cs.CL cs.AI cs.LG 82%

Better Alignment with Instruction Back-and-Forth Translation

Thao Nguyen, Jeffrey Li, Sewoong Oh, Ludwig Schmidt, Jason Weston, Luke Zettlemoyer, Xian Li

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13229 2024-06-21 cs.CL cs.AI cs.LG 82%

Probing the Emergence of Cross-lingual Alignment during LLM Training

Hetong Wang, Pasquale Minervini, Edoardo M. Ponti

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to Findings of the Association for Computational Linguistics: ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14863 2024-05-24 cs.CL cs.AI cs.LG 82%

A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns

Asaf Yehudai, Taelin Karidi, Gabriel Stanovsky, Ariel Goldstein, Omri Abend

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments CogSci

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00688 2024-05-03 cs.RO cs.AI cs.CL cs.HC cs.LG 82%

Understanding Social Perception, Interactions, and Safety Aspects of Sidewalk Delivery Robots Using Sentiment Analysis

Yuchen Du, Tho V. Le

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 34 pages, 7 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09704 2024-03-18 cs.CL cs.AI cs.LG 82%

Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations

Swapnaja Achintalwar, Ioana Baldini, Djallel Bouneffouf, Joan Byamugisha, Maria Chang, Pierre Dognin, Eitan Farchi, Ndivhuwo Makondo, Aleksandra Mojsilovic, Manish Nagireddy, Karthikeyan Natesan Ramamurthy, Inkit Padhi, Orna Raz, Jesus Rios, Prasanna Sattigeri, Moninder Singh, Siphiwe Thwala, Rosario A. Uceda-Sosa, Kush R. Varshney

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02896 2024-02-06 cs.CL cs.AI cs.CY cs.MA 82%

LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models

Ivar Frisch, Mario Giulianelli

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments To appear in Proceedings of the 1st Personalization of Generative AI Workshop, EACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08379 2023-11-29 cs.CY cs.AI cs.LG 82%

Scheming AIs: Will AIs fake alignment during training in order to get power?

Joe Carlsmith

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY、cs.LG

Comments 127 pages, 8 figures. Revised again to correct typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10111 2023-11-20 cs.CV cs.AI cs.CL cs.LG 82%

VideoCon: Robust Video-Language Alignment via Contrast Captions

Hritik Bansal, Yonatan Bitton, Idan Szpektor, Kai-Wei Chang, Aditya Grover

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 22 pages, 19 Figures, 7 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.07172 2023-09-15 cs.AI cs.CL cs.LG 82%

Exploring Large Language Models for Ontology Alignment

Yuan He, Jiaoyan Chen, Hang Dong, Ian Horrocks

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ISWC 2023 (Posters and Demos)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05382 2023-09-07 cs.CL cs.AI cs.LG 82%

ChatGPT is on the Horizon: Could a Large Language Model be Suitable for Intelligent Traffic Safety Research and Applications?

Ou Zheng, Mohamed Abdel-Aty, Dongdong Wang, Zijin Wang, Shengxuan Ding

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Submitted to Nature - Machine Intelligence (Revised and Extended)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04275 2023-08-09 cs.CL cs.AI cs.LG 82%

In-Context Alignment: Chat with Vanilla Language Models Before Fine-Tuning

Xiaochuang Han

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.15952 2022-06-13 cs.CL cs.AI cs.LG 82%

Knowledge Graph - Deep Learning: A Case Study in Question Answering in Aviation Safety Domain

Ankush Agarwal, Raj Gite, Shreya Laddha, Pushpak Bhattacharyya, Satyanarayan Kar, Asif Ekbal, Prabhjit Thind, Rajesh Zele, Ravi Shankar

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments LREC 2022 Main Conference Accepted Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.04948 2020-06-11 cs.CY cs.AI cs.LG 82%

AI Research Considerations for Human Existential Safety (ARCHES)

Andrew Critch, David Krueger

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.00783 2018-04-24 cs.AI cs.CY cs.LG 82%

Brief Notes on Hard Takeoff, Value Alignment, and Coherent Extrapolated Volition

Gopal P. Sarma

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY、cs.LG

Comments 3 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21311 2026-08-05 cs.CV cs.AI cs.LG 版本更新 82%

Efficient unsupervised domain adaptation via self-supervised vision transformer and synergistic cross-domain alignment

基于自监督视觉Transformer与协同跨域对齐的高效无监督域适应

Ali Abedi, Q. M. Jonathan Wu, Ning Zhang, Farhad Pourpanah

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 提出EUDA框架,以冻结的DINOv2为特征提取器,结合SDAL损失,在多数据集上实现高效无监督域适应,可训练参数减少42%至99.7%,适配资源受限环境。

Comments 22 pages, 4 figures

Journal ref Abedi, A., Wu, Q.M.J., Zhang, N. et al. Efficient unsupervised domain adaptation via self-supervised vision transformer and synergistic cross-domain alignment. Int. J. Mach. Learn. & Cyber. 17, 423 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06966 2025-09-10 eess.SP cs.AI cs.LG 82%

Cross-device Zero-shot Label Transfer via Alignment of Time Series Foundation Model Embeddings

Neal G. Ravindra, Arijit Sehanobish

机构 * Independent Researcher(独立研究者)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 5 pages, 3 figures, 1 table. tl;dr: Adversarial alignment of Time-Series Foundation Model (TSFM) embeddings enables transfer of high-quality clinical labels from medical-grade to consumer-grade wearables, enabling zero-shot prediction of gestational age without requiring paired data

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04231 2025-06-03 cs.MA cs.AI cs.CY cs.GT 82%

Quantifying Misalignment Between Agents: Towards a Sociotechnical Understanding of Alignment

Aidan Kierans, Avijit Ghosh, Hananel Hazan, Shiri Dori-Hacohen

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments 7 pages, 8 figures, 3 tables, forthcoming at the AAAI-25 Special Track on AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05282 2025-04-10 cs.CY cs.AI 82%

International Scientific Report on the Safety of Advanced AI (Interim Report)

Yoshua Bengio, Sören Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Danielle Goldfarb, Hoda Heidari, Leila Khalatbari, Shayne Longpre, Vasilios Mavroudis, Mantas Mazeika, Kwan Yee Ng, Chinasa T. Okolo, Deborah Raji, Theodora Skeadas, Florian Tramèr, Bayo Adekanmbi, Paul Christiano, David Dalrymple, Thomas G. Dietterich, Edward Felten, Pascale Fung, Pierre-Olivier Gourinchas, Nick Jennings, Andreas Krause, Percy Liang, Teresa Ludermir, Vidushi Marda, Helen Margetts, John A. McDermid, Arvind Narayanan, Alondra Nelson, Alice Oh, Gopal Ramchurn, Stuart Russell, Marietje Schaake, Dawn Song, Alvaro Soto, Lee Tiedrich, Gaël Varoquaux, Andrew Yao, Ya-Qin Zhang

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.CY

Comments Available under the open government license at https://www.gov.uk/government/publications/international-scientific-report-on-the-safety-of-advanced-ai

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01252 2024-09-04 cs.CL cs.AI stat.ML 82%

Towards Scalable Automated Alignment of LLMs: A Survey

Boxi Cao, Keming Lu, Xinyu Lu, Jiawei Chen, Mengjie Ren, Hao Xiang, Peilin Liu, Yaojie Lu, Ben He, Xianpei Han, Le Sun, Hongyu Lin, Bowen Yu

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Paper List: https://github.com/cascip/awesome-auto-alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15645 2024-07-23 cs.CL cs.AI 82%

Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models

Joy He-Yueya, Wanjing Anya Ma, Kanishk Gandhi, Benjamin W. Domingue, Emma Brunskill, Noah D. Goodman

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Code and data: https://github.com/joyheyueya/psychometric-alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10635 2023-10-17 cs.LG cs.AI cs.CV 82%

Towards Scenario-based Safety Validation for Autonomous Trains with Deep Generative Models

Thomas Decker, Ananta R. Bhattarai, Michael Lebacher

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments International Conference on Computer Safety, Reliability, and Security 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.03057 2018-12-10 cs.CY cs.LG stat.ML 82%

Open Problems in Engineering and Quality Assurance of Safety Critical Machine Learning Systems

Hiroshi Kuwajima, Hirotoshi Yasuoka, Toshihiro Nakae

专题命中 其他安全 :safety(title,abstract);分类 cs.CY、cs.LG

Comments DISE1: Joint Workshop on Deep (or Machine) Learning for Safety-Critical Applications in Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07748 2026-05-11 cs.CL 81%

TextLDM: Language Modeling with Continuous Latent Diffusion

TextLDM:基于连续潜在扩散的语言建模

Jiaxiu Jiang, Jingjing Ren, Wenbo Li, Bo Wang, Haoze Sun, Yijun Yang, Jianhui Liu, Yanbing Zhang, Shenghe Zheng, Yuan Zhang, Haoyang Huang, Nan Duan, Wangmeng Zuo

机构 * Joy Future Academy(京东探索研究院) HIT(Harbin Institute of Technology) HKUST(GZ)(Hong Kong University of Science and Technology (Guangzhou))

专题命中 其他安全 :alignment(summary_cn,abstract);分类 cs.CL

AI总结 TextLDM将视觉潜在扩散框架应用于文本生成,通过Representation Alignment提升文本表示质量,在OpenWebText2上训练后优于现有扩散语言模型,匹配GPT-2性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15549 2026-05-01 cs.CL 81%

M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning Datasets

M-DaQ:用于指令微调数据集的多语言多样性与质量样本检索

Chunguang Zhao, Yilun Liu, Pufan Zeng, Yuanchang Luo, Shimin Tao, Minggui He, Weibin Meng, Song Xu, Chen Liu, Hongxia Ma, Li Zhang, Boxing Chen, Daimeng Wei

机构 * Huawei Technologies Ltd.(华为技术有限公司) University of Science and Technology of China(中国科学技术大学)

专题命中 其他安全 :alignment(summary_cn,abstract);分类 cs.CL

AI总结 M-DaQ通过联合优化指令-响应质量与跨语言语义多样性,构建高质量平衡训练数据,验证了多语言设置下的Superficial Alignment Hypothesis,并在18种语言上展示出超过60%的胜率。

Comments Accepted by SIGIR 2026 Short

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12372 2026-08-14 cs.AI cs.CY 新提交 81%

Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning

立场:我们需要实用的AI对齐方法来镜像人类推理

Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins, Vincent Conitzer, Walter Sinnott-Armstrong, Jana Schaich Borg

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY

AI总结 该论文提出需实用AI对齐方法镜像人类推理,指出认知对齐可提升AI可理解性与可信赖度,提出研究议程以缩小现有对齐方法与认知对齐需求的差距,助力用户信赖AI系统。

Comments Accepted in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05678 2026-08-14 cs.CL cs.AI 版本更新 81%

Gradual Code-Switching as Inference-Time Cross-Lingual Representational Alignment for LLMs

渐进式代码切换作为大语言模型推理时的跨语言表征对齐

Haneul Yoo, Jiho Jin, Kyunghyun Cho, Alice Oh

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 针对大语言模型跨语言性能不均衡问题,提出推理时的代码切换上下文学习机制,经多模型、多语言、多数据集验证,可显著提升目标及未见语言任务性能,尤其在低资源场景表现突出。

Comments COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11167 2026-08-12 cs.CV cs.CL cs.LG 新提交 81%

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

多模态代码切换:将视觉对象交织入语言以实现显式对象级对齐

Changhao Xiang, Shangyu Xing, Zhen Wu, Jianbing Zhang, Xinyu Dai

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

AI总结 针对现有多模态大语言模型的图像级对齐存在指称歧义的问题,提出多模态代码切换(MMCS)范式,构建含77.3万样本的数据集,仅用5万样本即可匹配或超越60万图像-文本对训练的模型,提升了视觉基础与感知能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06227 2026-08-12 cs.CL cs.LG 版本更新 81%

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

基于开源LLM的自信意识细粒度辩论用于心理健康和在线安全的自动化数据增强

Junyu Mao, Anthony Hills, Talia Tseriotou, Maria Liakata, Aya Shamir, Dan Sayda, Dana Atzil-Slonim, Natalie Djohari, Pamela Ugwudike, Mahesan Niranjan, Stuart E. Middleton

机构 * University of Southampton, UK(英国南安普顿大学) Queen Mary University of London, UK(伦敦大学玛丽女王学院) The Alan Turing Institute, UK(阿兰·图灵研究所) Bar Ilan University, Israel(巴伊兰大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.LG

AI总结 本研究提出CFD框架,利用开源LLM的细粒度辩论实现心理健康和在线安全领域的自动化数据增强,通过实验验证其在多标签增强任务中的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏