arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9324 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9324 篇

2504.04994 2025-04-22 cs.CL cs.AI 73%

Following the Whispers of Values: Unraveling Neural Mechanisms Behind Value-Oriented Behaviors in LLMs

Ling Hu, Yuemei Xu, Xiaoyang Gu, Letao Han

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14520 2025-04-22 cs.AI cs.CL 73%

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Muhammad Awais Khan Bangash, Muhammad Ali Jamshed

机构 * University of Oklahoma(俄克拉荷马大学) Stanford University(斯坦福大学) School of Electrical and Computer Engineering, Oklahoma State University(电气与计算机工程学院,俄克拉荷马州立大学) University of Glasgow(格拉斯哥大学)

专题命中 安全评测 :RLHF(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments Submitted to IEEE Transactions on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14448 2025-04-22 cs.AI cs.LG math.OC stat.ME 73%

Seeing Through Risk: A Symbolic Approximation of Prospect Theory

Ali Arslan Yousaf, Umair Rehman, Muhammad Umair Danish

机构 * Tradeweb Markets Western University(西方大学)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07123 2025-04-21 cs.CY cs.LG 73%

Transforming disaster risk reduction with AI and big data: Legal and interdisciplinary perspectives

Kwok P Chun, Thanti Octavianti, Nilay Dogulu, Hristos Tyralis, Georgia Papacharalampous, Ryan Rowberry, Pingyu Fan, Mark Everard, Maria Francesch-Huidobro, Wellington Migliari, David M. Hannah, John Travis Marshall, Rafael Tolosana Calasanz, Chad Staddon, Ida Ansharyani, Bastien Dieppois, Todd R Lewis, Juli Ponce, Silvia Ibrean, Tiago Miguel Ferreira, Chinkie Peliño-Golle, Ye Mu, Manuel Delgado, Elizabeth Silvestre Espinoza, Martin Keulertz, Deepak Gopinath, Cheng Li

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CY、cs.LG

Comments 20 pages, 2 figures

Journal ref Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 15 (2025) e70011

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12491 2025-04-01 cs.AI cs.LG 73%

AI in radiological imaging of soft-tissue and bone tumours: a systematic review evaluating against CLAIM and FUTURE-AI guidelines

Douwe J. Spaanderman, Matthew Marzetti, Xinyi Wan, Andrew F. Scarsbrook, Philip Robinson, Edwin H. G. Oei, Jacob J. Visser, Robert Hemke, Kirsten van Langevelde, David F. Hanff, Geert J. L. H. van Leenders, Cornelis Verhoef, Dirk J. Gruühagen, Wiro J. Niessen, Stefan Klein, Martijn P. A. Starmans

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments 25 pages, 6 figures, 8 supplementary figures

Journal ref eBioMedicine(2025), Volume 114, 105642

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18935 2025-02-27 cs.CL cs.AI 73%

JailBench: A Comprehensive Chinese Security Assessment Benchmark for Large Language Models

Shuyi Liu, Simiao Cui, Haoran Bu, Yuming Shang, Xi Zhang

专题命中 安全评测 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Comments 12 pages, 5 figures, accepted at PAKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06867 2025-02-12 cs.CL cs.AI 73%

Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests

David Noever, Forrest McKee

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00786 2025-01-24 cs.CY cs.AI cs.HC 73%

Whether to trust: the ML leap of faith

Tory Frame, Julian Padget, George Stothart, Elizabeth Coulthard

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

Comments 12 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.04308 2025-01-07 cs.AI cs.CL cs.HC 73%

Personality testing of Large Language Models: Limited temporal stability, but highlighted prosociality

Bojana Bodroza, Bojana M. Dinic, Ljubisa Bojic

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03642 2024-12-17 cs.CL cs.AI cs.HC 73%

Aligning LLMs with Individual Preferences via Interaction

Shujin Wu, May Fung, Cheng Qian, Jeonghwan Kim, Dilek Hakkani-Tur, Heng Ji

专题命中 安全评测 :alignment(abstract);harmlessness(abstract);分类 cs.CL、cs.AI

Comments Accepted to COLING 2025. The code and dataset are made public at https://github.com/ShujinWu-0814/ALOE

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01002 2024-12-17 cs.CL cs.AI 73%

Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries

Zelalem Gero, Chandan Singh, Yiqing Xie, Sheng Zhang, Praveen Subramanian, Paul Vozila, Tristan Naumann, Jianfeng Gao, Hoifung Poon

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments Published in ML4H Findings 2024, 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02512 2024-12-16 cs.CY cs.AI 73%

Pre-Deployment Information Sharing: A Zoning Taxonomy for Precursory Capabilities

Matteo Pistillo, Charlotte Stix

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02145 2024-12-04 cs.CY cs.AI 73%

Effective Mitigations for Systemic Risks from General-Purpose AI

Risto Uuk, Annemieke Brouwer, Tim Schreier, Noemi Dreksler, Valeria Pulignano, Rishi Bommasani

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments 78 pages, 7 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15374 2024-10-22 cs.LG cs.AI cs.CV 73%

Explainability of Point Cloud Neural Networks Using SMILE: Statistical Model-Agnostic Interpretability with Local Explanations

Seyed Mohammad Ahmadi, Koorosh Aslansefat, Ruben Valcarce-Dineiro, Joshua Barnfather

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14281 2024-08-27 cs.LG cs.AI cs.CV 73%

Uncertainties of Latent Representations in Computer Vision

Michael Kirchhof

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments Doctoral thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05919 2024-07-09 cs.LG cs.CY 73%

Fostering Trust and Quantifying Value of AI and ML

Dalmo Cirne, Veena Calambur

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17864 2024-06-27 cs.CY cs.AI 73%

AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies

Yi Zeng, Kevin Klyman, Andy Zhou, Yu Yang, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, Bo Li

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00027 2024-06-11 cs.CR cs.AI cs.CL 73%

Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections

Yuanpu Cao, Bochuan Cao, Jinghui Chen

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07822 2024-06-06 cs.LG cs.AI 73%

Prototypical Self-Explainable Models Without Re-training

Srishti Gautam, Ahcene Boubekki, Marina M. C. Höhne, Michael C. Kampffmeyer

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10893 2024-05-20 cs.CL cs.AI 73%

COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain

Dimitrios P. Panagoulias, Persephone Papatheodosiou, Anastasios P. Palamidas, Mattheos Sanoudos, Evridiki Tsoureli-Nikita, Maria Virvou, George A. Tsihrintzis

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments Technical Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01858 2024-05-06 cs.CL cs.CY 73%

SUKHSANDESH: An Avatar Therapeutic Question Answering Platform for Sexual Education in Rural India

Salam Michael Singh, Shubhmoy Kumar Garg, Amitesh Misra, Aaditeshwar Seth, Tanmoy Chakraborty

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13716 2024-04-19 cs.LG cs.AI 73%

Can I trust my fake data -- A comprehensive quality assessment framework for synthetic tabular data in healthcare

Vibeke Binz Vallevik, Aleksandar Babic, Serena Elizabeth Marshall, Severin Elvatun, Helga Brøgger, Sharmini Alagaratnam, Bjørn Edwin, Narasimha Raghavan Veeraragavan, Anne Kjersti Befring, Jan Franz Nygård

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref Int. J. Med. Inform.185 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14680 2024-04-05 cs.CY cs.AI 73%

Trust in AI: Progress, Challenges, and Future Directions

Saleh Afroogh, Ali Akbari, Evan Malone, Mohammadali Kargar, Hananeh Alambeigi

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02510 2024-04-04 cs.LG cs.AI 73%

An Interpretable Client Decision Tree Aggregation process for Federated Learning

Alberto Argente-Garrido, Cristina Zuheros, M. Victoria Luzón, Francisco Herrera

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments Submitted to Information Science Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13375 2024-03-29 cs.LG cs.AI stat.ML 73%

Optimal Transport Perturbations for Safe Reinforcement Learning with Robustness Guarantees

James Queeney, Erhan Can Ozcan, Ioannis Ch. Paschalidis, Christos G. Cassandras

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments Transactions on Machine Learning Research (TMLR), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15394 2024-03-26 cs.CY cs.LG 73%

"Model Cards for Model Reporting" in 2024: Reclassifying Category of Ethical Considerations in Terms of Trustworthiness and Risk Management

DeBrae Kennedy-Mayo, Jake Gord

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CY、cs.LG

Comments 14 pages, 2 figures, submitted to ACM Conference on Fairness, Accountability, and Transparency 2024 (ACM FAccT '24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09676 2024-03-18 cs.CL cs.AI 73%

Unmasking the Shadows of AI: Investigating Deceptive Capabilities in Large Language Models

Linge Guo

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments AI deception, Large Language Models, ChatGPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16540 2024-03-08 cs.CL cs.LG stat.ML 73%

Unsupervised Pretraining for Fact Verification by Language Model Distillation

Adrián Bazaga, Pietro Liò, Gos Micklem

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.LG

Comments ICLR 2024 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01055 2024-01-15 cs.CL cs.AI 73%

LLaMA Beyond English: An Empirical Study on Language Capability Transfer

Jun Zhao, Zhihao Zhang, Luhui Gao, Qi Zhang, Tao Gui, Xuanjing Huang

专题命中 安全评测 :alignment(abstract);harmlessness(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06674 2023-12-13 cs.CL cs.AI 73%

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, Madian Khabsa

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏