arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9331 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9331 篇

2510.17910 2025-10-22 cs.CY cs.AI cs.CL 67%

Interpretability Framework for LLMs in Undergraduate Calculus

Sagnik Dakshit, Sushmita Sinha Roy

机构 * University of Texas at Tyler(德克萨斯理工大学) Florida Gulf Coast University(佛罗里达盖恩斯维尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15891 2025-10-21 cs.HC cs.AI cs.CL cs.LG 67%

Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System

Ziv Ben-Zion, Paul Raffelhüschen, Max Zettl, Antonia Lüönd, Achim Burrer, Philipp Homan, Tobias R Spiller

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15866 2025-10-20 cs.CV cs.NE 67%

BiomedXPro: Prompt Optimization for Explainable Diagnosis with Biomedical Vision Language Models

Kaushitha Silva, Mansitha Eashwara, Sanduni Ubayasiri, Ruwan Tennakoon, Damayanthi Herath

机构 * University of Peradeniya(珀德尼亚大学) RMIT University(皇家墨尔本理工大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract)

Comments 10 Pages + 15 Supplementary Material Pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21603 2025-10-20 cs.CL cs.CY cs.LG 67%

Operationalizing Automated Essay Scoring: A Human-Aware Approach

Yenisel Plasencia-Calaña

机构 * Brigthlands Institute for a Smart Society(智能社会研究院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04426 2025-10-17 cs.CL cs.AI cs.CY 67%

The simulation of judgment in LLMs

Edoardo Loru, Jacopo Nudo, Niccolò Di Marco, Alessandro Santirocchi, Roberto Atzeni, Matteo Cinelli, Vincenzo Cestari, Clelia Rossi-Arnaud, Walter Quattrociocchi

机构 * Department of Computer, Control and Management Engineering, Sapienza University of Rome(计算机、控制与管理工程系,罗马萨皮恩扎大学) Department of Computer Science, Sapienza University of Rome(计算机科学系,罗马萨皮恩扎大学) Department of Legal, Social, and Educational Sciences, Tuscia University(法律、社会与教育科学系,图斯西亚大学) Department of Psychology, Sapienza University of Rome(心理学系,罗马萨皮恩扎大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Please refer to published version: https://doi.org/10.1073/pnas.2518443122

Journal ref Proc. Natl. Acad. Sci. U.S.A. 122 (42) e2518443122, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19195 2025-10-10 cs.CL cs.AI cs.LG 67%

Can Small-Scale Data Poisoning Exacerbate Dialect-Linked Biases in Large Language Models?

Chaymaa Abbas, Mariette Awad, Razane Tajeddine

机构 * Department of Electrical and Computer Engineering, Maroun Semaan Faculty of Engineering and Architecture(电气与计算机工程系,马鲁恩·塞马安工程与建筑学院) American University of Beirut(贝鲁特美国大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22203 2025-10-08 cs.LG cs.AI cs.CL 67%

From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning

Yuzhen Huang, Weihao Zeng, Xingshan Zeng, Qi Zhu, Junxian He

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03688 2025-10-07 cs.AI cs.CL cs.LG 67%

AgentBench: Evaluating LLMs as Agents

Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, Jie Tang

机构 * Tsinghua University(清华大学) The Ohio State University(俄亥俄州立大学) UC Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published in ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00038 2025-10-06 cs.LG cs.AI cs.CY 67%

DM-Bench: Benchmarking LLMs for Personalized Decision Making in Diabetes Management

Maria Ana Cardei, Josephine Lamp, Mark Derdzinski, Karan Bhatia

机构 * University of Virginia(弗吉尼亚大学) Dexcom(德科姆公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01691 2025-10-03 cs.CV 67%

MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs

Jiyao Liu, Jinjie Wei, Wanying Qu, Chenglong Ma, Junzhi Ning, Yunheng Li, Ying Chen, Xinzhe Luo, Pengcheng Chen, Xin Gao, Ming Hu, Huihui Xu, Xin Wang, Shujian Gao, Dingkang Yang, Zhongying Deng, Jin Ye, Lihao Liu, Junjun He, Ningsheng Xu

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Imperial College London(帝国理工学院) University of Cambridge(剑桥大学)

专题命中 安全评测 :alignment(abstract);safety(abstract)

Comments 26 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03748 2025-10-02 cs.LG cs.AI cs.CL 67%

TDBench: A Benchmark for Top-Down Image Understanding with Reliability Analysis of Vision-Language Models

Kaiyuan Hou, Minghui Zhao, Lilin Xu, Yuang Fan, Xiaofan Jiang

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22376 2025-09-30 cs.LG cs.AI cs.CL cs.CV 67%

Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance

Dongmin Park, Sebin Kim, Taehong Moon, Minkyu Kim, Kangwook Lee, Jaewoong Cho

机构 * KRAFTON Seoul National University(首尔国立大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICLR 2025 (spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14279 2025-09-30 cs.CL cs.CY cs.LG 67%

GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs

Adrian-Marius Dumitran, Alexandra-Mihaela Danila, Angela-Liliana Dumitran

机构 * University of Bucharest Faculty of Mathematics and Computer Science(布加勒斯大学数学与计算机科学学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.CY、cs.LG

Comments Accepted as long paper @RANLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05418 2025-09-29 cs.CL cs.AI cs.LG 67%

Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning

Jaedong Hwang, Kumar Tanmay, Seok-Jin Lee, Ayush Agrawal, Hamid Palangi, Kumar Ayush, Ila Fiete, Paul Pu Liang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14944 2025-09-24 cs.HC 67%

LEKIA: Expert-Aligned AI Behavior Design for High-Risk Human-AI Interactions

Boning Zhao, Yutong Hu, Xinnuo Li

专题命中 安全评测 :alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14199 2025-09-19 cs.CV cs.AI cs.CL cs.LG 67%

Dense Video Understanding with Gated Residual Tokenization

Haichao Zhang, Wenhao Chai, Shwai He, Ang Li, Yun Fu

机构 * Northeastern University(东北大学) Princeton University(普林斯顿大学) University of Maryland(马里兰大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23804 2025-09-18 cs.CL cs.AI cs.LG 67%

Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequencies

Terrance Liu, Shuyi Wang, Daniel Preotiuc-Pietro, Yash Chandarana, Chirag Gupta

机构 * Carnegie Mellon University(卡内基梅隆大学) Bloomberg(彭博)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09397 2025-09-12 cs.CV 67%

Decoupling Clinical and Class-Agnostic Features for Reliable Few-Shot Adaptation under Shift

Umaima Rahman, Raza Imam, Mohammad Yaqub, Dwarikanath Mahapatra

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德人工智能大学) Khalifa University(卡利法大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08803 2025-09-11 cs.SI cs.AI cs.CL cs.CY 67%

Scaling Truth: The Confidence Paradox in AI Fact-Checking

Ihsan A. Qazi, Zohaib Khan, Abdullah Ghani, Agha A. Raza, Zafar A. Qazi, Wassay Sajjad, Ayesha Ali, Asher Javaid, Muhammad Abdullah Sohail, Abdul H. Azeemi

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 65 pages, 26 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18933 2025-09-10 cs.AI cs.CR cs.CY cs.LG 67%

VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation

David Egea, Barproda Halder, Sanghamitra Dutta

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07017 2025-09-10 cs.AI cs.CL cs.LG 67%

From Eigenmodes to Proofs: Integrating Graph Spectral Operators with Symbolic Interpretable Reasoning

Andrew Kiruluta, Priscilla Burity

机构 * Andrew Kiruluta and Priscilla Burity(独立研究者)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04657 2025-09-08 cs.CL cs.AI cs.DB cs.LG 67%

Evaluating NL2SQL via SQL2NL

Mohammadtaher Safarzadeh, Afshin Oroojlooyjadid, Dan Roth

机构 * Oracle AI

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03986 2025-09-05 cs.CV cs.AI cs.CL cs.LG 67%

Promptception: How Sensitive Are Large Multimodal Models to Prompts?

Mohamed Insaf Ismithdeen, Muhammad Uzair Khattak, Salman Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学) Swiss Federal Institute of Technology Lausanne (EPFL)(洛桑联邦理工学院) Australian National University(澳大利亚国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00806 2025-09-03 cs.CL cs.AI cs.LG 67%

CaresAI at BioCreative IX Track 1 -- LLM for Biomedical QA

Reem Abdel-Salam, Mary Adewunmi, Modinat A. Abayomi

机构 * Cairo University, Faculty of Engineering, Computer Engineering Department(开罗大学工程学院计算机工程系) Menzies School of Health Research, Charles Darwin University, NT, Australia(梅恩兹健康研究学院,查尔斯达尔文大学,澳大利亚NT) Department of Biology, Boston College, Massachusetts, USA(生物学系,波士顿学院,马萨诸塞州,美国) Peoples' Friendship University of Russia (RUDN University)(俄罗斯人民友谊大学(RUDN大学)) Joint Institute for Nuclear Research(联合核研究中心) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) University of Skövde(斯德哥尔摩大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Proceedings of the BioCreative IX Challenge and Workshop (BC9): Large Language Models for Clinical and Biomedical NLP at the International Joint Conference on Artificial Intelligence (IJCAI), Montreal, Canada, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21271 2025-09-01 cs.RO cs.CV 67%

Mini Autonomous Car Driving based on 3D Convolutional Neural Networks

Pablo Moraes, Monica Rodriguez, Kristofer S. Kappel, Hiago Sodre, Santiago Fernandez, Igor Nunes, Bruna Guterres, Ricardo Grando

机构 * Technological University of Uruguay, UTEC, Uruguay(乌拉圭技术大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14377 2025-08-26 cs.CL cs.AI cs.CY 67%

ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities

Wenhan Dong, Zhen Sun, Yuemeng Zhao, Zifan Peng, Jun Wu, Jingyi Zheng, Yule Liu, Xinlei He, Yu Wang, Ruiming Wang, Xinyi Huang, Lei Mo

机构 * School of Psychology, South China Normal University(南方科技大学心理学院) Information Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心) School of AI, Guangzhou University(广州大学人工智能学院) College of Cyber Security, Jinan University(济南大学网络安全学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15797 2025-08-25 cs.CL cs.AI cs.LG 67%

Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks

Nouar AlDahoul, Yasir Zaki

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15418 2025-08-22 cs.CL cs.AI cs.LG cs.MM cs.SD 67%

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

Yirong Sun, Yizhong Geng, Peidong Wei, Yanjun Chen, Jinghan Yang, Rongfei Chen, Wei Zhang, Xiaoyu Shen

机构 * Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT) Logic Intelligence Technology(逻辑智能技术) BUPT(北京邮电大学) Xiamen University(厦门大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15085 2025-08-22 cs.CL cs.AI cs.IR cs.LG 67%

LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text

MohamamdJavad Ardestani, Ehsan Kamalloo, Davood Rafiei

机构 * Department of Computing Science University of Alberta(计算科学系阿尔伯塔大学) ServiceNow Research(ServiceNow研究)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10028 2025-08-15 cs.CL cs.AI cs.HC cs.LG 67%

PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs

Xiao Fu, Hossein A. Rahmani, Bin Wu, Jerome Ramos, Emine Yilmaz, Aldo Lipani

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏