arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9324 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9324 篇

2505.17114 2025-09-08 cs.CL cs.CV cs.LG cs.MM 76%

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language

Subrata Biswas, Mohammad Nur Hossain Khan, Bashima Islam

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24671 2025-09-03 cs.CL cs.AI 76%

Multiple LLM Agents Debate for Equitable Cultural Alignment

Dayeon Ki, Rachel Rudinger, Tianyi Zhou, Marine Carpuat

机构 * University of Maryland(马里兰大学)

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

Comments ACL 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22940 2025-08-05 cs.CL cs.AI 76%

Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes

Rui Jiao, Yue Zhang, Jinku Li

机构 * School of Cyber Engineering, Xidian University(西安电子科技大学电子工程学院) School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院)

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17010 2025-07-24 cs.CR cs.AI cs.LG 76%

Towards Trustworthy AI: Secure Deepfake Detection using CNNs and Zero-Knowledge Proofs

H M Mohaimanul Islam, Huynh Q. N. Vo, Aditya Rane

机构 * School of Industrial Engineering and Management(工业工程与管理学院)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments Submitted for peer-review in TrustXR - 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13524 2025-07-21 cs.HC cs.AI cs.CY 76%

Humans learn to prefer trustworthy AI over human partners

Yaomin Jiang, Levin Brinkmann, Anne-Marie Nussberger, Ivan Soraperra, Jean-François Bonnefon, Iyad Rahwan

机构 * Toulouse School of Economics, Centre National de la Recherche Scientifique (TSM-R), Université Toulouse Capitole(图卢兹经济学院,法国国家科学研究中心(TSM-R),图卢兹大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10240 2025-07-15 cs.HC cs.AI cs.LG 76%

Visual Analytics for Explainable and Trustworthy Artificial Intelligence

Angelos Chatzimparmpas

机构 * Utrecht University(乌特勒支大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Journal ref IEEE CG&A 2025, vol. 45, pp. 100-111

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07576 2025-07-11 cs.AI cs.LG cs.LO 76%

On Trustworthy Rule-Based Models and Explanations

Mohamed Siala, Jordi Planes, Joao Marques-Silva

机构 * LAAS-CNRS, Université de Toulouse, CNRS, INSA Toulouse, France(法国图卢兹大学、CNRS、INSA图卢兹分校) Universitat de Lleida(莱里达大学) ICREA & University of Lleida(ICREA与莱里达大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01281 2025-07-03 cs.CL cs.AI 76%

Rethinking All Evidence: Enhancing Trustworthy Retrieval-Augmented Generation via Conflict-Driven Summarization

Juan Chen, Baolong Bi, Wei Zhang, Jingyan Sui, Xiaofei Zhu, Yuanzhuo Wang, Lingrui Mei, Shenghua Liu

机构 * University of Chinese Academy of Sciences(中国科学院大学) Chinese Academy of Sciences(中国科学院) National University of Defense Technology(国防科技大学) Chongqing University of Technology(重庆理工大学)

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02080 2025-06-12 cs.CV cs.CL cs.LG 76%

EMMA: Efficient Visual Alignment in Multi-Modal LLMs

Sara Ghazanfari, Alexandre Araujo, Prashanth Krishnamurthy, Siddharth Garg, Farshad Khorrami

机构 * Department of Electronic and Computer Engineering, New York University(电子与计算机工程系,纽约大学)

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00519 2025-06-04 cs.CL cs.AI 76%

CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention

Yuxi Sun, Aoqi Zuo, Wei Gao, Jing Ma

机构 * Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系) School of Mathematics and Statistics, The University of Melbourne(墨尔本大学数学与统计学学院) School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算与信息系统学院)

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI

Comments Accepted to Association for Computational Linguistics Findings (ACL) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02017 2025-06-02 cs.SI cs.AI cs.LG 76%

Multi-Domain Graph Foundation Models: Robust Knowledge Transfer via Topology Alignment

Shuo Wang, Bokui Wang, Zhixiang Shen, Boyan Deng, Zhao Kang

专题命中 安全评测 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16103 2025-05-23 cs.LG cs.AI 76%

Towards Trustworthy Keylogger detection: A Comprehensive Analysis of Ensemble Techniques and Feature Selections through Explainable AI

Monirul Islam Mahmud

机构 * Dept. of Computer & Information Science(计算机与信息科学系) Fordham University(福特汉姆大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13028 2025-05-21 cs.CR cs.AI cs.CL 76%

Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset

Sayon Palit, Daniel Woods

机构 * School of Informatics University of Edinburgh(信息学院爱丁堡大学)

专题命中 安全评测 :safety(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17017 2025-03-17 cs.CL cs.IT cs.LG math.IT math.LO 76%

Quantifying Logical Consistency in Transformers via Query-Key Alignment

Eduard Tulchinskii, Anastasia Voznyuk, Laida Kushnareva, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07212 2025-02-27 cs.CL cs.AI cs.HC 76%

Trustworthy and Practical AI for Healthcare: A Guided Deferral System with Large Language Models

Joshua Strong, Qianhui Men, Alison Noble

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI

Comments AAAI-AISI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08353 2025-02-13 cs.LG cs.AI 76%

Trustworthy GNNs with LLMs: A Systematic Review and Taxonomy

Ruizhan Xue, Huimin Deng, Fang He, Maojun Wang, Zeyu Zhang

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments Submitted to IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03467 2025-02-07 cs.CY cs.AI cs.SE 76%

Where AI Assurance Might Go Wrong: Initial lessons from engineering of critical systems

Robin Bloomfield, John Rushby

专题命中 安全评测 :safety(abstract,comments);AI safety(abstract,comments);分类 cs.AI、cs.CY

Comments Presented at UK AI Safety Institute (AISI) Conference on Frontier AI Safety Frameworks (FAISC 24), Berkeley CA, November 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11932 2025-01-07 cs.CL cs.AI 76%

CrossIn: An Efficient Instruction Tuning Approach for Cross-Lingual Knowledge Alignment

Geyu Lin, Bin Wang, Zhengyuan Liu, Nancy F. Chen

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07429 2024-12-11 cs.CL cs.AI 76%

Optimizing Alignment with Less: Leveraging Data Augmentation for Personalized Evaluation

Javad Seraj, Mohammad Mahdi Mohajeri, Mohammad Javad Dousti, Majid Nili Ahmadabadi

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06477 2024-11-06 cs.CL cs.AI 76%

Kun: Answer Polishment for Chinese Self-Alignment with Instruction Back-Translation

Tianyu Zheng, Shuyue Guo, Xingwei Qu, Jiawei Guo, Xinrun Du, Qi Jia, Chenghua Lin, Wenhao Huang, Jie Fu, Ge Zhang

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

Comments 12 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15454 2024-09-25 cs.CL cs.AI 76%

In-Context Learning May Not Elicit Trustworthy Reasoning: A-Not-B Errors in Pretrained Language Models

Pengrui Han, Peiyang Song, Haofei Yu, Jiaxuan You

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01596 2024-08-06 cs.LG cs.AI cs.GT 76%

Trustworthy Machine Learning under Social and Adversarial Data Sources

Han Shao

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18928 2024-07-30 cs.CY cs.AI 76%

The Voice: Lessons on Trustworthy Conversational Agents from "Dune"

Philip Feldman

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.CY

Comments 5 pages, 2 figures

Journal ref ACM Conversational User Interfaces 2024 (CUI '24), July 8--10, 2024, Luxembourg, Luxembourg

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07921 2024-07-12 cs.CR cs.AI cs.LG eess.SP 76%

A Trustworthy AIoT-enabled Localization System via Federated Learning and Blockchain

Junfei Wang, He Huang, Jingze Feng, Steven Wong, Lihua Xie, Jianfei Yang

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04766 2024-07-12 cs.CL cs.AI 76%

SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning

Bin Wang, Zhengyuan Liu, Xin Huang, Fangkai Jiao, Yang Ding, AiTi Aw, Nancy F. Chen

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

Comments Published at NAACL 2024. Code: https://seaeval.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16424 2024-05-28 cs.HC cs.AI cs.LG 76%

Improving Health Professionals' Onboarding with AI and XAI for Trustworthy Human-AI Collaborative Decision Making

Min Hun Lee, Silvana Xin Yi Choo, Shamala D/O Thilarajah

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04343 2023-12-08 cs.LG cs.AI 76%

Causality and Explainability for Trustworthy Integrated Pest Management

Ilias Tsoumas, Vasileios Sitokonstantinou, Georgios Giannarakis, Evagelia Lampiri, Christos Athanassiou, Gustau Camps-Valls, Charalampos Kontoes, Ioannis Athanasiadis

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments Accepted at NeurIPS 2023 Workshop on Tackling Climate Change with Machine Learning: Blending New and Existing Knowledge Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05470 2023-12-08 cs.CL cs.AI 76%

Generative Judge for Evaluating Alignment

Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, Pengfei Liu

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

Comments Fix typos in Table 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00848 2023-10-11 cs.SE cs.AI cs.LG cs.PL 76%

Toward Trustworthy Neural Program Synthesis

Darren Key, Wen-Ding Li, Kevin Ellis

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04760 2023-09-12 cs.LG cs.AI cs.CV 76%

RR-CP: Reliable-Region-Based Conformal Prediction for Trustworthy Medical Image Classification

Yizhe Zhang, Shuo Wang, Yejia Zhang, Danny Z. Chen

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments UNSURE2023 (Uncertainty for Safe Utilization of Machine Learning in Medical Imaging) at MICCAI2023; Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏