arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9331 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9331 篇

2403.08984 2024-03-15 cs.RO cs.AI cs.MA 70%

Safe Road-Crossing by Autonomous Wheelchairs: a Novel Dataset and its Experimental Evaluation

Carlo Grigioni, Franca Corradini, Alessandro Antonucci, Jérôme Guzzi, Francesco Flammini

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

Comments 14 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10083 2024-02-16 cs.AI 70%

Fine-tuning Large Language Model (LLM) Artificial Intelligence Chatbots in Ophthalmology and LLM-based evaluation using GPT-4

Ting Fang Tan, Kabilan Elangovan, Liyuan Jin, Yao Jie, Li Yong, Joshua Lim, Stanley Poh, Wei Yan Ng, Daniel Lim, Yuhe Ke, Nan Liu, Daniel Shu Wei Ting

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

Comments 13 Pages, 1 Figure, 8 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10534 2023-12-19 cs.LG cs.CR cs.CV 70%

Rethinking Robustness of Model Attributions

Sandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N Balasubramanian

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.LG

Comments Accepted AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10059 2023-12-19 cs.CY 70%

A collection of principles for guiding and evaluating large language models

Konstantin Hebenstreit, Robert Praas, Matthias Samwald

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CY

Comments Accepted at Socially Responsible Language Modelling Research (SoLaR) workshop, NeurIPS 2023 (https://openreview.net/forum?id=8iXdNXW34d). Based on previous manuscript version: doi:10.2139/ssrn.4446991

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05392 2023-12-12 cs.AI 70%

The logic of NTQR evaluations of noisy AI agents: Complete postulates and logically consistent error correlations

Andrés Corrada-Emmanuel

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI

Comments 18 pages, 9 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15888 2023-11-28 cs.CR cs.AI 70%

Towards Adaptive RF Fingerprint-based Authentication of IIoT devices

Emmanuel Lomba, Ricardo Severino, Ana Fernández Vilas

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

Journal ref IEEE 27th International Conference on Emerging Technologies and Factory Automation (ETFA), Stuttgart, Germany, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12004 2023-11-21 cs.LG 70%

Risk-averse Batch Active Inverse Reward Design

Panagiotis Liampas

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.LG

Comments 14 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04813 2023-11-09 cs.CV cs.LG 70%

Be Careful When Evaluating Explanations Regarding Ground Truth

Hubert Baniecki, Maciej Chrabaszcz, Andreas Holzinger, Bastian Pfeifer, Anna Saranti, Przemyslaw Biecek

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.11835 2023-10-30 cs.LG math.AT stat.ML 70%

Topological Parallax: A Geometric Specification for Deep Perception Models

Abraham D. Smith, Michael J. Catanzaro, Gabrielle Angeloro, Nirav Patel, Paul Bendich

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.LG

Comments 18 pages, 6 figures. Preprint submitted to NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12945 2023-10-27 cs.CL 70%

ExplainCPE: A Free-text Explanation Benchmark of Chinese Pharmacist Examination

Dongfang Li, Jindi Yu, Baotian Hu, Zhenran Xu, Min Zhang

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL

Comments EMNLP 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13625 2023-10-23 cs.CY 70%

Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers

Janet Egan, Lennart Heim

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01269 2023-09-06 cs.RO cs.CY cs.HC 70%

Outlining the design space of eXplainable swarm (xSwarm): experts perspective

Mohammad Naiseh, Mohammad D. Soorati, Sarvapali Ramchurn

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CY

Comments In the 16th International Symposium on Distributed Autonomous Robotic Systems 2022, November 28-30, 2022, Montbeliard, France

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01301 2023-07-07 cs.AI quant-ph 70%

Reliable AI: Does the Next Generation Require Quantum Computing?

Aras Bacho, Holger Boche, Gitta Kutyniok

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12173 2023-03-10 cs.CV cs.LG 70%

Towards Human-Interpretable Prototypes for Visual Assessment of Image Classification Models

Poulami Sinhamahapatra, Lena Heidemann, Maureen Monnet, Karsten Roscher

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.LG

Journal ref Proceedings of the 18th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - Volume 5: VISAPP, 878-887, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.10729 2022-11-18 cs.LG cs.CV 70%

Fair Robust Active Learning by Joint Inconsistency

Tsung-Han Wu, Hung-Ting Su, Shang-Tse Chen, Winston H. Hsu

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.LG

Comments 11 pages, 2 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05459 2022-09-13 cs.CY 70%

How Do AI Timelines Affect Existential Risk?

Stephen McAleese

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.01642 2022-09-07 cs.LG 70%

Fraud Detection Using Optimized Machine Learning Tools Under Imbalance Classes

Mary Isangediok, Kelum Gajamannage

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.LG

Comments 10 pages, 10 figures, submitted to IEEE BigData 2022 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.14157 2022-07-29 cs.SE cs.AI 70%

A Hazard Analysis Framework for Code Synthesis Large Language Models

Heidy Khlaaf, Pamela Mishkin, Joshua Achiam, Gretchen Krueger, Miles Brundage

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.11471 2021-12-23 cs.AI cs.CL cs.CY cs.HC cs.LG 70%

Towards a Science of Human-AI Decision Making: A Survey of Empirical Studies

Vivian Lai, Chacha Chen, Q. Vera Liao, Alison Smith-Renner, Chenhao Tan

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 36 pages, 2 figures, see https://haidecisionmaking.github.io for website

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.07445 2021-09-16 cs.CL cs.AI cs.CY cs.LG 70%

Challenges in Detoxifying Language Models

Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, Po-Sen Huang

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 23 pages, 6 figures, published in Findings of EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09928 2026-08-11 cs.CV cs.AI cs.CL cs.LG 新提交 69%

Multimodal Model Diffing for Feature Discovery and Control

用于特征发现与控制的多模态模型差异分析

Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr, Christian Schroeder de Witt, Constantin Venhoff, Ronald Clark

机构 * University of Oxford(牛津大学) Microsoft(微软公司)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG;trustworthy(comments)

AI总结 本研究提出MMDiff多模态模型差异分析框架,训练多模态SAEs以识别多模态训练改变的特征,实现特征隔离、检测与控制,在空间、OCR任务及多模态安全攻击评估中展现出良好效果。

Comments Preprint. Accepted at ICML 2026 Trustworthy AI for Good Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26886 2026-07-30 cs.CV cs.AI cs.CL cs.CY 新提交 69%

Hearsay: Vision-Language Medical Diagnoses Without an Image

Hearsay:无需图像的视觉-语言医学诊断

Siddharth Vohra

机构 * Carnegie Mellon University(卡内基梅隆大学) Amazon Web Services AI Native(亚马逊网络服务AI原生部门)

专题命中 安全评测 :trustworthy(abstract,comments);分类 cs.CL、cs.AI、cs.CY

AI总结 该研究发现前沿视觉-语言模型在无图像时会基于人口统计学特征编造医学诊断,存在不同失败模式,提出需直接审计其结构化输出通道并将探针词敏感性作为核心评估维度。

Comments Peer-reviewed and presented at the 1st Workshop on Toward Trustworthy Vision-Language Models in the Wild (TrustVLM), co-located with ACM ICMR 2026, Amsterdam. Non-archival workshop. Reviews public on OpenReview. 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25258 2026-05-26 cs.IR cs.AI cs.CY cs.LG 69%

First, do no harm: Breaking suicidogenic echo chambers in media recommendation

首先,不伤害:打破媒体推荐中的自杀性回音室

Alberto Díaz-Álvarez, Raúl Lara-Cabrera, Fernando Ortega-Requena, Víctor Ramos-Osuna

机构 * E.T.S.I. Sistemas Informáticos (Universidad Politécnica de Madrid)(马德里理工大学信息系统工程系)

专题命中 安全评测 :safety(abstract,comments);分类 cs.AI、cs.CY、cs.LG

AI总结 针对推荐系统在心理健康场景中可能加剧用户自杀倾向的问题,提出RankAid重排序方法,通过惩罚有害内容并提升治疗性内容,在保持推荐准确性的同时确保临床安全。

Comments 10 pages, 5 figures. Research on safety-aware recommender systems and algorithmic ethics

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16993 2026-05-19 cs.CY cs.AI cs.LG 69%

Adversarial Fragility and Language Vulnerability in Clinical AI: A Systematic Audit of Diagnostic Collapse Under Imperceptible Perturbations and Cross-Lingual Drift in Low-Resource Healthcare Settings

临床AI中的对抗脆弱性与语言脆弱性:在低资源医疗环境中对诊断崩溃的系统审计及不可察觉扰动和跨语言漂移的影响

Anthonio Oladimeji Gabriel, Ahmad Rufai Yusuf

机构 * Centre for Clinical Intelligence & Safety(临床智能与安全中心) Tomorrow University of Applied Sciences(明天应用科学大学)

专题命中 安全评测 :safety(abstract,comments);分类 cs.AI、cs.CY、cs.LG

AI总结 本文系统地审计了临床AI在不可察觉扰动和跨语言漂移下的诊断崩溃问题,揭示了对抗脆弱性和语言脆弱性对低资源医疗环境中的临床AI系统的影响。

Comments 23 pages, 9 figures, 3 tables. Code and data available at https://github.com/anthoniooladimeji11-coder/clinical-ai-safety-audit

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12767 2024-08-23 cs.CL cs.AI cs.LG 69%

Can we trust the evaluation on ChatGPT?

Rachith Aiyappa, Jisun An, Haewoon Kwak, Yong-Yeol Ahn

专题命中 安全评测 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG;trustworthy(journal_ref)

Journal ref Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023) (July 2023) 47-54

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16891 2024-04-29 cs.CR cs.AI cs.CL cs.CY 69%

Attacks on Third-Party APIs of Large Language Models

Wanru Zhao, Vidit Khazanchi, Haodi Xing, Xuanli He, Qiongkai Xu, Nicholas Donald Lane

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY;trustworthy(comments)

Comments ICLR 2024 Workshop on Secure and Trustworthy Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12095 2023-08-30 cs.AI cs.CL cs.LG 69%

On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective

Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, Binxin Jiao, Yue Zhang, Xing Xie

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG;trustworthy(comments)

Comments Highlighted paper at ICLR 2023 workshop on Trustworthy and Reliable Large-Scale Machine Learning Models; code is at: https://github.com/microsoft/robustlearn; more works: https://llm-eval.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.05853 2023-04-13 cs.CL cs.AI cs.CY 69%

Measuring Reliability of Large Language Models through Semantic Consistency

Harsh Raj, Domenic Rosati, Subhabrata Majumdar

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY;safety(comments)

Comments NeurIPS 2022 ML Safety Workshop, https://neurips2022.mlsafety.org

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18671 2026-08-20 cs.CV 新提交 67%

Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI

用于第一人称视角视频的视觉-语言模型:从手-物交互到具身智能

Mohammad Zamani, Fatemeh Ziaeetabar

机构 * University of Tehran(德黑兰大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract)

AI总结 本综述梳理了用于第一人称视角视频理解的视觉-语言模型的发展,分析其挑战、研究方向与局限,明确了可部署具身智能的关键优先方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04448 2026-08-20 cs.AI cs.CL cs.CV cs.LG cs.MA 版本更新 67%

SkillNet: Create, Evaluate, and Connect AI Skills

SkillNet: 创建、评估和连接AI技能

Yuan Liang, Ruobin Zhong, Haoming Xu, Chen Jiang, Yi Zhong, Runnan Fang, Jia-Chen Gu, Shumin Deng, Yunzhi Yao, Mengru Wang, Shuofei Qiao, Yida Xue, Xin Xu, Tongtong Wu, Kun Wang, Yang Liu, Zhen Bi, Jungang Lou, Yuchen Eleanor Jiang, Hangcheng Zhu, Gang Yu, Haiwen Hong, Longtao Huang, Hui Xue, Chenxi Wang, Yijun Wang, Zifei Shan, Xi Chen, Zhaopeng Tu, Feiyu Xiong, Xin Xie, Peng Zhang, Zhengke Gui, Lei Liang, Jun Zhou, Chiyu Wu, Jin Shang, Yu Gong, Junyu Lin, Changliang Xu, Hongjie Deng, Wen Zhang, Keyan Ding, Qiang Zhang, Fei Huang, Ningyu Zhang, Jeff Z. Pan, Guilin Qi, Haofen Wang, Huajun Chen

机构 * Zhejiang University(浙江大学) Tongji University(同济大学) Southeast University(东南大学) Alibaba Group(阿里巴巴集团) Tencent(腾讯) Fudan University(复旦大学) The University of Edinburgh(爱丁堡大学) Monash University(墨尔本大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Huzhou University(湖州大学) Hornor Device Co., Ltd(Hornor设备有限公司) Hangzhou Institute for Advanced Study, UCAS(杭州先进研究所,UCAS)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 SkillNet通过统一的本体和多维评估机制,大规模创建、评估和连接AI技能,提升代理性能并促进技能的持久掌握。

Comments http://skillnet.openkg.cn/; add SkillNet-Gym, a benchmark for evaluating skill retrieval, utilization, composition, and SkillNet-Fabric for task-specific skill routing through lightweight Wikis

详情

展开后加载摘要…

URL PDF HTML 收藏