arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9324 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9324 篇

2305.16822 2023-10-24 cs.LG cs.DC cs.SE 74%

Rethinking Certification for Trustworthy Machine Learning-Based Applications

Marco Anisetti, Claudio A. Ardagna, Nicola Bena, Ernesto Damiani

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Accepted in IEEE Internet Computing; 6 pages, 1 figure, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10147 2023-07-26 cs.IT cs.LG math.IT 74%

TEFL: Turbo Explainable Federated Learning for 6G Trustworthy Zero-Touch Network Slicing

Swastika Roy, Hatim Chergui, Christos Verikoukis

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Overlapes with the new version

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01902 2023-05-23 cs.LG stat.ML 74%

Barycentric-alignment and reconstruction loss minimization for domain generalization

Boyang Lyu, Thuan Nguyen, Prakash Ishwar, Matthias Scheutz, Shuchin Aeron

专题命中 安全评测 :alignment(title);分类 cs.LG

Comments This article has been accepted for publication in IEEE Access

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11597 2023-05-22 cs.AI 74%

Flexible and Inherently Comprehensible Knowledge Representation for Data-Efficient Learning and Trustworthy Human-Machine Teaming in Manufacturing Environments

Vedran Galetić, Alistair Nottle

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.01204 2023-04-05 cs.AI 74%

Automatic Geo-alignment of Artwork in Children's Story Books

Jakub J. Dylag, Victor Suarez, James Wald, Aneesha Amodini Uvara

专题命中 安全评测 :alignment(title);分类 cs.AI

Comments Master's project

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06045 2022-12-13 cs.LG 74%

PERFEX: Classifier Performance Explanations for Trustworthy AI Systems

Erwin Walraven, Ajaya Adhikari, Cor J. Veenman

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.00572 2022-10-04 cs.CL 74%

Risk-graded Safety for Handling Medical Queries in Conversational AI

Gavin Abercrombie, Verena Rieser

专题命中 安全评测 :safety(title);分类 cs.CL

Comments Accepted for publication at AACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05257 2022-09-13 cs.HC cs.LG 74%

TruVR: Trustworthy Cybersickness Detection using Explainable Machine Learning

Ripan Kumar Kundu, Rifatul Islam, Prasad Calyam, Khaza Anuarul Hoque

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Accepted copy, to be published in ISAMR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.04608 2022-08-10 cs.IR cs.AI 74%

Using Sentence Embeddings and Semantic Similarity for Seeking Consensus when Assessing Trustworthy AI

Dennis Vetter, Jesmin Jahan Tithi, Magnus Westerlund, Roberto V. Zicari, Gemma Roig

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.12630 2022-05-26 cs.CL cs.CV 74%

Multimodal Knowledge Alignment with Reinforcement Learning

Youngjae Yu, Jiwan Chung, Heeseung Yun, Jack Hessel, JaeSung Park, Ximing Lu, Prithviraj Ammanabrolu, Rowan Zellers, Ronan Le Bras, Gunhee Kim, Yejin Choi

专题命中 安全评测 :alignment(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.04816 2022-05-13 cs.CR cs.LG 74%

Towards a trustworthy, secure and reliable enclave for machine learning in a hospital setting: The Essen Medical Computing Platform (EMCP)

Hendrik F. R. Schmidt, Jörg Schlötterer, Marcel Bargull, Enrico Nasca, Ryan Aydelott, Christin Seifert, Folker Meyer

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments 9 pages, 5 figures, to be published in the proceedings of the 2021 IEEE CogMI conference. Christin Seifert and Folker Meyer are co-senior authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.15112 2021-10-01 cs.LG cs.CR 74%

Interpretability in Safety-Critical FinancialTrading Systems

Gabriel Deza, Adelin Travers, Colin Rowat, Nicolas Papernot

专题命中 安全评测 :safety(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.08258 2021-06-16 cs.CY 74%

Identifying Roles, Requirements and Responsibilities in Trustworthy AI Systems

Iain Barclay, Will Abramson

专题命中 安全评测 :trustworthy(title);分类 cs.CY

Comments Pre-print of paper submitted to Workshop on Reviewable and Auditable Pervasive Systems (WRAPS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.08359 2021-03-16 cs.LG 74%

Explaining Credit Risk Scoring through Feature Contribution Alignment with Expert Risk Analysts

Ayoub El Qadi, Natalia Diaz-Rodriguez, Maria Trocan, Thomas Frossard

专题命中 安全评测 :alignment(title);分类 cs.LG

Comments 11 pages, 4, figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.01284 2020-07-02 cs.LG stat.ML 74%

Independent Component Analysis for Trustworthy Cyberspace during High Impact Events: An Application to Covid-19

Zois Boukouvalas, Christine Mallinson, Evan Crothers, Nathalie Japkowicz, Aritran Piplai, Sudip Mittal, Anupam Joshi, Tülay Adalı

专题命中 安全评测 :trustworthy(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.03618 2020-06-09 cs.LG cs.RO stat.ML 74%

Efficient Black-box Assessment of Autonomous Vehicle Safety

Justin Norden, Matthew O'Kelly, Aman Sinha

专题命中 安全评测 :safety(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.05895 2019-05-16 cs.LG cs.CV stat.ML 74%

Addressing the Loss-Metric Mismatch with Adaptive Loss Alignment

Chen Huang, Shuangfei Zhai, Walter Talbott, Miguel Angel Bautista, Shih-Yu Sun, Carlos Guestrin, Josh Susskind

专题命中 安全评测 :alignment(title);分类 cs.LG

Comments Accepted to ICML 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.02471 2018-12-07 cs.AI 74%

The Role of Normware in Trustworthy and Explainable AI

Giovanni Sileno, Alexander Boer, Tom van Engers

专题命中 安全评测 :trustworthy(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1401.1123 2014-01-07 cs.LG 74%

Exploration vs Exploitation vs Safety: Risk-averse Multi-Armed Bandits

Nicolas Galichet, Michèle Sebag, Olivier Teytaud

专题命中 安全评测 :safety(title);分类 cs.LG

Comments 16 pages

Journal ref Asian Conference on Machine Learning 2013, Canberra : Australia (2013)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20379 2026-08-20 cs.AI cs.CL 版本更新 73%

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

训练模型,而非读者:可验证激活解释的可解码性监督

Hiskias Dingeto

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

AI总结 研究自然语言自动编码器对隐藏激活解释评分的问题,提出RECAP等方法,能使指定内容可解码,提高解释忠实度与安全性,如在预训练模型上实现可靠探测解码,提升对真实声明评分及识别谎言能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16894 2026-08-19 cs.CY cs.CL 新提交 73%

An Investigation of the NeurIPS and ICML 2025 Position Tracks

NeurIPS与ICML 2025立场论文 track 的调查

Fan Yang, Wenkai Li, Jun Liu

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.CY

AI总结 本文审计了NeurIPS和ICML 2025立场论文 track 的公开投稿,发现其以改革派批判为主,建议增设方向设定类工作,并提出四项CFP层面干预措施以优化投稿结构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16539 2026-08-18 cs.SD cs.AI cs.CL eess.AS 新提交 73%

Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization

听、推理与分段:使大型音频语言模型(LALMs)与编辑判断对齐以实现媒体章节化

Tony Alex, Wish Suharitdamrong, Sara Atito, Armin Mustafa, Muhammad Awais, Philip J. B. Jackson, Jiankang Deng, Ismail Elezi

机构 * University of Surrey(萨里大学)

专题命中 安全评测 :alignment(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 该研究针对LALMs在实际媒体章节化部署的不足,提出基于GRPO与CoT的AudioChaps框架,构建三类数据集,使AudioChaps-R1的平均F1较SOTA提升49个百分点,实现非结构化听觉流到结构化媒体的可靠转换。

Comments 19 pages, 9 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14577 2026-08-18 cs.CL cs.AI 新提交 73%

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

HarmProfile:刻画前沿大语言模型中的有害分布

Zhouyuan Ma, Yutao Wu, Hanxun Huang, Xiang Zheng, Xiao Liu, Yixin Cao, Zuxuan Wu, Xingjun Ma, Yu-Gang Jiang

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出HarmProfile基准数据集,收集23个前沿LLMs的超8万条有害输出,发现其有害性与多样性随模型能力增长,潜藏对齐表面下的危险知识。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08296 2026-08-17 cs.AI cs.LG 版本更新 73%

Revisiting the shutdown problem

重新审视关机问题

David Thorstad

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 本文重新评估了AI关机问题的难度,指出现有论证未能证明其难以解决,且相关技术方案对模型性能造成了高安全代价。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10300 2026-08-13 cs.AI cs.CR cs.LG 版本更新 73%

Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability

Logit边界几何置信接口与稀疏层束 enclaves 协议:安全网络电子健康记录(EHR)互操作性的自包含底层架构

Alvin Spivey, Yu Huang

机构 * Light Imaging Technologies, Inc.(光成像技术公司)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 该研究提出Logit边界几何置信接口与稀疏层束 enclaves 协议,构建EHR互操作的自包含底层架构,经GBI BoundaryBench v0.1评估,Qwen3-4B-Instruct-2507在该接口下无符合要求的输出,为接受边界提供实证证据。

Comments 39 pages; executable Julia verification code included as ancillary material; companion public benchmark: https://github.com/AlvinSpivey/GBI-BoundaryBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23146 2026-08-12 cs.LG cs.AI 版本更新 73%

Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness

Infra-Bayesian 强化学习智能体在最坏情况鲁棒性上优于经典强化学习

Manish Aryal, Faiyaz Azam, Agnivo Banerjee, Syed Mahir Ahamed, Sai Sidhanth Manoharan Jayanthi, Allegra Laro, Clément Legentilhomme, Andrew Lin, Florian Lorkowski, Marina Pérez del Valle, Radman Rakhshandehroo, Patric Rommel, Emanuel Ruzak, Nathan Theng, Paul Yushin Rapoport

机构 * Purdue University(普渡大学) Carnegie Mellon University(卡内基梅隆大学) WorldQuant University(WorldQuant大学) UC Berkeley(加州大学伯克利分校) Aix-Marseille University(阿维尼翁-马赛大学) MIT(麻省理工学院) University of Zurich(苏黎世大学) University of British Columbia(不列颠哥伦比亚大学) University of Stuttgart(斯图加特大学) University of Buenos Aires(布宜诺斯艾利斯大学) California State University, Fresno(弗雷斯诺加州州立大学) University of Chicago(芝加哥大学)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

AI总结 针对经典强化学习在模型误设和策略依赖不确定性下的失败,提出 Infra-Bayesian 框架,通过区分概率不确定性和 Knightian 不确定性并采用最坏情况决策,在有限状态无状态决策问题中实现更低的最坏情况遗憾。

Comments RLC 2026: Accepted to Finding the Frame Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10015 2026-08-11 cs.AI cs.CE cs.CL cs.MM 版本更新 73%

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

FinTrace: LLM工具调用在长周期金融任务中的轨迹级综合评估

Yupeng Cao, Haohang Li, Weijin Liu, Wenbo Cao, Anke Xu, Lingfei Qian, Xueqing Peng, Minxue Tang, Zhiyuan Yao, Jimin Huang, K. P. Subbalakshmi, Zining Zhu, Jordan W. Suchow, Yangyang Yu

机构 * Stevens Institute of Technology(斯蒂文斯理工学院) Independent Researcher(独立研究者) The FinAI(FinAI) Duke University(杜克大学)

专题命中 安全评测 :DPO(abstract,abstract_cn);分类 cs.CL、cs.AI

AI总结 本文提出FinTrace基准,通过800个专家标注的轨迹评估LLM工具调用,揭示模型在信息利用和最终答案质量上的不足,并通过FinTrace-Training数据集提升中间推理能力,但最终答案质量仍受限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06105 2026-08-07 cs.LG cs.AI cs.MA 新提交 73%

Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping

潜在上下文是否有帮助?北极航运中逆强化学习的受控评估

Vaishnav Vaidheeswaran, Dilith Jayakody, Biruk Ambaw, Jaswanth Kumar, Md Mahbub Alam, Gabriel Spadon

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本研究通过对北极航运航次的受控评估,发现非线性共享奖励模型优于线性基线,添加船舶特定潜在上下文反而降低性能,表明可观测特征已能解释行为差异,为安全关键领域的AI部署提供支持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06020 2026-08-07 cs.AI cs.LG 新提交 73%

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

从经济智能体到智能体经济:经济世界模型的系统蓝图

Jiale Han, Xiang Li, Jing Qian, Wenyuan Gu, Pin Gao, Ye Luo, Hongyuan Zha, Dacheng Tao, Benyou Wang, Lin William Cong

机构 * Shenzhen Loop Area Institute(深圳河套学院) School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院) University of Hong Kong(香港大学) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出经济世界模型(EWM)的六级能力阶梯实施蓝图,旨在加速可作为人类决策沙箱与AI智能体基础的下一代高保真经济模拟环境开发。

Comments Project page: https://github.com/FreedomIntelligence/Awesome-Economic-World-Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05439 2026-08-07 cs.AI cs.LG 新提交 73%

SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications

SCP-NL2TL:用于自然语言到时间逻辑规范的带语义验证的选择性共形预测

Yixuan Wang, Licheng Luo, Yu Fu, Kaidi Xu, Yue Dong, Mingyu Cai

机构 * University of California, Riverside(加州大学河滨分校) City University of Hong Kong(香港城市大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 该研究提出SCP-NL2TL框架,结合选择性共形预测与语义验证,可生成自然语言到时间逻辑的规范并判断其可靠性,在三类逻辑上提升了翻译可靠性与鲁棒性,为可信自然语言接口提供了基础。

详情

展开后加载摘要…

URL PDF HTML 收藏