arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9311 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9311 篇

2206.11981 2022-06-27 cs.AI cs.CY 81%

Never trust, always verify : a roadmap for Trustworthy AI?

Lionel Nganyewou Tidjon, Foutse Khomh

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.06591 2022-06-22 cs.SI cs.CY cs.LG 81%

An Interpretable Graph-based Mapping of Trustworthy Machine Learning Research

Noemi Derzsy, Subhabrata Majumdar, Rajat Malik

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

Comments Accepted in CompleNet-2021 (oral presentation)

Journal ref In: Teixeira, A.S., Pacheco, D., Oliveira, M., Barbosa, H., Gonçalves, B., Menezes, R. (eds) Complex Networks XII. CompleNet-Live 2021. Springer Proceedings in Complexity. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.13229 2022-06-14 cs.CV cs.AI cs.LG cs.SI 81%

Network-level Safety Metrics for Overall Traffic Safety Assessment: A Case Study

Xiwen Chen, Hao Wang, Abolfazl Razi, Brendan Russo, Jason Pacheco, John Roberts, Jeffrey Wishart, Larry Head, Alonso Granados Baca

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01167 2022-05-27 cs.AI cs.LG 81%

Trustworthy AI: From Principles to Practices

Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, Bowen Zhou

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.13373 2022-05-10 cs.HC cs.AI cs.LG 81%

Trustworthy AI and Robotics and the Implications for the AEC Industry: A Systematic Literature Review and Future Potentials

Newsha Emaminejad, Reza Akhavian

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.03413 2022-04-27 cs.AI cs.CY 81%

Systems Challenges for Trustworthy Embodied Systems

Harald Rueß

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments 57 pages, 7 figures, 3 tables. Public project deliverable, fortiss whitepaper

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.06046 2022-04-13 cs.LG cs.AI cs.CR 81%

Information Theoretic Evaluation of Privacy-Leakage, Interpretability, and Transferability for Trustworthy AI

Mohit Kumar, Bernhard A. Moser, Lukas Fischer, Bernhard Freudenthaler

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments arXiv admin note: text overlap with arXiv:2105.04615, arXiv:2104.07060

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06228 2022-03-15 cs.CL cs.AI 81%

CoDA21: Evaluating Language Understanding Capabilities of NLP Models With Context-Definition Alignment

Lütfi Kerem Senel, Timo Schick, Hinrich Schütze

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments To appear in ACL 2022, 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.07447 2022-03-11 cs.CY cs.AI 81%

Trustworthy Autonomous Systems (TAS): Engaging TAS experts in curriculum design

Mohammad Naiseh, Caitlin Bentley, Sarvapali D. Ramchurn

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03718 2022-03-09 cs.CY cs.AI 81%

Towards User-Centered Metrics for Trustworthy AI in Immersive Cyberspace

Pengyuan Zhou, Benjamin Finley, Lik-Hang Lee, Yong Liao, Haiyong Xie, Pan Hui

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.06704 2022-02-16 cs.LG cs.CL 81%

Food safety risk prediction with Deep Learning models using categorical embeddings on European Union data

Alberto Nogales, Rodrigo Díaz Morón, Álvaro J. García-Tejedor

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.LG

Comments 20 pages,8 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07773 2021-12-16 cs.AI cs.CY 81%

Filling gaps in trustworthy development of AI

Shahar Avin, Haydn Belfield, Miles Brundage, Gretchen Krueger, Jasmine Wang, Adrian Weller, Markus Anderljung, Igor Krawczuk, David Krueger, Jonathan Lebensold, Tegan Maharaj, Noa Zilberman

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Journal ref Science (2021) Vol 374, Issue 6573, pp. 1327-1329

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06912 2021-11-01 cs.LG cs.AI 81%

Blockchain-based Trustworthy Federated Learning Architecture

Sin Kit Lo, Yue Liu, Qinghua Lu, Chen Wang, Xiwei Xu, Hye-Young Paik, Liming Zhu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.11844 2021-08-27 cs.CY cs.AI 81%

AI at work -- Mitigating safety and discriminatory risk with technical standards

Nikolas Becker, Pauline Junginger, Lukas Martinez, Daniel Krupka, Leonie Beining

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.07334 2021-07-16 cs.HC cs.CR cs.CY cs.LG 81%

Tournesol: A quest for a large, secure and trustworthy database of reliable human judgments

Lê-Nguyên Hoang, Louis Faucon, Aidan Jungo, Sergei Volodin, Dalia Papuc, Orfeas Liossatos, Ben Crulis, Mariame Tighanimine, Isabela Constantin, Anastasiia Kucherenko, Alexandre Maurer, Felix Grimberg, Vlad Nitu, Chris Vossen, Sébastien Rouault, El-Mahdi El-Mhamdi

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

Comments 27 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.09051 2021-03-17 cs.CY cs.AI cs.SE 81%

Exploring the Assessment List for Trustworthy AI in the Context of Advanced Driver-Assistance Systems

Markus Borg, Joshua Bronson, Linus Christensson, Fredrik Olsson, Olof Lennartsson, Elias Sonnsjö, Hamid Ebabi, Martin Karsberg

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments Accepted for publication in the Proc. of the 2nd Workshop on Ethics in Software Engineering Research and Practice

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.05620 2021-01-15 cs.LG cs.CY 81%

A Framework for Assurance of Medication Safety using Machine Learning

Yan Jia, Tom Lawton, John McDermid, Eric Rojas, Ibrahim Habli

专题命中 安全评测 :safety(title,abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.15911 2021-01-06 cs.AI cs.LG stat.ML 81%

The role of explainability in creating trustworthy artificial intelligence for health care: a comprehensive survey of the terminology, design choices, and evaluation strategies

Aniek F. Markus, Jan A. Kors, Peter R. Rijnbeek

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Journal ref Journal of Biomedical Informatics, 113 (2021), 103655

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06373 2020-12-14 cs.LG cs.AI cs.AR cs.NE stat.ML 81%

Hardware Beyond Backpropagation: a Photonic Co-Processor for Direct Feedback Alignment

Julien Launay, Iacopo Poli, Kilian Müller, Gustave Pariente, Igor Carron, Laurent Daudet, Florent Krzakala, Sylvain Gigan

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 6 pages, 2 figures, 1 table. Oral at the Beyond Backpropagation Workshop, NeurIPS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.02272 2020-11-05 cs.CY cs.CR cs.CV cs.LG 81%

Trustworthy AI

Richa Singh, Mayank Vatsa, Nalini Ratha

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG

Comments ACM CODS-COMAD 2021 Tutorial

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.07768 2020-09-01 cs.SE cs.AI cs.CY 81%

Opening the Software Engineering Toolbox for the Assessment of Trustworthy AI

Mohit Kumar Ahuja, Mohamed-Bachir Belaid, Pierre Bernabé, Mathieu Collet, Arnaud Gotlieb, Chhagan Lal, Dusica Marijan, Sagar Sen, Aizaz Sharif, Helge Spieker

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY

Comments 1st International Workshop on New Foundations for Human-Centered AI @ ECAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.01263 2020-08-05 cs.SE cs.AI cs.LG cs.SY eess.SY 81%

Safety design concepts for statistical machine learning components toward accordance with functional safety standards

Akihisa Morikawa, Yutaka Matsubara

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.03964 2020-07-09 math.OC cs.AI cs.LG 81%

Responsive Safety in Reinforcement Learning by PID Lagrangian Methods

Adam Stooke, Joshua Achiam, Pieter Abbeel

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments ICML 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20691 2026-07-24 cs.CV cs.AI 新提交 80%

Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis

用于可靠乳腺超声诊断的空间基础概念瓶颈模型

Moshiur Rahman Tonmoy, Dunren Che, Haitham Y. Adarbah, Afzel Noore

专题命中 安全评测 :trustworthy(title,comments);alignment(abstract);分类 cs.AI

AI总结 研究乳腺超声诊断中概念瓶颈模型可信度受监督限制问题,提出空间基础概念瓶颈模型(SG-CBM),利用病变轮廓弱监督,通过导出特定区域训练概念图,经交叉验证等提升诊断指标与概念证据空间对齐,强调数据质量监督设计及可信度验证的必要。

Comments Accepted to the Workshop on Data Quality Aware, High-Performance, and Trustworthy AI Systems for Healthcare at IEEE/ACM CHASE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25256 2026-06-05 cs.AI 80%

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

谁的对齐?比较不同组织决策情境下的大语言模型过程对齐

Niklas Weller, Emilio Barkett

机构 * University of Cambridge(剑桥大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出一种决策策略捕获方法测量过程对齐,发现LLM在ECHR第6条决策中过程对齐与输出准确性高度相关,但在德国消费信贷决策中关系消失,揭示了多元对齐挑战。

Comments Accepted to Pluralistic Alignment Workshop @ ICML 2026, Seoul, South Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19869 2026-05-20 cs.CV cs.AI 80%

Passive Construction Site Safety Monitoring via Persona-Scaffolded Adversarial Chain-of-Thought VLM Verification

通过基于人设的对抗性链式思考视觉语言模型验证实现被动施工现场安全监控

Ananth Sriram, Neel Mokaria, Rajveer Singh

机构 * Department of Computer Science, University of Maryland, College Park, MD, USA(大学马里兰学院计算机科学系,马里兰州科利尔帕克,MD,美国)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

AI总结 本文提出了一种被动的施工现场安全监控方法,通过三阶段架构处理视频数据,结合细调的YOLO11、SAM 3和Qwen3-VL-8B-Instruct模型,利用基于人设的对抗性链式思考协议提高合规性验证和幻觉控制,主要贡献是第三阶段提示设计,提升了12%的精度。

Comments 10 pages, 4 figures. First place, Ironsite.ai Spatial Intelligence Hackathon, University of Maryland, February 2026. Code available at https://github.com/ananthsriram1/ironsite-hackathon-project-safety_assistant

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28087 2026-05-12 cs.LO cs.AI 80%

Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles

迈向基于法律和安全原则的神经符号因果规则合成、验证与评估

Zainab Rehan, Christian Medeiros Adriano, Sona Ghahremani, Holger Giese

机构 * Hasso Plattner Institute \ of Potsdam Prof.-Dr.-Helmert Str. 2-3, D-14482 Potsdam, Germany Hasso Plattner Institute \ of Potsdam

专题命中 安全评测 :safety(title,abstract);分类 cs.AI;trustworthy(journal_ref)

AI总结 本文提出一种神经符号因果框架,结合一阶逻辑抽象树、结构因果模型和深度强化学习,通过Meta层缓解目标误指定问题,实现可扩展的规则维护。

Journal ref Neurosymbolic eXplainable Trustworthy Systems @ AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19457 2026-04-22 cs.AI 80%

Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents

四轴决策对齐用于长周期企业AI代理

Vasundra Srininvasan

机构 * Vasundra Srinivasan

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出四轴对齐框架,用于评估长周期企业AI代理的决策行为,涵盖事实精度、推理连贯性、合规重建和校准回避,通过实验揭示了决策对齐的重要性。

Comments 21 pages, 5 figures, 8 tables. PDFLaTeX. Code and artifacts: https://github.com/vasundras/decision-alignment-long-horizon-agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18850 2026-04-22 cs.HC cs.AI cs.SI 80%

The Triadic Loop: A Framework for Negotiating Alignment in AI Co-hosted Livestreaming

三元循环:一种协商AI共播直播中对齐的框架

Katherine Wang, Nadia Berthouze, Aneesha Singh

机构 * University College London(伦敦大学学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出三元循环框架,用于协商AI共播直播中的对齐问题,通过三者之间的双向适应过程,解决多用户社交环境中的动态反馈循环问题。

Comments 6 pages, 1 figure, Proceedings the Human-AI Interaction Alignment Workshop at CHI 2026 (CHI26 BiAlign Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01816 2026-01-06 cs.AI 80%

Admissibility Alignment

可接受性对齐

Chris Duffey

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 本文提出MAP-AI架构,通过蒙特卡洛方法评估决策策略的可接受性,实现AI对齐的动态决策理论属性。

Comments 24 pages, 2 figures, 2 tables.. Decision-theoretic alignment under uncertainty

详情

展开后加载摘要…

URL PDF HTML 收藏