FormalAlign: Automated Alignment Evaluation for Autoformalization
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 23 pages, 13 tables, 3 figures
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 23 pages, 13 tables, 3 figures
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Working in progress
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 安全评测 :alignment(title,abstract);trustworthy(abstract)
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to ACL 2024
专题命中 安全评测 :safety(title,abstract);alignment(abstract)
Comments 18 pages, 3 figures, accepted for publication at the International Symposium on Leveraging Applications of Formal Methods (ISoLA 2024)
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Dataset and models are available at https://github.com/jihaonew/MM-Instruct
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments 9 pages (excluding references), accepted to ACL 2024 Main Conference
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted at ACL 2024 Findings
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to ACL
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments v3 is the camera-ready version for AAAI-24 Workshop on Public Sector LLMs: Algorithmic and Sociotechnical Design. 7 pages of main content, 1 page of references, 3 pages of appendices, and 7 figures. Our full prompts are released in the repo: https://github.com/zowiezhang/HVAE
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments A preprint work. Benchmark link: https://github.com/whitzard-ai/jade-db. Website link: https://whitzard-ai.github.io/jade.html
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments To be published in GEM workshop. Conference on Empirical Methods in Natural Language Processing (EMNLP). 2023
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments main paper p.1-29, 5 figures, 2 tables
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments 20 pages, 5 figures
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 安全评测 :safety(title,abstract);AI safety(abstract)
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments Full version (with appendix) of the accepted paper in 36th AAAI Conference on Artificial Intelligence 2022
专题命中 安全评测 :trustworthy(title);safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments This submission has been accepted in ACM FAT* 2020 Conference
大语言模型对齐的可验证性:行为评估下的规范不可区分性
机构 * Igor Santos-Grueiro(独立研究者)
专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG
AI总结 本文研究了大语言模型对齐的可验证性,指出行为评估无法唯一确定潜在对齐,提出了规范不可区分性概念,并通过实验验证了在评估意识下行为基准的局限性。
Comments 10 pages. Theoretical analysis of behavioral alignment evaluation
MTMCS-Bench: 多轮对话中多模态大语言模型上下文安全性的评估
机构 * University of Notre Dame(诺丁汉大学) ; University of California, Los Angeles(加州大学洛杉矶分校) ; Georgia Institute of Technology(佐治亚理工学院) ; University of Montreal(蒙特利尔大学)
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI
AI总结 MTMCS-Bench评估多模态大语言模型在多轮对话中的上下文安全性,揭示了安全与效用之间的权衡及现有防护措施的不足。
Comments A benchmark of realistic images and multi-turn conversations that evaluates contextual safety in MLLMs under two complementary settings
机构 * Jesse C. Cresswell(独立研究者)
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG;alignment(comments)
Comments Presented at the ICLR 2025 Workshop on Bidirectional Human-AI Alignment
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG
Comments ICLR 2025 Workshop on Representational Alignment (Re-Align)
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI
Comments The first work to benchmark Large Multimodal Models in safety insight on social media
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY
Comments Updates: Added disclaimer about USA's recent U-turn on Trustworthy AI Executive Order. Improved Fairness and Group size in section 6.5. Fixed typos. Added a few new references. Updated title
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI;alignment(comments)
Comments Accepted at AAAI 2025 (Special Track on AI Alignment)
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY
Comments Accepted at [TAS '23]{First International Symposium on Trustworthy Autonomous Systems}